Fusion proteins
Patent Information
- Application Number
- JP2024549597
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-05
- Filing Date
- 2023-02-23
- Publication Date
- 2026-02-13
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 313,151, filed February 23, 2022, and U.S. Provisional Patent Application No. 63 / 478,690, filed January 5, 2023, each of which is incorporated by reference in its entirety.
[0002] Sequence Listing The contents of the electronic sequence listing (MTG-009WO_SL.xml, size: 223,389 bytes, and creation date: February 23, 2023) are incorporated herein by reference in their entirety. [Background technology]
[0003] Cas enzymes, along with their associated clustered regularly interspaced short palindromic repeats (CRISPR) guide ribonucleic acid (RNA), appear to be widespread (~45% in bacteria, ~84% in archaea) components of the prokaryotic immune system, helping to protect such microorganisms from non-self nucleic acids, such as infectious viruses and plasmids, by CRISPR-RNA-guided nucleic acid cleavage. While deoxyribonucleic acid (DNA) elements encoding CRISPR RNA elements may be relatively conserved in structure and length, their CRISPR-associated (Cas) proteins are highly diverse and contain a wide variety of nucleic acid-interacting domains. Although CRISPR DNA elements have been observed as early as 1987, the programmable endonuclease cleavage capabilities of CRISPR / Cas complexes have only been recognized relatively recently, leading to the use of recombinant CRISPR / Cas systems in a variety of DNA engineering and gene editing applications. Summary of the Invention
[0004] In some aspects, the disclosure provides a fusion protein comprising (a) a class 2 V-type Cas effector and (b) a functional domain comprising a DNA binding domain (DBD) or a chromatin modulating domain (CMD). In some embodiments, the functional domain is derived from human histone 1 central globular domain, HMGN1, cbx5, or Saccharolobus solfataricus sso7d. In some embodiments, the Cas effector is derived from the CAST locus. In some embodiments, the Cas effector comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to a Cas domain of any one of SEQ ID NOs: 113-116, or a variant thereof. In some embodiments, the functional domain comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 109-112, or a variant thereof. In some embodiments, the fusion protein comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 113-116, or a variant thereof.
[0005] In some aspects, the disclosure provides a fusion protein comprising: (a) a TniQ protein; and (b) a functional domain comprising a DNA binding domain (DBD) or a chromatin modulating domain (CMD). In some embodiments, the TniQ protein is derived from the CAST locus. In some embodiments, the TniQ protein comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the TniQ domain of any one of SEQ ID NOs: 117-120, or a variant thereof. In some embodiments, the functional domain comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 109-112, or a variant thereof. In some embodiments, the fusion protein comprises a sequence having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 117-120, or a variant thereof.
[0006] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site, the system comprising: a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a Tn7-type transposase complex; a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleotide sequence; and a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit. In some embodiments, the cargo nucleotide sequence is flanked by a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further comprises a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the target nucleic acid comprises a PAM sequence compatible with the Cas effector complex flanking the target nucleic acid site. In some embodiments, the PAM sequence is located 3' of the target nucleic acid site. In some embodiments, the PAM sequence is located 5' of the target nucleic acid site. In some embodiments, the engineered guide polynucleotide is configured to bind to the class 2 V-type Cas effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NO: 1, 12, 16, 20-30, 64, 80-85, and 200, or a variant thereof. In some embodiments, the TnsB subunit comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NO: 2, 13, 17, or 65, or a variant thereof. In some embodiments, the Tn7-type transposase complex comprises at least one or at least two, three polypeptides comprising a sequence having at least 80% identity to any one of SEQ ID NO: 3-4, 14-15, 18-19, or 66-67, or a variant thereof.In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, or 104-105, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, or 96-103, or a variant thereof. In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs: 9, 11, 36-38, 76, or 78, or a variant thereof. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NOs: 8, 10, 39-44, 77, 79, or 93, or a variant thereof. In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by polynucleotide sequences comprising less than about 10 kilobases.
[0007] In some aspects, the present disclosure provides a method for translocating a cargo nucleotide sequence into a target nucleic acid site comprising a target nucleotide sequence, comprising expressing in a cell or introducing into a cell a system of any of the aspects or embodiments described herein.
[0008] In some aspects, the disclosure provides a method for transposing a cargo nucleotide sequence into a target nucleic acid site, comprising contacting a first double-stranded nucleic acid comprising the cargo nucleotide sequence with a Cas effector complex comprising a class 2 V-type Cas effector and at least one engineered guide polynucleotide configured to hybridize to the target nucleotide sequence, a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit, and a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the cargo nucleotide sequence is flanked by a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the target nucleic acid comprises a PAM sequence compatible with the Cas effector complex flanking the target nucleic acid site. In some embodiments, the PAM sequence is located 3' of the target nucleic acid site. In some embodiments, the engineered guide polynucleotide is configured to bind to the class 2 V-type Cas effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NO: 1, 12, 16, 20-30, 64, 80-85, and 200, or a variant thereof. In some embodiments, the TnsB subunit comprises a polypeptide comprising a sequence having at least 80% identity to SEQ ID NO: 2, 13, 17, or 65, or a variant thereof. In some embodiments, the Tn7-type transposase complex comprises at least one or at least two polypeptides comprising a sequence having at least 80% identity to any one of SEQ ID NO: 3-4, 14-15, 18-19, or 66-67, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, or 104-105, or a variant thereof.In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NO: 9, 11, 36-38, 76, or 78, or a variant thereof. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NO: 8, 10, 39-44, 77, 79, or 93, or a variant thereof. In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.
[0009] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site, the system comprising: a first double-stranded nucleic acid comprising a cargo nucleotide sequence configured to interact with a Tn7-type transposase complex; a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleotide sequence; and a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components. and a Tn7-type transposase complex comprising a polypeptide having a sequence with at least 80% sequence identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200, or a variant thereof, or (b) the Tn7-type transposase complex comprises a TnsB, TnsC, or TniQ component having a sequence with at least 80% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, and 65-67, or a variant thereof. In some embodiments, the transposase complex is non-covalently linked to the Cas effector complex. In some embodiments, the transposase complex is covalently linked to the Cas effector complex. In some embodiments, the transposase complex is fused to the Cas effector complex in a single polypeptide. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide having a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200, or a variant thereof. In some embodiments, the Tn7-type transposase complex comprises a TnsB, TnsC, or TniQ component having a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, and 65-67, or a variant thereof. In some embodiments, the class 2 V-type Cas effector is a Cas12k effector.In some embodiments, the cargo nucleotide sequence is flanked by a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the system further comprises a second double-stranded nucleic acid comprising the target nucleic acid site. In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' or 3' of the target nucleic acid site. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the engineered guide polynucleotide is configured to bind to the class 2 V-type Cas effector. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, or 104-105, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of the non-degenerate nucleotides of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, or 96-103, or a variant thereof. In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, or 78, or a variant thereof. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, or 93. In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.In some embodiments, (a) the class 2 V-type Cas effector comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, or 85, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 9, 11, 36, 37, or 38, or a variant thereof; or (c) the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 39, 40, 41, 42, 43, 44, or 93, or a variant thereof. or (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO:6, or a variant thereof, or (ii) a sequence having at least 80% identity to the non-degenerate nucleotides of any one of SEQ ID NOs:5, 45-63, 68-75, or 96-103, or a variant thereof; (e) the TnsB, TnsC, and TniQ components comprise polypeptides having at least 80% identity to SEQ ID NOs:2-4, or a variant thereof; or (f) the PAM sequence comprises SEQ ID NO:31. In some embodiments, (a) the class 2 V-type Cas effector comprises a sequence having at least 80% sequence identity to SEQ ID NO: 12, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 76, or a variant thereof; (c) the right recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NO: 77, or a variant thereof; (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO: 32 or 104, or (ii) a sequence having at least 80% identity to any one of the non-degenerate nucleotides of SEQ ID NO: 107 or 102, or a variant thereof; or (e) the TnsB, TnsC, and TniQ components comprise a polypeptide having a sequence having at least 80% identity to SEQ ID NO: 13-15, or a variant thereof.In some embodiments, (a) the class 2 V-type Cas effector comprises a sequence having at least 80% sequence identity to SEQ ID NO: 16, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to SEQ ID NO: 78, or a variant thereof; (c) the right recombinase sequence comprises a sequence having at least 80% identity to SEQ ID NO: 79, or a variant thereof; (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO: 33 or 105, or (ii) a sequence having at least 80% identity to any one of the non-degenerate nucleotides of SEQ ID NO: 108 or 103, or a variant thereof; or (e) the TnsB, TnsC, and TniQ components comprise a polypeptide having a sequence having at least 80% identity to SEQ ID NO: 17-19, or a variant thereof.
[0010] In some aspects, the disclosure provides an engineered nuclease system comprising an endonuclease comprising a RuvC domain, the endonuclease being derived from an uncultured microorganism, the endonuclease being a class 2 VK-type Cas effector having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200, or a variant thereof, and an engineered guide RNA, the engineered guide RNA being configured to form a complex with the endonuclease, the engineered guide RNA comprising a spacer sequence configured to hybridize within a target nucleic acid sequence. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, or 104-105, or a variant thereof. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% identity to a non-degenerate nucleotide of any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, or 96-103, or a variant thereof. In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with the Cas effector complex adjacent to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' of the target nucleic acid site. In some embodiments, the PAM sequence comprises SEQ ID NO: 31.In some embodiments, (a) the class 2 VK-type Cas effector comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, or 85, or a variant thereof; (b) the left recombinase sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 9, 11, 36, 37, or 38, or a variant thereof; or (c) the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 39, 40, 41, 42, 43, 44, or 93, or a variant thereof. or (d) the engineered guide polynucleotide comprises (i) a sequence having at least 80% sequence identity to at least 46-80 nucleotides of SEQ ID NO:6, or a variant thereof, or (ii) a sequence having at least 80% identity to the non-degenerate nucleotides of any one of SEQ ID NOs:5, 45-63, 68-75, or 96-103, or a variant thereof; (e) the TnsB, TnsC, and TniQ components comprise polypeptides having at least 80% identity to SEQ ID NOs:2-4, or a variant thereof; or (f) the PAM sequence comprises SEQ ID NO:31.
[0011] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0012] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a small prokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site; a Tn7-type transposase complex configured to bind to the Cas effector complex and comprising TnsB, TnsC, and TniQ components; a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex and comprising a cargo nucleotide sequence; and a functional domain comprising a DNA binding domain (DBD) or a chromatin modulating domain (CMD).
[0013] In some embodiments, the Cas effector complex is non-covalently linked to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is covalently linked to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7-type transposase complex.
[0014] In some embodiments, the cargo nucleotide sequence is flanked by left and right transposase recognition sequences recognized by a Tn7-type transposase complex. In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.
[0015] In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' to the target nucleic acid site.
[0016] In some embodiments, the class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.
[0017] In some embodiments, the TnsB component comprises a polypeptide having a sequence having at least 80% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having a sequence having at least 90% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the TnsB component comprises a polypeptide having a sequence having at least one of SEQ ID NOs: 2, 13, 17, and 65. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some embodiments, the Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some embodiments, the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising any one of the sequences of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.
[0018] In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 165.
[0019] In some embodiments, the functional domain is derived from human histone 1 central globular domain, HMGN1, cbx5, or Saccharolobus solfataricus sso7d. In some embodiments, the functional domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 109-112. In some embodiments, a class 2 V-type Cas effector is fused to the functional domain to form a fusion protein. In some embodiments, the fusion protein comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 113-116.
[0020] In some embodiments, the Tn7-type transposase complex comprises a TniQ protein. In some embodiments, the TniQ protein is fused to a functional domain to form a fusion protein. In some embodiments, the TniQ protein comprises a sequence having at least about 80% sequence identity with any one of the TniQ domains of SEQ ID NOs: 117-120.
[0021] In some embodiments, the small prokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 167-169. In some embodiments, the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 161-163.
[0022] In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.
[0023] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising a Cas effector complex comprising a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site, the Cas effector complex comprising a polypeptide comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200, and an engineered guide polynucleotide configured to bind to the Cas effector complex. and a Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components, the TnsB, TnsC, and TniQ components comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, and 65-67; a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising a cargo nucleotide sequence; and a functional domain comprising a DNA-binding domain (DBD) or a chromatin modulating domain (CMD).
[0024] In some embodiments, the Cas effector complex is non-covalently linked to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is covalently linked to the Tn7-type transposase complex. In some embodiments, the Cas effector complex is fused to the Tn7-type transposase complex.
[0025] In some embodiments, the cargo nucleotide sequence is flanked by left and right transposase recognition sequences recognized by a Tn7-type transposase complex. In some embodiments, the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some embodiments, the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.
[0026] In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO:31.
[0027] In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' to the target nucleic acid site.
[0028] In some embodiments, the class 2 V-type Cas effector is a Cas12k effector. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some embodiments, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence of any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.
[0029] In some embodiments, the TnsB, TnsC, or TniQ component comprises a sequence having at least 90% sequence identity to any one of SEQ ID NOs: 2-4, 13-15, 17-19, and 65-67.
[0030] In some embodiments, the TnsB, TnsC, or TniQ component comprises any one of SEQ ID NOs: 2-4, 13-15, 17-19, and 65-67. In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.
[0031] In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 165.
[0032] In some embodiments, the functional domain is derived from human histone 1 central globular domain, HMGN1, cbx5, or Saccharolobus solfataricus sso7d. In some embodiments, the functional domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 109-112. In some embodiments, a class 2 V-type Cas effector is fused to the functional domain to form a fusion protein. In some embodiments, the fusion protein comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 113-116.
[0033] In some embodiments, the Tn7-type transposase complex comprises a TniQ protein. In some embodiments, the TniQ protein is fused to a functional domain to form a fusion protein. In some embodiments, the TniQ protein comprises a sequence having at least about 80% sequence identity with any one of the TniQ domains of SEQ ID NOs: 117-120.
[0034] In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit S15. In some embodiments, the small prokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 167-169. In some embodiments, the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 161-163.
[0035] In some embodiments, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.
[0036] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising a Cas effector complex comprising: i) a class 2 V-type Cas effector comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1, 81, 82, 83, and 85; and ii) an engineered guide polynucleotide comprising at least 80% identity to any one of SEQ ID NOs: 5, 6, 45-63, 68-75, and 96-103; and a Tn7-type transposase complex configured to bind to the Cas effector complex, comprising TnsB, TnsC, and TniQ components, wherein the TnsB, TnsC, and the TniQ component comprises a Tn7-type transposase complex comprising a sequence having at least 80% identity to any one of SEQ ID NOs:2-4; a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase sequence comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs:9, 11, 36, 37, and 38, a cargo nucleotide sequence, and a right recombinase sequence comprising a sequence having at least 80% identity to any one of SEQ ID NOs:8, 39-44, and 93; and a functional domain comprising a DNA binding domain (DBD) or a chromatin modulating domain (CMD).
[0037] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising a Cas effector complex configured to hybridize to the target nucleic acid site, the Cas effector complex comprising: i) a class 2 V-type Cas effector comprising a sequence having at least 80% sequence identity to SEQ ID NO: 12; and ii) an engineered guide polynucleotide comprising one having at least 80% identity to any one of SEQ ID NOs: 32, 102, 104, and 107; and a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components. The system includes a Tn7-type transposase complex, wherein the TnsB, TnsC, or TniQ component comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 13-15; a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase sequence comprising a sequence having at least 80% sequence identity to SEQ ID NO: 76, a cargo nucleotide sequence, and a right recombinase sequence comprising a sequence having at least 80% identity to SEQ ID NO: 77; and a functional domain comprising a DNA binding domain (DBD) or a chromatin modulating domain (CMD).
[0038] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising a Cas effector complex configured to hybridize to the target nucleic acid site, the Cas effector complex comprising: i) a class 2 V-type Cas effector comprising a sequence having at least 80% sequence identity to SEQ ID NO: 16; and ii) an engineered guide polynucleotide comprising one having at least 80% identity to any one of SEQ ID NOs: 33, 103, 105, and 108; and a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising TnsB, TnsC, and TniQ components. The system includes a Tn7-type transposase complex, wherein the TnsB, TnsC, or TniQ component comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 17-19; a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex, the double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase sequence comprising a sequence having at least 80% sequence identity to SEQ ID NO: 78, a cargo nucleotide sequence, and a right recombinase sequence comprising a sequence having at least 80% identity to SEQ ID NO: 79; and a functional domain comprising a DNA binding domain (DBD) or a chromatin modulating domain (CMD).
[0039] In some embodiments, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex. In some embodiments, the PAM sequence comprises SEQ ID NO: 31. In some embodiments, the PAM sequence is located about 50 to about 70 base pairs from the target nucleic acid site. In some embodiments, the PAM sequence is located 3' to the target nucleic acid site. In some embodiments, the PAM sequence is located 5' to the target nucleic acid site.
[0040] In some embodiments, the functional domain is derived from human histone 1 central globular domain, HMGN1, cbx5, or Saccharolobus solfataricus sso7d. In some embodiments, the functional domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 109-112. In some embodiments, a class 2 V-type Cas effector is fused to the functional domain to form a fusion protein. In some embodiments, the fusion protein comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 113-116.
[0041] In some embodiments, the Tn7-type transposase complex comprises a TniQ protein. In some embodiments, the TniQ protein is fused to a functional domain to form a fusion protein. In some embodiments, the TniQ protein comprises a sequence having at least about 80% sequence identity with any one of the TniQ domains of SEQ ID NOs: 117-120.
[0042] In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit S15. In some embodiments, the small prokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 167-169. In some embodiments, the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 161-163.
[0043] In some aspects, the disclosure provides a system for transposing a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, the system comprising: a Cas effector complex comprising a class 2 V-type Cas effector, a small prokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide, wherein the engineered guide polynucleotide can hybridize to the target nucleic acid; a Tn7-type transposase complex operably linked to the Cas effector complex and comprising TnsB, TnsC, and TniQ components; a double-stranded nucleic acid comprising, in 5' to 3' order, a left recombinase recognition sequence, a cargo nucleotide sequence, and a right recombinase recognition sequence, wherein the left recombinase recognition sequence and the right recombinase recognition sequence can be recognized by the Tn7-type transposase complex; and a functional domain comprising a DNA binding domain (DBD) or a chromatin modulating domain (CMD).
[0044] In some aspects, the disclosure provides an engineered nuclease system comprising an endonuclease comprising a RuvC domain, the endonuclease being derived from an uncultured microorganism and being a class 2 VK-type Cas effector having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200; and an engineered guide RNA configured to form a complex with the endonuclease and comprising a spacer sequence configured to hybridize within a target nucleic acid sequence.
[0045] In some embodiments, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some embodiments, the engineered guide polynucleotide comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 165.
[0046] In some aspects, the present disclosure provides a method for translocating a cargo nucleotide sequence into a target nucleic acid site comprising introducing a system of the present disclosure into a cell.
[0047] In some aspects, the disclosure provides a cell comprising the system of the disclosure. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell.
[0048] In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is an A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the cell is an engineered cell. In some embodiments, the cell is a stable cell. [Brief description of the drawings]
[0049] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings (also referred to herein as "Figure" and "FIG.").
[0050] [Figure 1] Illustrates typical organizations of different classes and types of CRISPR / Cas loci. [Diagram 2]FIG. 1 illustrates the structure of a natural class 2 type II crRNA / tracrRNA pair as shown, for example, for Cas9, compared to a hybrid sgRNA in which the crRNA and tracrRNA are joined. [Diagram 3] The two pathways found in Tn7 and Tn7-like elements are illustrated. [Figure 4A] The genomic context of the MG64 family's type V Tn7 CAST is illustrated. Figure 4A illustrates that the MG64-1 CAST system contains a CRISPR array (CRISPR repeats), a type V nuclease, and three predicted transposase protein sequences. A tracrRNA was predicted within the intergenic region between the CAST effector array and the CRISPR array. Bottom: Multiple sequence alignment of the catalytic domain of transposase TnsB. Catalytic residues are indicated by boxes. [Figure 4B] Figure 4B illustrates the genomic context of MG64 family type V Tn7 CAST. Figure 4B illustrates that two transposon ends were predicted for the MG64-1 CAST system. [Diagram 5] The predicted structure of the corresponding sgRNA of the CAST system described herein is illustrated. Panel A of Figure 5 (left) shows the predicted MG64-1 tracrRNA and crRNA duplex in the repeat-repeat suppressor stem. The loop was truncated and a tetraloop of GAAA was added to the stem-loop structure to produce the designed sgRNA shown in Panel B of Figure 5 (right). [Figure 6] Illustrated are the results of transposition reactions targeted with a plasmid library consisting of NNNNNNNN at the 5' of the target spacer sequence. Reaction #1 shows the presence of the targeted library, #2 shows the presence of the donor fragment in both transposition reactions, and #3-5 show the sg-specific PCR bands corresponding to proper transposition reactions. [Figure 7A]Illustrating the results of Sanger sequencing. Figure 7A shows Sanger sequencing of the donor-target junction on the left end (LE) of the transposon in a LE transposition reaction closer to the PAM. The expected sequence is at the top of the panel, with the predicted transposition event 61 bp away from the PAM. The top chromatogram is the sequencing result starting within the donor fragment. A clear signal is seen on the far right up to the donor / target junction (dotted line). This shows a mixture of transposition products. The bottom chromatogram of the panel is sequencing from the target to the donor / target junction. The signal from the left is a clear signal up to the junction. [Figure 7B] The results of Sanger sequencing are illustrated. Figure 7B shows Sanger sequencing of the donor-target junction on the right end (RE) of the transposon in the product of the LE closer to the PAM. The expected sequence is at the top of the panel, with the predicted transposition event 61 bp away from the PAM. The top chromatogram is the sequencing result starting within the donor fragment. A clear signal is seen to the left all the way to the donor / target junction (dotted line). [Figure 7C] The results of Sanger sequencing are illustrated in Figure 7C. SeqLogo analysis of NGS of LE events closer to the PAM, showing a very strong preference for NGTN at the PAM motif. [Figure 8] Figure 1 illustrates the phylogenetic gene tree of Cas12k effector sequences. This tree was inferred from a multiple sequence alignment of 64 Cas12k sequences recovered here (orange and black branches) and 229 reference Cas12k sequences from public databases (gray branch). The orange branch represents Cas12k effectors that have been confirmed to be associated with CAST transposon components. [Figure 9]MG64 family CRISPR repeat alignment is shown. Cas12k CAST CRISPR repeat contains the conserved motif 5'-GNNGGNNTGAAAG-3'. In MG64-1, the short repeat-repeat suppressor (RAR) within the CRISPR repeat motif aligns with the tracrRNA. The MG64 RAR motif appears to define the start and end of the tracrRNA (5' end: RAR1 (TTTC), 3' end: RAR2 (CCNNC)). [Figure 10A-1] Illustrates the secondary structure predicted from folding of CRISPR repeats+tracrRNA for the MG64 system. [Figure 10A-2] Illustrates the secondary structure predicted from folding of CRISPR repeats+tracrRNA for the MG64 system. [Figure 10B-1] Illustrates the secondary structure predicted from folding of CRISPR repeats+tracrRNA for the MG64 system. [Figure 10B-2] Illustrates the secondary structure predicted from folding of CRISPR repeats+tracrRNA for the MG64 system. [Figure 11A] The MG64-3 CRISPR locus is illustrated. The tracrRNA is encoded upstream from the CRISPR array, while the transposon ends are encoded downstream (inner black box). Sequences corresponding to partial 3' CRISPR repeats and partial spacers are encoded within the transposon (outer box). Self-matching spacers are encoded outside the transposon ends. [Figure 11B] Illustrating tracrRNA sequence alignments to various CASTs provided herein. Alignment of tracrRNA sequences shows regions of conservation. In particular, the sequence "TGCTTTC" (upper box) at sequence positions 92-98 is suggested to be important for sgRNA tertiary structure and for non-contiguous repeat-repeat suppressor pairing with the crRNA. The hairpin "CYCC(n6)GGRG" (lower box) at positions 265-278 may be important for function, potentially positioning downstream sequences for crRNA pairing. [Figure 12A] The predicted structure of the MG64-1 sgRNA is illustrated. [Figure 12B] The predicted structure of the MG64-3 sgRNA is illustrated. [Figure 12C] The predicted structure of the MG64-5 sgRNA is illustrated. [Figure 13A] Figure 13 illustrates PCR data demonstrating that MG64-1 is active with sgRNA v2-1. Using the protocol described for in vitro target integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. After translation, target DNA, cargo DNA, and sgRNA were added in the reaction buffer. Integration was assayed by PCR across the target / donor junction. Figure 13A illustrates a schematic showing the potential orientations of the integrated donor DNA. PCR reactions 3, 4, 5, and 6 represent each integrated ligation product depending on the orientation in which the donor was integrated at the target site. [Figure 13B] Figure 13B illustrates PCR data demonstrating that MG64-1 is active with sgRNA v2-1. Using the protocol described for in vitro target integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. After translation, target DNA, cargo DNA, and sgRNA were added in the reaction buffer. Integration was assayed by PCR across the target / donor junction. Figure 13B illustrates a gel image of PCR4 (detecting the RE junction to the donor) of transposition showing lane 1) apo (no sgRNA), lane 2) with sgRNA 1, and lane 3) with sgRNA v2-1. [Figure 13C]Figure 13C illustrates PCR data demonstrating that MG64-1 is active with sgRNA v2-1. Using the protocol described for in vitro target integrase activity, the effector protein and its TnsB, TnsC, and TniQ proteins were expressed in an in vitro transcription / translation system. After translation, target DNA, cargo DNA, and sgRNA were added in the reaction buffer. Integration was assayed by PCR across the target / donor junction. Figure 13C illustrates a gel image of PCR5 (detecting the LE junction to the donor) of transposition showing lane 1) apo (no sgRNA), lane 2) with sgRNA 1, and lane 3) with sgRNA v2-1. [Figure 14] Illustrated are PCR reaction 5 (LE proximal to the PAM, top half of the plot) and PCR reaction 4 (RE distal to the PAM, bottom half of the plot) plotted against sequence and distance from the PAM of MG64-1. Analysis of the integration window indicates that 95% of integrations occurring at the spacer PAM site fall within a 10 bp window 58-68 nucleotides away from the PAM. The difference in integration distance between distal and proximal frequencies reflects integration site overlap, i.e., a 3-5 base pair overlap as a result of the staggered nuclease activity of the transposase upon integration. [Figure 15] Illustrated are the results of colony PCR screening of transposition efficiency. After incubation, 18 colony forming units (CFU) were visible on the plates, 8 on plate A (no IPTG, lane labeled as A) and 10 on plate B (100 μM IPTG during recovery, lane labeled as B). All 18 were analyzed by colony PCR, which gave rise to product bands (arrows) indicating a successful transposition reaction. [Figure 16]Sequencing results of selected colony PCR products are shown, confirming that they represent transposition events since they span the junction between the LE and PAM at the engineered target site within the lacZ gene. The minimal LE sequence is shown in blue at the top of the screen (Min LE), while the target and PAM are shown in grey. Some sequence variation is observed in the PCR products, but this variation is expected considering that insertions can occur at variable distances upstream of the PAM. [Figure 17]Figure 17 illustrates the results of testing engineered single guides for 64-1 transposition activity. Black boxes are lanes not relevant to this experiment. Panel A of Figure 17 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-1, lane 4 = sgRNA v1-2, lane 5 = sgRNA v1-3. Panel B of Figure 17 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-1, lane 4 = sgRNA v1-2, lane 5 = sgRNA v1-3. Panel C of Figure 17 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-4, lane 4 = sgRNA v1-6, lane 5 = sgRNA v1-7, lane 6 = sgRNA v1-8, lane 7 = sgRNA v1-9. Panel D of Figure 17 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-4, lane 4 = sgRNA v1-6, lane 5 = sgRNA v1-7, lane 6 = sgRNA v1-8, lane 7 = sgRNA v1-9. Panel E of Figure 17 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-5, lane 4 = skip, lane 5 = sgRNA v1-10. Panel F of Figure 17 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNA v1-5, lane 4 = skip, lane 5 = sgRNA v1-10.Panel G of Figure 17 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNAv1-17, lane 4 = sgRNA v1-18, lane 5 = skip, lane 6 = sgRNA v1-19, lane 7 = skip, lane 8 = sgRNA v1-20. Panel H of Figure 17 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = sgRNAv1-17, lane 4 = sgRNA v1-18, lane 5 = skip, lane 6 = sgRNA v1-19, lane 7 = skip, lane 8 = sgRNA v1-20. [Figure 18]Figure 18 shows the results of testing engineered LE and RE on 64-1 transposition activity. Black boxes are lanes not relevant to this experiment. Panel A of Figure 18 illustrates a gel image of PCR4 of transposition (detecting RE junction to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = LE 86bp, lane 4 = LE 105bp, lane 5 = RE 196bp, lane 6 = RE 242bp, lane 7 = RE internal deletion 50, lane 8 = RE internal deletion 81. Panel B of Figure 18 illustrates a gel image of PCR 5 of the transposition (detecting the LE junction to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = LE 86 bp, lane 4 = LE 105 bp, lane 5 = RE 196 bp, lane 6 = RE 242 bp, lane 7 = RE internal deletion 50, lane 8 = RE internal deletion 81. Panel C of Figure 18 illustrates a gel image of PCR 4 of the transposition (detecting the RE junction to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletion 81 and 178 bp, lane 4 = skip, lane 5 = RE internal deletion 81 and 196 bp, lane 6 = skip, lane 7 = RE internal deletion 81 and 212 bp, lane 8 = skip. Panel D of Figure 18 shows a gel image of PCR 5 of the transposition (detecting the LE junction to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletion 81 and 178bp, lane 4 = skip, lane 5 = RE internal deletion 81 and 196bp, lane 6 = skip, lane 7 = RE internal deletion 81 and 212bp, lane 8 = skip. Panel E of Figure 18 illustrates a gel image of PCR 4 of the transposition (detecting the RE junction to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletion 81 and 178bp + LE 68bp, lane 4 = RE internal deletion 81 and 178bp + LE 86bp, lane 5 = skip, lane 6 = RE internal deletion 81 and 178bp + LE 105bp, lane 7 = skip.Panel F of Figure 18 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = RE internal deletions 81 and 178bp + LE 68bp, lane 4 = RE internal deletions 81 and 178bp + LE 86bp, lane 5 = skip, lane 6 = RE internal deletions 81 and 178bp + LE 105bp, lane 7 = skip. Panel G of Figure 18 illustrates a gel image of PCR 6 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = 0bp overhang, lane 4 = 1bp overhang, lane 5 = 2bp overhang, lane 6 = 3bp overhang, lane 7 = 5bp overhang, lane 8 = 10bp overhang. [Figure 19]Figure 19 illustrates the results of testing engineered CAST components containing NLS for transposition activity. Black boxes are lanes not relevant to this experiment. Panel A of Figure 19 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = skip, lane 6 = NLS-TnsB, lane 7 = skip, lane 8 = TnsB-NLS. Panel B of Figure 19 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = skip, lane 6 = NLS-TnsB, lane 7 = skip, lane 8 = TnsB-NLS. Panel C of Figure 19 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = skip, lane 6 = NLS-TniQ, lane 7 = skip, lane 8 = TniQ-NLS. Panel D of Figure 19 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = skip, lane 6 = NLS-TniQ, lane 7 = skip, lane 8 = TniQ-NLS. Panel E of Figure 19 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = NLS-Cas12k, lane 6 = Cas12k-NLS, lane 7 = NLS-TnsC, lane 8 = TnsC-NLS. Panel F of Figure 19 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = skip, lane 4 = skip, lane 5 = NLS-Cas12k, lane 6 = Cas12k-NLS, lane 7 = NLS-TnsC, lane 8 = TnsC-NLS.Panel G of Figure 19 illustrates a gel image of PCR 4 of transposition (detecting RE junction to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-HA-TnsC, lane 4 = NLS-TnsC-FLAG, lane 5 = NLS-TnsC-HA, lane 6 = NLS-TnsC-Myc, lane 7 = NLS-FLAG-TnsC, lane 8 = NLS-Myc-TnsC. Panel H of Figure 19 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-HA-TnsC, lane 4 = NLS-TnsC-FLAG, lane 5 = NLS-TnsC-HA, lane 6 = NLS-TnsC-Myc, lane 7 = NLS-FLAG-TnsC, lane 8 = NLS-Myc-TnsC. Panel I of Figure 19 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = Cas 2x NLS apo (no sgRNA), lane 4 = Cas 2x NLS holo (+sgRNA). Panel J of Figure 19 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = Cas 2x NLS apo (no sgRNA), lane 4 = Cas 2x NLS holo (+sgRNA). [Figure 20]The engineered CAST-NLS acting as a single suite is illustrated. All lanes have Cas12k-NLS and NLS-TniQ, TnsB, TnsC, and sgRNA unless otherwise stated. Panel A of Figure 20 illustrates a gel image of PCR4 of transposition (detecting RE junction to donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TnsB, lane 4 = TnsB-NLS, lane 5 = NLS-TnsB and NLS-TnsC, lane 6 = TnsB-NLS and NLS-TnsC. Panel B of Figure 20 shows a gel image of PCR 5 of the transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TnsB, lane 4 = TnsB-NLS, lane 5 = NLS-TnsB and NLS-TnsC, lane 6 = TnsB-NLS and NLS-TnsC. [Figure 21]The results of testing Cas effector and TniQ protein fusions on transposition activity are illustrated. Panel A of FIG. 21 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo with Cas-TniQ fusion (no sgRNA), lane 2 = holo with Cas-TniQ fusion (+sgRNA), lane 3 = apo with TniQ-Cas fusion (no sgRNA), lane 4 = holo with TniQ-Cas fusion (+sgRNA). Panel B of FIG. 21 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo with Cas-TniQ fusion (no sgRNA), lane 2 = holo with Cas-TniQ fusion (+sgRNA), lane 3 = apo with TniQ-Cas fusion (no sgRNA), lane 4 = holo with TniQ-Cas fusion (+sgRNA). Panel C of Figure 21 illustrates a gel image of PCR4 of transposition (detecting RE junction to donor): lane 1 = apo with TniQ-Cas fusion (no sgRNA), lane 2 = holo with TniQ-Cas fusion (+sgRNA), lane 3 = holo-Cas alone, lane 4 = apo with TniQ-48 linker-Cas fusion (no sgRNA), lane 5 = holo with TniQ-48 linker-Cas fusion (+sgRNA), lane 6 = apo with TniQ-68 linker-Cas fusion (no sgRNA), lane 7 = holo with TniQ-68 linker-Cas fusion (+sgRNA), lane 8 = holo with TniQ-72 linker-Cas fusion (+sgRNA). Figure 21, panel D illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo with TniQ-Cas fusion (no sgRNA), lane 2 = holo with TniQ-Cas fusion (+sgRNA), lane 3 = holo-Cas alone, lane 4 = apo with TniQ-48 linker-Cas fusion (no sgRNA), lane 5 = holo with TniQ-48 linker-Cas fusion (+sgRNA), lane 6 = apo with TniQ-68 linker-Cas fusion (no sgRNA), lane 7 = holo with TniQ-68 linker-Cas fusion (+sgRNA), lane 8 = holo with TniQ-72 linker-Cas fusion (+sgRNA).Panel E of Figure 21 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 4 = holo (+sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 5 = apo (no sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion, lane 6 = holo (+sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion. Panel F of Figure 21 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 4 = holo (+sgRNA) with NLS-TniQ-Cas-NLS fusion, lane 5 = apo (no sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion, lane 6 = holo (+sgRNA) with NLS-TniQ-77 linker-Cas-NLS fusion. Panel G of Figure 21 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TniQ-Cas-NLS apo (no sgRNA), lane 4 = NLS-TniQ-Cas-NLS holo (+sgRNA), lane 5 = Cas-NLS-P2A-NLS-TniQ apo (no sgRNA), lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA). Panel H of Figure 21 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = NLS-TniQ-Cas-NLS apo (no sgRNA), lane 4 = NLS-TniQ-Cas-NLS holo (+sgRNA), lane 5 = Cas-NLS-P2A-NLS-TniQ apo (no sgRNA), lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA). [Figure 22]Figure 22 illustrates the expression of TnsB and TnsC in human cells, followed by cell fractionation and the results of an in vitro transposition reaction. Panel A of Figure 22 illustrates a gel image of PCR4 (detecting RE junctions to the donor) of transposition: lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) with untreated (no TnsB) cytoplasm, lane 4 = holo (+sgRNA) with untreated nucleoplasm, lane 5 = holo (+sgRNA) with NLS-TnsB cytoplasm, lane 6 = holo (+sgRNA) with NLS-TnsB nucleoplasm, lane 7 = holo (+sgRNA) with TnsB-NLS cytoplasm, lane 8 = holo (+sgRNA) with TnsB-NLS nucleoplasm, lane 9 = holo (+sgRNA) with NLS-TniQ cytoplasm, lane 10 = holo (+sgRNA) with NLS-TniQ nucleoplasm. Panel B of Figure 22 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) with untreated (no TnsB) cytoplasm, lane 4 = holo (+sgRNA) with untreated nucleoplasm, lane 5 = holo (+sgRNA) with NLS-TnsB cytoplasm, lane 6 = holo (+sgRNA) with NLS-TnsB nucleoplasm, lane 7 = holo (+sgRNA) with TnsB-NLS cytoplasm, lane 8 = holo (+sgRNA) with TnsB-NLS nucleoplasm, lane 9 = holo (+sgRNA) with NLS-TniQ cytoplasm, lane 10 = holo (+sgRNA) with NLS-TniQ nucleoplasm. Panel C of Figure 22 illustrates a gel image of PCR 4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) without TnsC, lane 4 = holo (+sgRNA) with untreated (no TnsC) cytoplasm, lane 5 = holo (+sgRNA) with untreated nucleoplasm, lane 6 = holo (+sgRNA) with NLS-HA-TnsC cytoplasm, lane 7 = holo (+sgRNA) with NLS-HA-TnsC nucleoplasm, lane 8 = holo (+sgRNA) with TnsC-NLS cytoplasm, lane 9 = holo (+sgRNA) with TnsC-NLS nucleoplasm.Panel D of Figure 22 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = holo (+sgRNA) without TnsC, lane 4 = holo (+sgRNA) with untreated (no TnsC) cytoplasm, lane 5 = holo (+sgRNA) with untreated nucleoplasm, lane 6 = holo (+sgRNA) with NLS-HA-TnsC cytoplasm, lane 7 = holo (+sgRNA) with NLS-HA-TnsC nucleoplasm, lane 8 = holo (+sgRNA) with TnsC-NLS cytoplasm, lane 9 = holo (+sgRNA) with TnsC-NLS nucleoplasm. Panel E of Figure 22 illustrates a gel image of PCR4 of the transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, lane 4 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, lane 5 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, lane Lane 6 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, lane 7 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 8 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 9 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, lane 10 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm.Panel F of Figure 22 illustrates a gel image of PCR 5 of the transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, lane 4 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC cytoplasm, lane 5 = apo (no sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, lane Lane 6 = holo (+sgRNA) NLS-TnsB-IRES-NLS-TnsC nucleoplasm, lane 7 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 8 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC cytoplasm, lane 9 = apo (no sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm, lane 10 = holo (+sgRNA) TnsB-NLS-IRES-NLS-TnsC nucleoplasm. [Figure 23]Figure 23 illustrates the expression of Cas12k and TniQ combined constructs in human cells, followed by the results of an in vitro transposition study. Panel A of Figure 23 illustrates a gel image of PCR5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = Cas-NLS holo (+sgRNA) cytoplasm, lane 4 = Cas-NLS holo (+sgRNA) nucleoplasm, lane 5 = Cas-NLS holo (+sgRNA) nucleoplasm + additional sgRNA, lane 6 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm, lane 7 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm, lane 8 = Cas-NLS-P2A-NLS-TniQ holo (+sgRNA) cytoplasm + additional sgRNA. Panel B of Figure 23 illustrates a gel image of PCR4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 4 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 5 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 6 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 7 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + additional holo Cas-NLS, lane 8 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ.Panel C of Figure 23 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 4 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ cytoplasm, lane 5 = apo (no sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 6 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm, lane 7 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + additional holo Cas-NLS, lane 8 = holo (+sgRNA) Cas-NLS-P2A-NLS-TniQ nucleoplasm + NLS-TniQ. Panel D of Figure 23 illustrates a gel image of PCR4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 4 = holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 5 = apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 6 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 7 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional holoCas-NLS, lane 8 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ. Panel E of Figure 23 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 4 = holo (+sgRNA) NLS-TniQ-Cas-NLS cytoplasm, lane 5 = apo (no sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 6 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm, lane 7 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + additional holo-Cas-NLS, lane 8 = holo (+sgRNA) NLS-TniQ-Cas-NLS nucleoplasm + NLS-TniQ.Panel F of Figure 23 shows a gel image of PCR4 of transposition (detecting RE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, lane 4 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, lane 5 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 6 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 7 = apo (no sgRNA) Cas-NLS Lane 5 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 6 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, lane 9 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 10 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 11 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 12 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ.Panel G of Figure 23 illustrates a gel image of PCR 5 of transposition (detecting LE junctions to the donor): lane 1 = apo (no sgRNA), lane 2 = holo (+sgRNA), lane 3 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, lane 4 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ cytoplasm, lane 5 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 6 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 7 = apo (no sgRNA) Cas- Lane 8 = apo (no sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ, lane 9 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm, lane 10 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional PURExpress, lane 11 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + additional Cas-NLS, lane 12 = holo (+sgRNA) Cas-NLS-IRES-NLS-TniQ nucleoplasm + NLS-TniQ. [Figure 24] Illustrates electrophoretic mobility shift assay (EMSA) results of 64-1 TnsB and its LE DNA sequence. EMSA results confirm binding and TnsB recognition. TnsB protein was expressed in an in vitro transcription / translation system, incubated with FAM-labeled DNA containing the LE sequence, and then separated on a native 5% TBE gel. Binding is observed as an upward shift in the labeled band. Multiple TnsB binding sites lead to multiple shifts in the EMSA. Lane 1: FAM-labeled DNA only. Lane 2: FAM DNA + in vitro transcription / translation system (no TnsB protein). Lane 3: FAM DNA + TnsB. [Diagram 25]The activity of Cas12k and TniQ fusions tested in vitro is illustrated. Panel A of Figure 25 shows a gel image showing the transposition activity of the left end to the donor. Lane 1 = apo (no sgRNA), lane 2 = holo (with sgRNA), lane 3 = (Skip) Cas12k-2xNLS apo (no sgRNA), lane 4 = (Skip) Cas12k-2xNLS holo (with sgRNA), lane 5 = Cas12k-Cbx5 fusion protein apo (no sgRNA), lane 6 = Cas12k-Cbx5 fusion protein holo (with sgRNA). Panel B of Figure 25 shows a gel image showing the transposition activity of the right end to the donor. Lane 1 = apo (no sgRNA), lane 2 = holo (with sgRNA), lane 3 = Cas12k-2xNLS apo (no sgRNA), lane 4 = Cas12k-2xNLS holo (with sgRNA), lane 5 = Cas12k-Cbx5 fusion protein apo (no sgRNA), lane 6 = Cas12k-Cbx5 fusion protein holo (with sgRNA). Panel C of Figure 25 illustrates a gel image showing the transposition activity of the left end into the donor. Lane 1 = apo (no sgRNA), lane 2 = holo (with sgRNA), lane 3 = Cbx5-TniQ fusion protein holo (with sgRNA), lane 4 = H1core-TniQ fusion protein holo (with sgRNA), lane 5 = HMGN1-TniQ fusion protein holo (with sgRNA), lane 6 = Cas12k-H1core fusion protein holo (with sgRNA), lane 7 = Cas12k-HMGN1 fusion protein holo (with sgRNA). Panel D of Figure 25 illustrates a gel image showing the transposition activity of the left end to the donor under the combinatorial sso7d fusion protein described in the lanes.Lane 1 = WT Cas12k with WT TniQ apo (without sgRNA), lane 2 = WT Cas12k with WT TniQ holo (with sgRNA), lane 3 = WT Cas12k with sso7d-TniQ fusion protein apo (with sgRNA), lane 4 = WT Cas12k with sso7d-TniQ fusion protein holo (with sgRNA), lane 5 = Cas12k-sso7d fusion protein with WT TniQ apo (without sgRNA), lane 6 = Cas12k-sso7d fusion protein with WT TniQ holo (with sgRNA), lane 7 = Cas12k-sso7d fusion protein with sso7d-TniQ fusion protein apo (without sgRNA), lane 8 = Cas12k-sso7d fusion protein with sso7d-TniQ holo (with sgRNA). [Figure 26A]Illustrating the function of sso7d with Cas12k. Figure 26A depicts a gel image of LE targeting transposition in nuclear extracts using Cas12k-sso7d-NLS. Lane 1 = apo (without sgRNA) in vitro transposition reaction with Cas12k-sso7d-NLS with other WT components; lane 2 = holo (with sgRNA) in vitro transposition reaction with Cas12k-sso7d-NLS with other WT components; lane 3 = cytoplasmic extract of Cas12k-sso7d-NLS cell line with apo condition (without sgRNA) of other in vitro expressed CAST components; lane 4 = cytoplasmic extract of Cas12k-sso7d-NLS cell line with holo condition (with sgRNA) of other in vitro expressed CAST components; lane 5 = nuclear extract of Cas12k-sso7d-NLS cell line with apo condition (without sgRNA) of other in vitro expressed CAST components; lane 6 = nuclear extract of Cas12k-sso7d-NLS cell line with holo condition (with sgRNA) of other in vitro expressed CAST components; lane 7 = additional WT Lane 6 = Nuclear extract of Cas12k-sso7d-NLS cell line in holo condition (with sgRNA) of other in vitro expressed CAST components with Cas12k, lane 7 = Nuclear extract of Cas12k-sso7d-NLS cell line in holo condition (with sgRNA) of other in vitro expressed CAST components with additional WT TniQ. Arrows indicate active transposition events. [Figure 26B] Figure 26B illustrates the function of sso7d with Cas12k. Figure 26B illustrates the sequence of the gel extracted band from lane 6 of 2A aligned to the predicted transposition product. [Figure 26C]Illustrating the function of sso7d with Cas12k. Figure 26C depicts a gel image of LE targeted transposition in nuclear extracts with WT or DBD TniQ fusions or the addition of WT Cas12k. Nuclear extracts contain WT NLS-tagged Cas12k, TnsB, TnsC, and TniQ. Lane 1 = nuclear extract with apoWT Cas12k and WT TniQ (no sgRNA), lane 2 = nuclear extract with holoCas12k (with sgRNA) and WT TniQ, lane 3 = nuclear extract with additional WT TniQ apo condition (no sgRNA), lane 4 = additional WT Lane 5 = nuclear extract from additional Cbx5-TniQ fusion protein apo condition (without sgRNA), lane 6 = nuclear extract from additional Cbx5-TniQ fusion protein holo condition (with sgRNA), lane 7 = nuclear extract from additional H1core-TniQ fusion protein apo condition (without sgRNA), lane 8 = nuclear extract from additional H1core-TniQ fusion protein holo condition (with sgRNA), lane 9 = nuclear extract from additional HMGN1-TniQ fusion protein apo condition (without sgRNA), lane 10 = nuclear extract from additional HMGN1-TniQ fusion protein holo condition (with sgRNA), lane 11 = nuclear extract from additional sso7d-TniQ fusion protein apo condition (without sgRNA), lane 12 = nuclear extract from additional sso7d-TniQ fusion protein holo condition (with sgRNA). [Figure 27]Gel images of transposition reactions are illustrated. Panel A of Figure 27 illustrates gel images of transposition activity of Cas12k-sso7d and HMGN1-TniQ from cell extracts. All lanes contain added in vitro expressed TnsB and TnsC. Lane 1 = in vitro CAST holo (with sgRNA), lane 2 = in vitro CAST apo (without sgRNA), lane 3 = cytoplasmic extract without additional Cas12k, TniQ protein, and sgRNA, lane 4 = cytoplasmic extract without Cas12k and TniQ protein and with additional sgRNA, lane 5 = nuclear extract without additional Cas12k and TniQ protein and sgRNA, lane 6 = nuclear extract without additional Cas12k and TniQ and with additional sgRNA, lane 7 = nuclear extract with additional Cas12k and sgRNA only, lane 8 = nuclear extract with additional TniQ and sgRNA only. Panel B of Figure 27 illustrates gel images of transposition activity of Cas12k-sso7d and H1core-TniQ from cell extracts. All lanes contain added in vitro expressed TnsB and TnsC. Lane 1 = in vitro CAST holo (with sgRNA), lane 2 = in vitro CAST apo (without sgRNA), lane 3 = cytoplasmic extract without added Cas12k, TniQ protein, and sgRNA, lane 4 = cytoplasmic extract without Cas12k and TniQ protein and with added sgRNA, lane 5 = nuclear extract without added Cas12k and TniQ protein and sgRNA, lane 6 = nuclear extract without added Cas12k and TniQ and with added sgRNA, lane 7 = nuclear extract with added Cas12k and sgRNA only, lane 8 = nuclear extract with added TniQ and sgRNA only. Figure 27, panel C, depicts a gel image of the transposition activity of Cas12k-sso7d, HMGN1-TniQ, TnsB, and TnsC cell extracts.Lane 1 = in vitro CAST apo (no sgRNA), lane 2 = in vitro CAST holo (with sgRNA), lane 3 = cytoplasmic extract without additional Cas12k, TniQ protein, and sgRNA, lane 4 = cytoplasmic extract without Cas12k and TniQ protein and with additional sgRNA, lane 5 = nuclear extract without additional Cas12k and TniQ protein and sgRNA, lane 6 = nuclear extract without additional Cas12k and TniQ and with additional sgRNA, lane 7 = nuclear extract with additional Cas12k and sgRNA only, lane 8 = nuclear extract with additional TniQ and sgRNA only. Panel D of Figure 27 depicts gel images of transposition activity of Cas12k-sso7d, H1core-TniQ, TnsB, and TnsC cell extracts. Lane 1 = in vitro CAST apo (no sgRNA), lane 2 = in vitro CAST holo (with sgRNA), lane 3 = cytoplasmic extract without added Cas12k, TniQ protein, and sgRNA, lane 4 = cytoplasmic extract without Cas12k and TniQ protein and with added sgRNA, lane 5 = nuclear extract without added Cas12k and TniQ protein and sgRNA, lane 6 = nuclear extract without added Cas12k and TniQ and with added sgRNA, lane 7 = nuclear extract with added Cas12k and sgRNA only, lane 8 = nuclear extract with added TniQ and sgRNA only. [Figure 28A] Schematic diagram of serial dilution of target DNA for in vitro transposition experiments. CAST components are expressed in PureExpress and added to reactions with in vitro transcribed sgRNA and donor plasmid. Target plasmid DNA is added at decreasing concentrations and tested for transposition experiments. When the minimum amount of target DNA is determined, the transposition reaction is assayed by adding increasing amounts of human genomic DNA. [Figure 28B]A diagram of PCR amplification of transposition reactions is shown. An 8N PAM plasmid library (8N-target, Rxn#1) is targeted with the CAST system to integrate donor DNA (Rxn#2). Upon successful integration, junction PCR reactions are performed with primers to amplify four putative integration reactions based on the orientation of cargo integration (Rxn#3, #4, #5, and #6). [Figure 28C] 28B shows PCR reaction products from an in vitro transposition assay using serial dilutions of target plasmid DNA. Target, donor, and reactions #3, #4, #5, and #6 correspond to the PCR integration products shown in FIG. 28B. [Figure 28D] PCR reaction products from an in vitro transposition assay using a fixed amount of target plasmid DNA (0.5 ng) with increasing amounts of human genomic DNA added to increase the search space. Target, donor, and reactions #3, #4, #5, and #6 correspond to the PCR integration products shown in Figure 28B. [Figure 29A] A schematic of a transposition reaction across a high copy element is shown. The target PCR product spans the wild type target element when assayed with CAST protein and an sgRNA targeting one of multiple aligned targets. Integration can occur in either the forward orientation, reverse orientation, or both. The forward transposition product is assayed by junction PCR, which amplifies the region encompassing the LE of the donor DNA to the 5' end of the target site (Fwd PCR). The reverse junction reaction assays the region encompassing the LE of the donor DNA to the 3' end of the target element (Rev PCR). [Figure 29B] 29A shows PCR reaction products from an in vitro transposition assay at 15 target sites (guides) in the line 1 3' element in human genomic DNA. The targets and Fwd and Rev PCR of the reactions correspond to the PCR integration products shown in FIG. 29A. [Figure 29C]29A-29C show PCR reaction products from an in vitro transposition assay at 15 target sites (guides) in the SVA element in human genomic DNA. The targets and Fwd and Rev PCR of the reactions correspond to the PCR integration products shown in FIG. 29A. Bands highlighted with arrows indicate successful targeted integration. [Figure 29D] Figure 29 shows PCR reaction products from an in vitro transposition assay at 15 target sites (guides) in HERV elements in human genomic DNA. The targets and Fwd and Rev PCR of the reactions correspond to the PCR integration products shown in Figure 29A. Bands highlighted with arrows indicate successful targeted integration. [Figure 29E] Sanger sequencing of Fwd PCR integration products at multiple target sites in the line 1 3' element. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 29F] Sanger sequencing of Rev PCR integration products at multiple target sites in the line 1 3' element is shown. The point where the sequencing trace stops matching the target DNA (gray vertical bar) is the point where integration occurs. [Figure 29G] Sanger sequencing of the Fwd PCR integration product at SVA target site 3 is shown. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 29H] Sanger sequencing of the Fwd PCR product at HERV target site 5 is shown. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Diagram 30] Shown are PCR reaction products from an in vitro transposition assay at line 1 target sites 12 and 15 in human genomic DNA with functional domains. The target, and Fwd and Rev PCR of the reaction correspond to the PCR integration products shown in Figure 28A. Bands highlighted with arrows indicate successful targeted integration. [Figure 31A]In vitro transposition experiments with CAST, S15, NLS-S15, and S15-NLS expressed from eukaryotic transcription / translation reactions are shown. Figure 31A shows an in vitro transposition reaction with MG64-1 CAST and S15. Wheat germ extract-expressed CAST components promote transposition without the addition of S15, albeit at a slower rate (faint band highlighted with an arrow). Addition of PURExpress reagent (spent PUREx) increases transposition efficiency as shown by the intensity of the band in Rxn#5 (PURExpress reagent contains S15). Independent addition of S15 and S15-NLS translated from wheat germ extract reactions increases transposition efficiency by MG64-1 in vitro compared to other conditions tested (intense band highlighted with an arrow). [Figure 31B] Figure 31B shows an in vitro transposition experiment using CAST, S15, NLS-S15, and S15-NLS expressed from a eukaryotic transcription / translation reaction. Figure 31B shows an in vitro reaction of transposition using the NLS-S15 construct. Addition of PURExpress reagent increases in vitro transposition (lane 3) compared to the CAST construct only condition (lane 2). The NLS-S15 construct did not improve transposition (lanes 4-5). The boxed Rxn#5 represents the expected band if transposition activity was detected. [Figure 32A] Schematic diagram of fusion plasmids for cell transposition. Figure 32A: Two targeting complex plasmids and one donor plasmid are assembled for high copy element line 1, targets 8, 12, and 15, and SVA target 3. [Figure 32B] Schematic diagram of fusion plasmid for cell transposition. Figure 32B shows cell transposition into high copy elements using H1core-TniQ or HMGN1-TniQ in line 1 targets 8, 12, 15, and SVA target 3. Arrows indicate amplified transposition junction reactions in either forward (Fwd PCR) or reverse (Rev PCR) orientation of transposition. Mock controls represent reactions without targeting or donor plasmid. [Figure 32C]Schematic diagram of fusion plasmid for cell transposition. Figure 32C shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 carrying NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Fig. 32D] Schematic diagram of fusion plasmid for cell transposition. Figure 32D shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 8. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 32E] Schematic diagram of fusion plasmid for cell transposition. Figure 32E shows Sanger sequencing of PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 with NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Fig. 32F] Schematic diagram of fusion plasmid for cell transposition. Figure 32F shows Sanger sequencing of PCR integration product Rev PCR at line 1 3' target site 12. Integration was mediated by MG64-1 carrying NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Fig. 32G] A schematic diagram of the fusion plasmid for cell transposition is shown. Figure 32G shows Sanger sequencing of the PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-H1core-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Fig. 32H]Schematic diagram of fusion plasmid for cell transposition. Figure 32H shows Sanger sequencing of PCR integration product Fwd PCR at line 1 3' target site 15. Integration was mediated by MG64-1 carrying the NLS-HMGN1-TniQ fusion. The point where the sequencing trace stops matching the donor DNA (gray vertical bar) is the point where integration occurs. [Figure 33A] Illustrating the in vitro screening of MG64-1 Cas12k CAST transposition. Figure 33A: Schematic representation of constructs used for MG64-1 holocomplex purification. [Figure 33B] Illustrates in vitro screening of MG64-1 Cas12k CAST transposition. Figure 33B: Schematic of junction PCR for detection of transposition products. A target substrate with 5'PAM followed by a protospacer (target, Rxn#1) is targeted with the CAST system to incorporate cargo DNA (Rxn#2). If the incorporation is successful, a junction PCR reaction is performed with primers to amplify the four putative incorporation reactions based on the orientation of the cargo incorporation. [Figure 33C] MG64-1 protein purification is illustrated in Figure 33C: Fractions collected during 2 L-scale purification of MG64-1 holocomplex running on a stain-free denaturing PAGE gel. [Figure 33D] Illustrating MG64-1 protein purification. Figure 33D: Chromatogram of size exclusion chromatography (SEC) performed on MG64-1 holocomplex. The peak centered at 29.3 mL (peak 1) was used for in vitro activity assays. [Figure 34A]Illustrated is in vitro transposition using peak 1 recovered holo complex supplemented with TnT expression components. Lane L) ladder, lane 1) TnT expression CAST components apo condition (-sgRNA), lane 2) TnT expression CAST components holo condition (+sgRNA), lane 3) purified peak 1 complemented with TnT CAST components without additional supplementation of Cas12k (-TnT Cas12k), lane 4) purified peak 1 complemented with TnT CAST components without additional supplementation of TnsC (-TnT TnsC), lane 5) peak 1 complemented with TnT CAST components without additional supplementation of TniQ (-TnT TniQ), lane 6) peak 1 complemented with TnT CAST components without additional supplementation of S15 (-TnT S15). [Figure 34B] Sanger sequencing of lane 3, lane 4, lane 5, and lane 6 from both the pDonor and target orientations of the amplified LE-PAM target-donor junction are illustrated. The vertical line defines the predicted translocation junction for MG64-1 in the reference sequence. Degradation of the signal from either orientation results in multiple signals reflected in the PCR amplification. [Diagram 35] FIG. 1 illustrates the identification of ribosomal protein S15 homologues in cyanobacterial genomic fragments. Candidate sequences from the same sample in which MG64-1 was recovered are highlighted with dark circles. The reference S15 from E. coli is indicated with an arrow.
[0051] Brief Description of the Sequence Listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.
[0052] MG64 SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200 show full-length peptide sequences of MG64 Cas effectors.
[0053] SEQ ID NOs: 2-4, 13-15, 17-19, and 65-67 show peptide sequences of MG64 translocation proteins that may comprise a recombinase complex associated with the MG64 Cas effector.
[0054] SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202 show the nucleotide sequences of MG64 tracrRNA derived from the same locus as the MG64 Cas effector.
[0055] SEQ ID NOs: 7 and 34-35 show the nucleotide sequences of MG64-targeted CRISPR repeats.
[0056] SEQ ID NOs: 106-108 and 201 show the nucleotide sequences of MG64 crRNA.
[0057] SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93 show the nucleotide sequences of the right transposase recognition sequences associated with the MG64 system.
[0058] SEQ ID NOs: 9, 11, 36-38, 76, and 78 show the nucleotide sequences of the left-side transposase recognition sequences associated with the MG64 system.
[0059] SEQ ID NO:31 shows the PAM sequence associated with the MG64 Cas effector described herein.
[0060] SEQ ID NOs: 45-63, 68-75, and 96-103 show the nucleotide sequences of single guide RNAs engineered to function with the MG64 Cas effector.
[0061] SEQ ID NOs: 113-120 show the nucleotide and peptide sequences of MG64 DNA binding domain CAST fusion proteins.
[0062] SEQ ID NO: 188 shows the nucleotide sequence of the MG64 expression construct.
[0063] MG190 SEQ ID NOs: 189 to 199 show the full-length peptide sequences of the MG190 ribosomal protein S15 homologues.
[0064] Other Arrays SEQ ID NOs: 86 to 87 and 172 to 187 show peptide sequences of nuclear localization signals.
[0065] SEQ ID NOs: 88 to 89 show the peptide sequences of linkers.
[0066] SEQ ID NOs: 90 to 92 show the peptide sequences of epitope tags.
[0067] SEQ ID NOs: 109 to 112 show the peptide sequences of the DNA binding domain.
[0068] SEQ ID NOs: 121 to 123 show genomic target sequences.
[0069] SEQ ID NOs: 124 to 160 show target guide sequences.
[0070] SEQ ID NOs: 161 to 163 show the nucleic acid sequences of S15 fusion proteins.
[0071] SEQ ID NO: 164 shows the donor construct.
[0072] SEQ ID NO: 165 shows the MG64-1 sgRNA sequence.
[0073] SEQ ID NO: 166 shows the linker sequence.
[0074] SEQ ID NOs: 167 to 169 show the amino acid sequences of S15 fusion proteins.
[0075] SEQ ID NOs: 170 to 171 show promoter sequences. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0076] While various embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be used.
[0077] The practice of some of the methods disclosed herein, unless otherwise indicated, employs techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012), the series Current Protocols in Molecular Biology (FMA Usubel, et al. eds.), the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)).
[0078] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variations thereof are used in any of the detailed description and / or claims, such terms are intended to be inclusive in the same manner as the term "comprising."
[0079] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0080] As used herein, "cell" refers to a biological cell. A cell may be the basic structural, functional, and / or biological unit of a living organism. A cell may originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-cell eukaryotic cells, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, cane, corn, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, mosses, cells from mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.), and cells from plants (e.g., cereals, fruits, vegetables, grains, soybeans, cane, corn, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, mosses, cells from mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.). C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Sometimes the cells are not derived from a naturally occurring organism (e.g., cells can be synthetically produced and sometimes referred to as artificial cells).
[0081] The term "nucleotide" as used herein refers to a base-sugar-phosphate combination. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides may be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules that contain them. The term nucleotide as used herein may refer to dideoxyribonucleoside triphosphates (ddNTPs) and derivatives thereof. Illustrative examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as by using a moiety that includes an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides may include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer (Foster City, Calif.); fluoro-conjugated deoxynucleotides, fluoro-conjugated Cy3-dCTP, fluoro-conjugated Cy5-dCTP, fluoro-conjugated fluoroX-dCTP, fluoro-conjugated Cy3-dUTP, and fluoro-conjugated Cy5-dUTP available from Amersham (Arlington Heights, Ill.); Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Mannheim (Indianapolis, Ind.); and Molecular Examples of chromosomal labeling nucleotides available from Probes (Eugene, Oreg.) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs can include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0082] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-stranded, double-stranded, or multiple-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and perform any function. In polynucleotides when referring to T, T means U (uracil) in RNA and T (thymine) in DNA. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, heterologous nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci defined from binding analyses, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers.The sequence of nucleotides may be interrupted by non-nucleotide components.
[0083] The term "transfection" or "transfected" refers to the introduction of a nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule can be a genetic sequence encoding a complete protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88.
[0084] The terms "peptide", "polypeptide", and "protein" are used interchangeably herein to refer to a polymer of at least two amino acid residues linked by peptide bonds. The term does not refer to a specific length of the polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers that contain at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins, and proteins with or without secondary and / or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified by any other manipulation, such as, for example, disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. The terms "amino acid" and "amino acids" as used herein refer to natural and unnatural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids can include natural and unnatural amino acids that have been chemically modified to include non-naturally occurring groups or chemical moieties on the amino acid. Amino acid analogs can refer to amino acid derivatives. The term "amino acid" includes both D- and L-amino acids.
[0085] As used herein, "non-natural" may refer to a nucleic acid or polypeptide sequence that is not found in a natural nucleic acid or protein. Non-natural may refer to an affinity tag. Non-natural may refer to a fusion. Non-natural may refer to a naturally occurring nucleic acid or polypeptide sequence that includes mutations, insertions, and / or deletions. A non-natural sequence may exhibit and / or encode an activity (e.g., an enzyme activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, an ubiquitination activity, etc.) that may also be exhibited by the nucleic acid and / or polypeptide sequence to which the non-natural sequence is fused. A non-natural nucleic acid or polypeptide sequence may be joined by genetic engineering to a naturally occurring nucleic acid and / or polypeptide sequence (or a variant thereof) to generate a chimeric nucleic acid and / or polypeptide sequence that encodes a chimeric nucleic acid or polypeptide.
[0086] The term "promoter" as used herein refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and may be located adjacent to or overlapping the nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often referred to as transcription factors, that facilitate the binding of RNA polymerase to DNA resulting in gene transcription. A "basal promoter", also referred to as a "core promoter", may refer to a promoter that contains all the basic and necessary elements to facilitate the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters typically, but not necessarily, contain a TATA-box and / or a CAAT box. In some embodiments, different promoters induce the expression of a gene in different tissues or cell types, or at different developmental stages, or in response to different environmental or physiological conditions or inducer molecules. A promoter that causes a gene to be expressed in most cell types at most times is commonly referred to as a "constitutive promoter". A promoter that causes a gene to be expressed in specific cell and tissue types is commonly referred to as a "cell-specific promoter" or a "tissue-specific promoter", respectively. Promoters that cause expression of a gene at a specific stage of development or cell differentiation are generally referred to as "development-specific promoters" or "cell differentiation-specific promoters." Promoters that induce and result in expression of a gene after exposure or treatment of cells with a promoter-inducing drug, biomolecule, chemical, ligand, light, etc. are generally referred to as "inducible promoters" or "regulatable promoters." It is further recognized that in some embodiments, DNA fragments of different lengths have the same promoter activity, since the exact boundaries of regulatory sequences are in most cases not completely defined.
[0087] The term "expression" as used herein refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which the transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as the "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0088] As used herein, "operably linked", "operable linkage", "operatively linked", or their grammatical equivalents refer to an arrangement of genetic elements, such as promoters, enhancers, polyadenylation sequences, etc., such that operation (e.g., movement or activation) of a first genetic element has some effect on a second genetic element. The effect on the second genetic element may be, but need not be, of the same type as the operation of the first genetic element. For example, two genetic elements are operably linked if movement of the first element causes activation of the second element. A regulatory element is operably linked to a coding region if the regulatory element helps to initiate transcription of the coding sequence, which may include, for example, promoter and / or enhancer sequences. There may be intervening residues between the regulatory element and the coding region, so long as this functional relationship is maintained.
[0089] As used herein, a "vector" refers to a polymer or association of polymers that contains or is associated with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally includes a genetic element, such as a regulatory element, operably linked to a gene to facilitate expression of the gene in a target.
[0090] As used herein, "expression cassette" and "nucleic acid cassette" are used interchangeably to refer to a combination of nucleic acid sequences or elements that are expressed together or operably linked for expression. In some cases, an expression cassette refers to a combination of regulatory elements and one or more genes that are operably linked for expression.
[0091] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence can be the ability to affect expression in a manner attributable to the full-length sequence.
[0092] The terms "engineered," "synthetic," and "artificial" are used interchangeably herein to refer to an entity that has been modified by human intervention. For example, the terms may refer to a polynucleotide or polypeptide that does not occur in nature. An engineered peptide may, but need not, have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) with a naturally occurring human protein. For example, the VPR domain and the VP64 domain are synthetic transactivation domains. By way of non-limiting examples, a nucleic acid may be modified by changing its sequence to one that does not occur in nature, a nucleic acid may be modified by ligating to a nucleic acid with which it is not naturally associated such that the ligated product possesses a function not present in the original nucleic acid, an engineered nucleic acid may be synthesized in vitro with a sequence that does not occur in nature, a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature, an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.
[0093] The term "tracrRNA" or "tracr sequence" refers to transactivating CRISPR RNA. tracrRNA interacts with CRISPR(cr)RNA to form a guide nucleic acid (e.g., guide RNA or gRNA) that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid. When tracrRNA is engineered, it can have about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% sequence identity and / or sequence similarity with a wild-type exemplary tracrRNA sequence (e.g., tracrRNA from S. pyogenes, S. aureus, etc.). tracrRNA can refer to modified forms of tracrRNA that can include nucleotide changes such as deletions, insertions, or substitutions, variants, mutations, or chimeras. tracrRNA may refer to a nucleic acid that may be at least about 60% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from S. pyogenes, S. aureus, etc.) sequence over a stretch of at least six consecutive nucleotides. For example, a tracrRNA sequence may be at least about 60% identical, at least about 65% identical, at least about 70% identical, at least about 75% identical, at least about 80% identical, at least about 85% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, or 100% identical to a wild-type exemplary tracrRNA (e.g., tracrRNA from S. pyogenes, S. aureus, etc.) sequence over a stretch of at least six consecutive nucleotides. Type II tracrRNA sequences can be predicted on genomic sequences by identifying regions that have complementarity to portions of the repeat sequences in adjacent CRISPR arrays.
[0094] As used herein, a "guide nucleic acid" or "guide polynucleotide" refers to a nucleic acid that can hybridize to a target nucleic acid and thereby direct an associated nuclease to the target nucleic acid. A guide nucleic acid can be an RNA (guide RNA or gRNA). A guide nucleic acid can be a DNA. A guide nucleic acid can be a mixture of RNA and DNA. A guide nucleic acid can include crRNA or tracrRNA, or a combination of both. A guide nucleic acid can be engineered. A guide nucleic acid can be programmed to specifically bind to a target nucleic acid. A portion of a target nucleic acid can be complementary to a portion of a guide nucleic acid. A strand of a double-stranded target polynucleotide that is complementary to a guide nucleic acid and hybridizes with the guide nucleic acid can be referred to as a complementary strand. A strand of a double-stranded target polynucleotide that is complementary to a complementary strand and therefore may not be complementary to the guide nucleic acid can be referred to as a non-complementary strand. A guide nucleic acid can include a polynucleotide strand and can be referred to as a "single guide nucleic acid." A guide nucleic acid can include two polynucleotide strands and can be referred to as a "double guide nucleic acid." Unless otherwise specified, the term "guide nucleic acid" is inclusive and can refer to both single and double guide nucleic acids. A guide nucleic acid can include a segment that can be referred to as a "nucleic acid targeting segment" or a "nucleic acid targeting sequence" or a "spacer sequence." A nucleic acid targeting segment can include a sub-segment that can be referred to as a "protein binding segment" or a "protein binding sequence" or a "Cas protein binding segment."
[0095] As used herein, the terms "gene editing" and "genome editing" can be used interchangeably. Gene editing or genome editing refers to changing the nucleic acid sequence of a gene or genome. Genome editing can include, for example, insertions, deletions, and mutations.
[0096] The term "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical, or have a certain percentage of identical amino acid residues or nucleotides, when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and presence of 11, gap cost at an extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of word length (W) of 2, expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (default parameters for BLASTP are available in BLAST at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using the Smith-Waterman homology search algorithm with parameters of match of 2, mismatch of -1, and gap of -1; MUSCLE with default parameters; MAFFT with parameters retree of 2 and maximum iterations of 1000; Novafold with default parameters; HMMER hmmalign with default parameters.
[0097] The present disclosure includes variants of any of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without destroying the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length for each other. Additionally or alternatively, by comparing aligned sequences of homologous proteins from different species, conservative substitutions can be identified by identifying amino acid residues that are not mutated between species (e.g., residues that are not conserved without altering the basic function of the encoded protein). Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity with any one of the systems described herein (e.g., the MG64 system described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants may include sequences with substitutions such that the activity of key active site residues of the endonuclease is not destroyed. In some embodiments, a functional variant of any of the systems described herein lacks at least one substitution of a conserved or functional residue called out in Figures 4 and 5. In some embodiments, a functional variant of any of the systems described herein lacks all substitutions of a conserved or functional residue called out in Figures 4 and 5.
[0098] Conservative substitution tables providing functionally similar amino acids are available in a variety of references (e.g., Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2 nd The following eight groups each contain amino acids that are conservative substitutions for one another: 1. Alanine (A), Glycine (G); 2. Aspartic acid (D), glutamic acid (E); 3. Asparagine (N), Glutamine (Q); 4. Arginine (R), Lysine (K); 5. Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6. Phenylalanine (F), Tyrosine (Y), Tryptophan (W); 7. Serine (S), Threonine (T); and 8. Cysteine (C), methionine (M).
[0099] As used herein, the term "RuvC_III domain" refers to the third discontinuous segment of the RuvC endonuclease domain (the RuvC nuclease domain is composed of three discontinuous segments, RuvC_I, RuvC_II, and RuvC_III). RuvC domains or segments thereof can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) built based on documented domain sequences (e.g., Pfam HMM PF18541 for RuvC_III).
[0100] As used herein, the term "HNH domain" refers to an endonuclease domain having characteristic histidine and asparagine residues. HNH domains can generally be identified by alignment to documented domain sequences, structural alignment to proteins with annotated domains, or comparison to hidden Markov models (HMMs) constructed based on documented domain sequences (e.g., Pfam HMM PF01844 for domain HNH).
[0101] As used herein, the term "recombinase" refers to an enzyme that mediates the recombination of DNA fragments located between recombinase recognition sequences, resulting in excision, insertion, inversion, exchange, or transposition of the DNA fragments located between the recombinase recognition sequences.
[0102] As used herein, the term "recombining" or "recombination" in the context of nucleic acid modification (e.g., genomic modification) refers to a process in which two or more nucleic acid molecules, or two or more regions of a single nucleic acid molecule, are modified by the action of a recombinase protein. Recombination can result in, among other things, excision, insertion, inversion, exchange, or rearrangement of a nucleic acid sequence within or between one or more nucleic acid molecules.
[0103] As used herein, the term "transposon" or "transposable element" refers to a nucleic acid sequence in a genome that is a mobile genetic element that can change its position in the genome. In some cases, the transposon transports additional "cargo DNA" that is excised from the genome. Transposons include, for example, retrotransposons, DNA transposons, autonomous and non-autonomous transposons, and class III transposons. The transposon nucleic acid sequence includes, for example, a gene encoding a cognate transposase, one or more recognition sequences for the transposase, or a combination thereof. In some cases, these transposons differ by the type of nucleic acid that they transpose, the type of repeats at the ends of the transposon, the type of cargo they carry, or the mode of transposition (i.e., self-repair or host repair). As used herein, the term "transposase" or "transposases" refers to an enzyme that binds to the recognition sequence of the transposon and catalyzes its movement to another part of the genome. In some cases, the movement is by a cut-and-paste mechanism or a copy-and-transposition mechanism.
[0104] As used herein, the term "Tn7" or "Tn7-like transposase" refers to a family of transposases that includes three major components: a heteromeric transposase (TnsA and / or TnsB) along with a regulatory protein (TnsC). In addition to the TnsABC transposition proteins, Tn7 elements can encode dedicated target site selection proteins, TnsD and TnsE. In conjunction with TnsABC, the sequence-specific DNA binding protein TnsD directs transposition into a conserved site called the "Tn7 attachment site", i.e., attTn7. TnsD is a member of a large family of proteins that also includes TniQ. TniQ has been shown to target transposition into the degradation site of a plasmid.
[0105] As used herein, the term "complex" refers to the joining of at least two components. Each of the two components may retain the properties / activity it had before forming the complex. The joining may be by covalent bonds, non-covalent bonds (i.e., hydrogen bonds, ionic interactions, van der Waals interactions, and hydrophobic bonds), the use of linkers, fusion, or any other suitable method. In some cases, the components in the complex are polynucleotides, polypeptides, or combinations thereof. For example, the complex may include a Cas protein and a guide nucleic acid.
[0106] In some cases, the CAST system described herein comprises one or more Tn7 or Tn7-like transposases. In certain exemplary embodiments, the Tn7 or Tn7-like transposase comprises a multimeric protein complex. In certain exemplary embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form a complex or fusion protein with each other.
[0107] In some cases, the CAST system described herein comprises one or more Tn5053 or Tn5053-like transposases. In certain exemplary embodiments, the Tn5053 or Tn5053-like transposase comprises a multimeric protein complex. In certain exemplary embodiments, the multimeric protein complex comprises TnsA, TnsB, TnsC, or TniQ. In these combinations, the transposases (TnsA, TnsB, TnsC, TniQ) may form a complex or fusion protein with each other.
[0108] As used herein, the term "Cas12k" (alternatively "Class 2 VK type") refers to a subtype of V type CRISPR system that has been found to be defective in nuclease activity (e.g., they may contain at least one defective RuvC domain that lacks at least one catalytic residue important for DNA cleavage). Such effector subtypes are generally associated with the CAST system.
[0109] As used herein, the term "functional domain (FD)" refers to a small protein that can facilitate protein interaction with DNA. Types of functional domains include, but are not limited to, DNA binding domains (DBD) and chromatin modulating domains (CMD). Non-limiting examples of functional domains include human histone 1 central globular domain (H1core), high mobility group nucleosome binding domain 1 (HMGN1), chromobox 5 (Cbx5), and Saccharolobus solfataricus sso7d. In some embodiments, the functional domains described herein can be included in a fusion protein with a system or component thereof described herein. In some embodiments, the fusion protein exhibits increased activity in a cell compared to the non-fused protein.
[0110] In accordance with IUPAC convention, the following abbreviations are used throughout the examples: A=Adenine C=Cytosine G=guanine T=Thymine R = adenine or guanine Y = cytosine or thymine S = guanine or cytosine W = adenine or thymine K = guanine or thymine M = adenine or cytosine B=C, G, or T D=A, G, or T H=A, C, or T V=A, C, or G
[0111] overview Discovery of new Cas enzymes with unique functionality and structure may confer the potential to further disrupt deoxyribonucleic acid (DNA) editing technologies, improving speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of clustered regularly interspaced short palindromic repeats (CRISPR) systems in microbes and the net diversity of microbial species, a relatively small number of functionally characterized CRISPR / Cas enzymes exist in the literature. This is in part because the vast number of microbial species are not easily cultured under laboratory conditions. Metagenomic sequencing from natural environmental niches representing a large number of microbial species could dramatically increase the number of documented new CRISPR / Cas systems and confer the potential to expedite the discovery of new oligonucleotide editing functions. A fruitful recent example of such an approach is demonstrated by the 2016 discovery of the CasX / CasY CRISPR system from metagenomic analysis of natural microbial communities.
[0112] CRISPR / Cas systems are RNA-directed nuclease complexes that have been described to function as adaptive immune systems in microorganisms. In their natural context, CRISPR / Cas systems occur in CRISPR (clustered regularly interspaced short palindromic repeats) operons or loci, which generally contain two parts: (i) an array of short repeat sequences (30-40 bp) separated by equally short spacer sequences that encode RNA-based targeting elements, and (ii) an ORF encoding a Cas that encodes a nuclease polypeptide that is directed by the RNA-based targeting element flanked by accessory proteins / enzymes. Efficient nuclease targeting of a specific target nucleic acid sequence generally requires both (i) complementary hybridization between the first 6-8 nucleic acids of the target (target seed) and the crRNA guide; and (ii) the presence of a protospacer adjacent motif (PAM) sequence within a defined vicinity of the target seed (PAM is usually a sequence that is not commonly represented in the host genome). Depending on the exact function and composition of the system, CRISPR-Cas systems are commonly organized into two classes, five types, and 16 subtypes based on shared functional characteristics and evolutionary similarities (see Figure 1).
[0113] Class 1 CRISPR-Cas systems have large multi-subunit effector complexes and include types I, III, and IV.
[0114] Type I CRISPR-Cas systems are considered to be of medium complexity in terms of components. In type I CRISPR-Cas systems, an array of RNA targeting elements is transcribed as a long precursor crRNA (pre-crRNA) that is processed at the repetitive elements, and a short mature crRNA that directs the nuclease complex to the nucleic acid target is then released following a suitable short consensus sequence called the protospacer adjacent motif (PAM). This processing occurs via the endoribonuclease subunit (Cas6) of a large endonuclease complex called Cascade, which also contains the nuclease (Cas3) protein component of the crRNA-directed nuclease complex. Cas I nuclease functions primarily as a DNA nuclease.
[0115] Type III CRISPR systems can be characterized by the presence of a central nuclease known as Cas10, along with repeat-associated mysterious proteins (RAMPs) that contain Csm or Cmr protein subunits. Similar to type I systems, mature crRNA is processed from pre-crRNA using a Cas6-like enzyme. Unlike type I and type II systems, type III systems appear to target and cleave DNA-RNA duplexes (such as the DNA strand used as a template for RNA polymerase).
[0116] Type IV CRISPR-Cas systems possess an effector complex that contains a highly reduced large subunit nuclease (csf1), two genes for RAMP proteins of the Cas5 (csf3) and Cas7 (csf2) family, and in some cases a predicted small subunit gene; such systems are commonly found on endogenous plasmids.
[0117] Class 2 CRISPR-Cas systems generally have a single polypeptide multi-domain nuclease effector and include Types II, V, and VI.
[0118] Type II CRISPR-Cas systems are considered the simplest in terms of components. In type II CRISPR-Cas systems, the processing of the CRISPR array into mature crRNA does not require the presence of a special endonuclease subunit, but rather a small transcoding crRNA (tracrRNA) with a region complementary to the array repeat sequence, which interacts with both its corresponding effector nuclease (e.g., Cas9) and the repeat sequence to form a precursor dsRNA structure that is cleaved by endogenous RNAse III to generate the mature effector enzyme loaded with both tracrRNA and crRNA. Type II nucleases are known as DNA nucleases. Type II effectors generally exhibit a structure that includes a RuvC-like endonuclease domain that fits into an RNase H fold with an unrelated HNH nuclease domain inserted into the fold of the RuvC-like nuclease domain. The RuvC-like domain is involved in cleavage of the target (e.g., crRNA-complementary) DNA strand, while the HNH domain is involved in cleavage of the variant DNA strand.
[0119] Type V CRISPR-Cas systems are characterized by a nuclease effector (e.g., Cas12) structure similar to that of type II effectors, including a RuvC-like domain. Like type II, most (but not all) type V CRISPR systems use tracrRNA to process pre-crRNA into mature crRNA, but unlike type II systems that require RNAse III to cleave pre-crRNA into multiple crRNAs, type V systems can cleave pre-crRNA using the effector nuclease itself. Like type II CRISPR-Cas systems, type V CRISPR-Cas systems are again known as DNA nucleases. Unlike type II CRISPR-Cas systems, some type V enzymes (e.g., Cas12a) appear to have robust single-stranded non-specific deoxyribonuclease activity that is activated by the first crRNA-directed cleavage of the double-stranded target sequence.
[0120] Type VI CRISPR-Cas systems have an RNA-guided RNA endonuclease. Instead of a RuvC-like domain, the single polypeptide effector of type VI systems (e.g., Cas13) contains two HEPN ribonuclease domains. Unlike both type II and type V systems, type VI systems also do not appear to require tracrRNA to process pre-crRNA into crRNA. However, similar to type V systems, some type VI systems (e.g., C2C2) appear to have robust single-stranded non-specific nuclease (ribonuclease) activity that is activated by the first crRNA-directed cleavage of the target RNA.
[0121] Due to their simpler structure, class 2 CRISPR-Cas have been most widely adopted for engineering and development as engineered nuclease / genome editing applications.
[0122] One of the initial adaptations of such a system for in vivo use involves the use of (i) purified recombinantly expressed full-length Cas9 (e.g., a class 2 type II Cas enzyme) isolated from S. pyogenes SF370, (ii) purified mature, approximately 42 nt crRNA (total crRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence) carrying an approximately 20 nt 5' sequence complementary to the target DNA sequence desired to be cleaved, followed by a 3' tracr binding sequence, (iii) purified tracrRNA transcribed in vitro from a synthetic DNA template carrying a T7 promoter sequence, and (iv) Mg 2+ Subsequent improved and engineered systems involved a crRNA (ii) joined to the 5' end of (iii) by a linker (e.g., GAAA) to form a single fusion synthetic guide RNA (sgRNA) that can itself guide Cas9 to the target (compare the top and bottom panels of FIG. 2).
[0123] Such engineered systems can be adapted for use in mammalian cells by providing a DNA vector encoding (i) an ORF encoding a codon-optimized Cas9 (e.g., a class 2 type II Cas enzyme) under a suitable mammalian promoter with a C-terminal nuclear localization sequence (e.g., SV40 NLS) and a suitable polyadenylation signal (e.g., TK pA signal), and (ii) an ORF encoding an sgRNA (having a 5' sequence starting with G, followed by 20 nt of complementary targeting nucleic acid sequence joined to the 3' tracr binding sequence, a linker, and the tracrRNA sequence) under a suitable polymerase III promoter (e.g., U6 promoter).
[0124] Transposons are mobile elements that can move between locations in a genome. Such transposons have evolved to limit the negative effects they have on the host. Various control mechanisms are used to maintain translocation at low frequency and sometimes to coordinate translocation with various cellular processes. Some prokaryotic transposons can also marshal functions that benefit the host or otherwise help maintain the element. Certain transposons may also have evolved mechanisms of strict control over target site selection, the most prominent example being the Tn7 family.
[0125] Transposon Tn7 and similar elements are reservoirs for antibiotic resistance and pathogenic functions in clinical settings, and may encode other adaptive functions in the natural environment. The Tn7 system, for example, has evolved mechanisms to almost completely avoid integration into critical host genes, but to maximize dispersal of the element by recognizing mobile plasmids and bacteriophages capable of transferring Tn7 between host bacteria.
[0126] Tn7 and Tn7-like elements control where and when they insert, and may possess one pathway that directs insertion into a single conserved location in the bacterial genome, and a second pathway that appears to be adapted to maximize targeting into mobile plasmids capable of transporting elements between bacteria (Figure 3). The link between Tn7-like transposons and CRISPR-Cas systems suggests that transposons may have hijacked CRISPR effectors that generate R-loops at target sites, facilitating the spread of transposons through plasmids and phages.
[0127] MG64 series In some embodiments, provided herein is an MG64 system for transposing a cargo nucleotide sequence into a target nucleic acid site. See Figures 4A-4B. In some embodiments, the system comprises a double-stranded nucleic acid comprising a cargo nucleotide sequence. In some embodiments, the cargo nucleotide sequence is configured to interact with a Tn7-type or Tn5053-type transposase complex. In some embodiments, the system comprises a Cas effector complex. In some embodiments, the Cas effector complex comprises a class 2 V-type Cas effector and an engineered guide polynucleotide configured to hybridize to a target nucleotide sequence. In some embodiments, the system comprises a Tn7-type or Tn5053-type transposase complex configured to bind to the Cas effector complex, the Tn7-type or Tn5053-type transposase complex comprising a TnsB subunit.
[0128] In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a right transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence.
[0129] In some cases, the target nucleic acid comprises a target nucleic acid site. In some cases, the target nucleic acid comprises a PAM sequence that is compatible with a Cas effector complex adjacent to the target nucleic acid site. In some cases, the PAM sequence is located 3' of the target nucleic acid site. In some cases, the PAM sequence is located 5' of the target nucleic acid site.
[0130] In some cases, the engineered guide polynucleotide is configured to bind to a class 2 V-type Cas effector. In some cases, the class 2 V-type Cas effector is a class 2 VK-type effector. In some cases, the Class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs:1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having 100% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.
[0131] In some cases, the TnsB subunit comprises a polypeptide having a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 2, 13, 17, and 65.In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having 100% identity to any one of SEQ ID NOs: 2, 13, 17, and 65.
[0132] In some cases, the Tn7-type transposase complex comprises at least one polypeptide (e.g., at least one, two, three, four, five, six, or more than six polypeptides) comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs:3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having 100% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.
[0133] In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having 100% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.
[0134] In some embodiments, the systems disclosed herein comprise at least one engineered guide polynucleotide, e.g., a gRNA.
[0135] In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs:5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 70% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 75% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 85% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 90% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 91% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 92% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 93% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 94% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 95% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 96% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 97% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 98% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having 100% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.
[0136] In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that are identical to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 70% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165.In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 75% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 80% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 85% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 90% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 91% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 92% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 93% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165.In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 94% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 95% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 96% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 97% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 98% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 99% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having 100% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165.
[0137] In some embodiments, the guide RNA comprises various structural elements, including but not limited to a spacer sequence that binds to a protospacer sequence (target sequence), a crRNA, and an optional tracrRNA. In some embodiments, the guide RNA comprises a crRNA that comprises a spacer sequence. In some embodiments, the guide RNA additionally comprises a tracrRNA or a modified tracrRNA.
[0138] In some embodiments, the systems provided herein include one or more guide RNAs. In some embodiments, the guide RNA includes a sense sequence. In some embodiments, the guide RNA includes an antisense sequence. In some embodiments, the guide RNA includes a nucleotide sequence other than a region that is complementary or substantially complementary to a region of a target sequence. For example, the crRNA is or is considered part of the guide RNA, or is included in the guide RNA, e.g., a crRNA:tracrRNA chimera.
[0139] In some embodiments, the guide RNA comprises synthetic or modified nucleotides. In some embodiments, the guide RNA comprises one or more internucleoside linkers modified from natural phosphodiester. In some embodiments, the internucleoside linker of the guide RNA, or all of its contiguous nucleotide sequence, is modified. For example, in some embodiments, the internucleoside linkage comprises sulfur (S), such as a phosphorothioate internucleoside linkage.
[0140] In some embodiments, the guide RNA comprises a modification to the ribose sugar or nucleobase. In some embodiments, the guide RNA comprises one or more nucleosides comprising a modified sugar moiety, which is a modification of the sugar moiety as compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is in the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acid (LNA)), or a non-linked ribose ring that typically lacks a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside in which the sugar moiety is replaced with a non-sugar moiety, e.g., a peptide nucleic acid (PNA) or a morpholino nucleic acid.
[0141] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, sugar modifications include modifications made by altering the substituent on the ribose ring to a group other than hydrogen or to the 2'-OH group naturally found in DNA and RNA nucleosides. In some embodiments, the substituent is introduced at the 2', 3', 4', or 5' position, or a combination thereof. In some embodiments, the nucleoside having a modified sugar moiety comprises a 2' modified nucleoside, e.g., a 2' substituted nucleoside. A 2' sugar modified nucleoside, in some embodiments, is a nucleoside having a substituent other than -H or -OH at the 2' position (2' substituted nucleoside) or comprises a 2' linked biradical, and includes 2' substituted nucleosides and LNA (2'-4' biradical bridged) nucleosides. Examples of 2'-substituted modified nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group comprises a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).
[0142] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the guide RNA comprises only modified sugars. In certain embodiments, the guide RNA comprises more than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises a 2'-O-methoxyethyl group. In some embodiments, the guide RNA comprises both an internucleoside linker modification and a nucleoside modification.
[0143] In some cases, the guide RNA comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a eukaryotic genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a fungal genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a plant genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a mammalian genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a human genomic polynucleotide sequence.
[0144] In some embodiments, the guide RNA is 30-250 nucleotides in length. In some embodiments, the guide RNA is more than 90 nucleotides in length. In some embodiments, the guide RNA is less than 245 nucleotides in length. In some embodiments, the guide RNA is 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides in length. In some embodiments, the guide RNA is about 30 to about 40, about 30 to about 50, about 30 to about 60, about 30 to about 70, about 30 to about 80, about 30 to about 90, about 30 to about 100, about 30 to about 120, about 30 to about 140, about 30 to about 160, about 30 to about 180, about 30 to about 200, about 30 to about 220, about 30 to about 240, about 50 to about 60, about 50 to about 70, about 50 to about 80, about 50 to about 90, about 50 to about 100, about 50 about 120, about 50 to about 140, about 50 to about 160, about 50 to about 180, about 50 to about 200, about 50 to about 220, about 50 to about 240, about 100 to about 120, about 100 to about 140, about 100 to about 160, about 100 to about 180, about 100 to about 200, about 100 to about 220, about 100 to about 240, about 160 to about 180, about 160 to about 200, about 160 to about 220, or about 160 to about 240 nucleotides in length.
[0145] In some cases, the left-hand recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78.In some cases, the left-hand recombinase sequence comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having 100% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78.
[0146] In some cases, the right recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.In some cases, the right recombinase sequence comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having 100% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.
[0147] In some cases, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence that comprises less than about 20 kilobases, less than about 15 kilobases, less than about 10 kilobases, or less than about 5 kilobases.
[0148] In some embodiments, the class 2 type V effector comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is at the N-terminus of the class 2 type V effector. In some embodiments, the NLS is at the C-terminus of the class 2 type V effector. In some embodiments, the NLS is at both the N-terminus and the C-terminus of the class 2 type V effector.
[0149] In some embodiments, the NLS comprises any one of SEQ ID NOs: 172-187, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 172-187.In some cases, the NLS comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having 100% identity to any one of SEQ ID NOs: 172-187.
[0150] [Table 1]
[0151] In some embodiments, the Cas effector complex further comprises a small prokaryotic ribosomal protein subunit, S15. In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 70% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 75% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 80% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 85% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 90% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 91% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 92% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 93% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 94% identity to any one of SEQ ID NOs: 161-163.In some cases, S15 is encoded by a sequence having at least about 95% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 96% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 97% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 98% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 99% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having 100% identity to any one of SEQ ID NOs: 161-163.
[0152] In some cases, S15 comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having 100% identity to any one of SEQ ID NOs:167-169.
[0153] In some embodiments, the Cas effector complex comprises one or more linkers linking the class 2 type V effector, the small prokaryotic ribosomal protein subunit S15, the transposase, the gRNA, or combinations thereof. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, or 400 amino acids. In some embodiments, the linker comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the linker is encoded by the sequence of SEQ ID NO: 166 or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NO: 166. In some embodiments, the linker is encoded by SEQ ID NO:166.
[0154] Fusion proteins In some embodiments, described herein is a system for translocating a cargo nucleotide sequence into a target nucleic acid site comprising a fusion protein or a nucleic acid encoding the fusion protein. In some embodiments, the fusion protein or the nucleic acid encoding the fusion protein comprises a class 2 type V effector, a small prokaryotic ribosomal protein subunit S15, a transposase, a gRNA, or a combination thereof. In some embodiments, the fusion protein comprises one or more transposases.
[0155] In some embodiments, a nuclear localization sequence (NLS) is fused to the class 2 type V effector. In some embodiments, the NLS is fused at the N-terminus of the class 2 type V effector. In some embodiments, the NLS is fused at the C-terminus of the class 2 type V effector. In some embodiments, the NLS is fused at both the N-terminus and the C-terminus of the class 2 type V effector.
[0156] In some embodiments, the NLS comprises any one of SEQ ID NOs: 172-187, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 172-187.In some cases, the NLS comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 172-187. In some cases, the NLS comprises a sequence having 100% identity to any one of SEQ ID NOs: 172-187.
[0157] In some embodiments, the fusion protein or a nucleic acid encoding the fusion protein comprises a fusion of S15 and a nuclear localization sequence (NLS). In some embodiments, the NLS is fused at the N-terminus of S15. In some embodiments, the NLS is fused at the C-terminus of S15. In some embodiments, the NLS is fused at both the N-terminus and the C-terminus of S15.
[0158] In some embodiments, the S15 fusion protein further comprises a cleavable peptide, hi some embodiments, the peptide is a 2A peptide.
[0159] In some embodiments, the S15 fusion protein is encoded by a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 70% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 75% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 80% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 85% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 90% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 91% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 92% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 93% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 94% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 95% identity to any one of SEQ ID NOs: 161-163.In some cases, S15 is encoded by a sequence having at least about 96% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 97% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 98% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having at least about 99% identity to any one of SEQ ID NOs: 161-163. In some cases, S15 is encoded by a sequence having 100% identity to any one of SEQ ID NOs: 161-163.
[0160] In some cases, S15 comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 167-169. In some cases, S15 comprises a sequence having 100% identity to any one of SEQ ID NOs:167-169.
[0161] In some embodiments, the NLS is fused to a transposase. In some embodiments, the transposase is TnsB, TnsC, or TniQ. In some embodiments, the transposase is TnsB. In some embodiments, the transposase is TnsC. In some embodiments, the transposase is TniQ. In some embodiments, the NLS is fused at the N-terminus of the transposase. In some embodiments, the NLS is fused at the C-terminus of the transposase. In some embodiments, the NLS is fused at both the N-terminus and the C-terminus of the transposase.
[0162] In some embodiments, the NLS is fused at the N-terminus of the transposase. In some embodiments, the NLS is fused at the C-terminus of the transposase. In some embodiments, the NLS is fused at the N- and C-terminus of the transposase.
[0163] In some embodiments, the fusion protein or a nucleic acid encoding the fusion protein comprises a gRNA (e.g., a dual gRNA or a single gRNA) described herein.
[0164] In some embodiments, the class 2 type V effector, small prokaryotic ribosomal protein subunit S15, transposase, gRNA, or fusion protein comprises a tag. In some embodiments, the tag is a polypeptide or a polynucleotide. In some embodiments, the tag is an affinity tag. Exemplary affinity tags include, but are not limited to, His tag, Flag tag, Myc tag, MBP tag, and GST tag.
[0165] In some embodiments, the fusion protein or gene editing system comprising class 2 V-type effector, small prokaryotic ribosomal protein subunit S15, transposase, single gRNA, or any combination thereof comprises a tag. In some embodiments, the tag is an affinity tag. Exemplary affinity tags include, but are not limited to, His tag, Flag tag, Myc tag, MBP tag, and GST tag.
[0166] In some embodiments, the class 2 type V effector, small prokaryotic ribosomal protein subunit S15, transposase, or fusion protein comprises a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, a TEV site, a C3 site, a factor Xa site, and an enterokinase site.
[0167] cell In certain embodiments, cells comprising the systems described herein are described herein.
[0168] In some embodiments, the cell is a eukaryotic cell (e.g., a plant cell, an animal cell, a protist cell, or a fungal cell), a mammalian cell (Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embryonic kidney (HEK), mouse myeloma (NS0), or a human retinal cell), an immortalized cell (e.g., a HeLa cell, a COS cell, a HEK-293T cell, an MDCK cell, a 3T3 cell, a PC12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, a N2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, a S2 cell, or a Heliothis virescens cell), a yeast cell (e.g., a Saccharomyces In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.
[0169] In some embodiments, the cells are A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.
[0170] Delivery and Vectors In some embodiments, disclosed herein are nucleic acid sequences encoding the MG64 system, including a class 2 type V effector, a small prokaryotic ribosomal protein subunit S15, a transposase, a gRNA, a fusion protein, or a gene editing system disclosed herein.
[0171] In some embodiments, the nucleic acid encoding the MG64 system is DNA, e.g., linear DNA, plasmid DNA, or minicircle DNA. In some embodiments, the nucleic acid encoding the MG64 system is RNA, e.g., mRNA.
[0172] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a nucleic acid-based vector. In some embodiments, the nucleic acid-based vector is a plasmid (e.g., a circular DNA molecule that can replicate autonomously inside a cell), a cosmid (e.g., a pWE or sCos vector), an artificial chromosome, a human artificial chromosome (HAC), a yeast artificial chromosome (YAC), a bacterial artificial chromosome (BAC), a P1-derived artificial chromosome (PAC), a phagemid, a phage derivative, a bacmid, or a virus. In some embodiments, the nucleic acid based vector is pSF-CMV-NEO-NH2-PPT-3XFLAG, pSF-CMV-NEO-COOH-3XFLAG, pSF-CMV-PURO-NH2-GST-TEV, pSF-OXB20-COOH-TEV-FLAG(R)-6His, pCEP4 pDEST27, pSF-CMV-Ub-KrYFP, pSF-CMV-FMDV-daGFP, pEF1a-mCherry-N1 vector, pEF1a-tdTomato vector, pSF-CMV-FMDV-Hygro, pSF-CMV-PGK-Puro, pMCP-tag(m), pSF-CMV-PURO-NH2-CMYC, pSF-OXB20-BetaGal, pSF-OXB20-Fluc, pSF-OXB20, pSF-Tac, pRI 101-AN The vector is selected from the list consisting of pCambia2301, pTYB21, pKLAC2, pAc5.1 / V5-His A, and pDEST8.
[0173] In some embodiments, the nucleic acid-based vector comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a minipromoter, an inducible promoter, a constitutive promoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments, the promoter is a U6 promoter. In some embodiments, the promoter is a CAG promoter. In some embodiments, the promoter is encoded by any one of SEQ ID NOs: 190-191, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 190-191.
[0174] In some embodiments, the nucleic acid based vector is a virus. In some embodiments, the virus is an alphavirus, parvovirus, adenovirus, AAV, baculovirus, dengue virus, lentivirus, herpes virus, poxvirus, anellovirus, bocavirus, vaccinia virus, or retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpes virus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the virus is a retrovirus.
[0175] In some embodiments, the AAV is AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-rh8, AAV-rh 10, AAV-rh20, AAV-rh39, AAV-rh74, AAV-rhM4-1, AAV-hu37, AAV-Anc80, AAV-Anc80L65, AAV-7m8, AAV-PHP-B, AAV-PHP-EB, AAV-2.5, AAV-2tYF, AAV-3B, AAV-LK03, AAV-HSC1, AAV-HSC2, AAV-HSC3, AAV-HSC4, AAV-HSC5, AAV-HSC6, AAV-HSC7, AAV-HSC8, AAV-HSC9, AAV-HSC10, AAV-HSC11, AAV-HSC12, AAV-HSC13, AAV-HSC14, AAV-HSC15, AAV-TT, AAV-DJ / 8, AAV-Myo, AAV-NP40, AAV-NP59, AAV-NP22, AAV-NP66, AAV-HSC16, or derivatives thereof. In some embodiments, the herpes virus is HSV type 1, HSV-2, VZV, EBV, CMV, HHV-6, HHV-7, or HHV-8.
[0176] In some embodiments, the virus is AAV1 or a derivative thereof. In some embodiments, the virus is AAV2 or a derivative thereof. In some embodiments, the virus is AAV3 or a derivative thereof. In some embodiments, the virus is AAV4 or a derivative thereof. In some embodiments, the virus is AAV5 or a derivative thereof. In some embodiments, the virus is AAV6 or a derivative thereof. In some embodiments, the virus is AAV7 or a derivative thereof. In some embodiments, the virus is AAV8 or a derivative thereof. In some embodiments, the virus is AAV9 or a derivative thereof. In some embodiments, the virus is AAV10 or a derivative thereof. In some embodiments, the virus is AAV11 or a derivative thereof. In some embodiments, the virus is AAV12 or a derivative thereof. In some embodiments, the virus is AAV13 or a derivative thereof. In some embodiments, the virus is AAV14 or a derivative thereof. In some embodiments, the virus is AAV15 or a derivative thereof. In some embodiments, the virus is AAV16 or a derivative thereof. In some embodiments, the virus is AAV-rh8 or a derivative thereof. In some embodiments, the virus is AAV-rh10 or a derivative thereof. In some embodiments, the virus is AAV-rh20 or a derivative thereof. In some embodiments, the virus is AAV-rh39 or a derivative thereof. In some embodiments, the virus is AAV-rh74 or a derivative thereof. In some embodiments, the virus is AAV-rhM4-1 or a derivative thereof. In some embodiments, the virus is AAV-hu37 or a derivative thereof. In some embodiments, the virus is AAV-Anc80 or a derivative thereof. In some embodiments, the virus is AAV-Anc80L65 or a derivative thereof. In some embodiments, the virus is AAV-7m8 or a derivative thereof. In some embodiments, the virus is AAV-PHP-B or a derivative thereof. In some embodiments, the virus is AAV-PHP-EB or a derivative thereof.In some embodiments, the virus is AAV-2.5 or a derivative thereof. In some embodiments, the virus is AAV-2tYF or a derivative thereof. In some embodiments, the virus is AAV-3B or a derivative thereof. In some embodiments, the virus is AAV-LK03 or a derivative thereof. In some embodiments, the virus is AAV-HSC1 or a derivative thereof. In some embodiments, the virus is AAV-HSC2 or a derivative thereof. In some embodiments, the virus is AAV-HSC3 or a derivative thereof. In some embodiments, the virus is AAV-HSC4 or a derivative thereof. In some embodiments, the virus is AAV-HSC5 or a derivative thereof. In some embodiments, the virus is AAV-HSC6 or a derivative thereof. In some embodiments, the virus is AAV-HSC7 or a derivative thereof. In some embodiments, the virus is AAV-HSC8 or a derivative thereof. In some embodiments, the virus is AAV-HSC9 or a derivative thereof. In some embodiments, the virus is AAV-HSC10 or a derivative thereof. In some embodiments, the virus is AAV-HSC11 or a derivative thereof. In some embodiments, the virus is AAV-HSC12 or a derivative thereof. In some embodiments, the virus is AAV-HSC13 or a derivative thereof. In some embodiments, the virus is AAV-HSC14 or a derivative thereof. In some embodiments, the virus is AAV-HSC15 or a derivative thereof. In some embodiments, the virus is AAV-TT or a derivative thereof. In some embodiments, the virus is AAV-DJ / 8 or a derivative thereof. In some embodiments, the virus is AAV-Myo or a derivative thereof. In some embodiments, the virus is AAV-NP40 or a derivative thereof. In some embodiments, the virus is AAV-NP59 or a derivative thereof. In some embodiments, the virus is AAV-NP22 or a derivative thereof. In some embodiments, the virus is AAV-NP66 or a derivative thereof. In some embodiments, the virus is AAV-HSC16 or a derivative thereof.
[0177] In some embodiments, the virus is HSV-1 or a derivative thereof. In some embodiments, the virus is HSV-2 or a derivative thereof. In some embodiments, the virus is VZV or a derivative thereof. In some embodiments, the virus is EBV or a derivative thereof. In some embodiments, the virus is CMV or a derivative thereof. In some embodiments, the virus is HHV-6 or a derivative thereof. In some embodiments, the virus is HHV-7 or a derivative thereof. In some embodiments, the virus is HHV-8 or a derivative thereof.
[0178] In some embodiments, the nucleic acid encoding the MG64 system is delivered by a non-nucleic acid based delivery system (e.g., a non-viral delivery system). In some embodiments, the non-viral delivery system is a liposome. In some embodiments, the nucleic acid is associated with a lipid. The nucleic acid associated with a lipid is in some embodiments encapsulated in the aqueous interior of the liposome, interspersed within the lipid bilayer of the liposome, attached to the liposome via a linking molecule associated with both the liposome and the nucleic acid, entrapped in the liposome, complexed with the liposome, dispersed in a solution containing lipid, mixed with lipid, combined with lipid, contained as a suspension in lipid, contained in or complexed with micelles, or otherwise associated with lipid. In some embodiments, the nucleic acid is included in a lipid nanoparticle (LNP).
[0179] In some embodiments, the fusion protein or genome editing system is introduced into the cell in any suitable manner, either stably or transiently. In some embodiments, the fusion protein or genome editing system is transfected into the cell. In some embodiments, the cell is transduced or transfected with a nucleic acid construct encoding the fusion protein or genome editing system. For example, the cell is transduced (e.g., with a virus encoding the fusion protein or genome editing system) or transfected with a nucleic acid encoding the fusion protein or genome editing system, or a translated fusion protein or genome editing system (e.g., with a plasmid encoding the fusion protein or genome editing system). In some embodiments, the transduction is stable or transient transduction. In some embodiments, the cell expressing or containing the fusion protein or genome editing system is transduced or transfected with one or more gRNA molecules, for example, when the fusion protein or genome editing system contains a CRISPR nuclease. In some embodiments, a plasmid expressing a fusion protein or a genome editing system is introduced into a cell through electroporation, transient (e.g., lipofection), stable genome integration (e.g., piggybac), and viral transduction (e.g., lentivirus or AAV), or other methods known to those skilled in the art. In some embodiments, the gene editing system is introduced into a cell as one or more polypeptides. In some embodiments, delivery is achieved through the use of an RNP complex. Methods for delivering polypeptides and / or RNPs into cells are known in the art, for example, by electroporation or by cell squeezing.
[0180] Exemplary methods of nucleic acid delivery include lipofection, nucleofection, electroporation, stable genomic integration (e.g., piggybac), microinjection, biolistec, virosomes, liposomes, immunoliposomes, polycations or lipid-nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Pat. Nos. 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam™, Lipofectin™, and SF Cell Line 4D-Nucleofector X Kit™ (Lonza)). Cationic and neutral lipids that are suitable for efficient receptor-recognition lipofection of polynucleotides include the lipids of WO91 / 17424 and WO91 / 16024. In some embodiments, delivery is to a cell (e.g., in vitro or ex vivo administration) or to a target tissue (e.g., in vivo administration). In some embodiments, the nucleic acid is contained in a liposome or nanoparticle that specifically targets the host cell.
[0181] Additional methods for delivery of nucleic acids into cells are known to those of skill in the art, see, e.g., US2003 / 0087817.
[0182] In some embodiments, the disclosure provides a cell comprising a vector or a nucleic acid described herein. In some embodiments, the cell expresses a gene editing system or a portion thereof. In some embodiments, the cell is a human cell. In some embodiments, the cell is genome edited ex vivo. In some embodiments, the cell is genome edited in vivo.
[0183] Methods for the rearrangement The present disclosure provides a method for translocating a cargo nucleotide sequence into a target nucleic acid site. In some embodiments, the method comprises expressing a system described herein in a cell or introducing a system described herein into a cell. In some embodiments, the method comprises contacting a cell with a system described herein.
[0184] In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising a cargo nucleotide sequence with a Cas effector complex comprising a class 2 V-type Cas effector and at least one engineered guide polynucleotide configured to hybridize to a target nucleotide sequence. In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising the cargo nucleotide sequence with a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB subunit. In some embodiments, the method comprises contacting a double-stranded nucleic acid comprising the cargo nucleotide sequence with a double-stranded target nucleic acid comprising a target nucleic acid site.
[0185] In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a right transposase recognition sequence. In some cases, the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence. In some embodiments, the method further comprises a PAM sequence compatible with the nuclease adjacent to the target nucleic acid site. In some cases, the PAM sequence is located 3' to the target nucleic acid site.
[0186] In some cases, the engineered guide polynucleotide is configured to bind to a Class 2 V-type Cas effector. In some cases, the Class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200, or a variant thereof. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200. In some cases, class 2 V-type Cas effectors include a polypeptide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.In some cases, the class 2 V-type Cas effector comprises a polypeptide comprising a sequence having 100% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.
[0187] In some cases, the TnsB subunit comprises a polypeptide having a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 2, 13, 17, and 65.In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 2, 13, 17, and 65. In some cases, the TnsB component comprises a polypeptide comprising a sequence having 100% identity to any one of SEQ ID NOs: 2, 13, 17, and 65.
[0188] In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs:3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, the Tn7-type transposase complex comprises at least one polypeptide comprising a sequence having 100% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.
[0189] In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 70% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 75% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 85% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 90% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 91% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 92% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 93% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 94% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 95% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 96% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 97% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 98% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least about 99% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67. In some cases, a Tn7-type transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having 100% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.
[0190] In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs:5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 70% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 75% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 85% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 90% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 91% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 92% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 93% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 94% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 95% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 96% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 97% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 98% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 99% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202. In some cases, the engineered guide polynucleotide comprises a sequence comprising at least about 46-80 contiguous nucleotides having 100% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.
[0191] In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that are identical to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 70% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165.In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 75% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 80% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 85% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 90% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 91% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 92% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having at least about 93% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165.In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 94% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 95% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that comprises at least about 46-80 contiguous nucleotides that have at least about 96% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 97% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 98% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence that includes at least about 46-80 contiguous nucleotides that have at least about 99% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165. In some cases, the engineered guide polynucleotide is a guide RNA and comprises a sequence comprising at least about 46-80 contiguous nucleotides having 100% identity to any one of SEQ ID NOs: 45-63, 68-75, 96-103, and 165.
[0192] In some cases, the left-hand recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78.In some cases, the left-hand recombinase sequence comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78. In some cases, the left-hand recombinase sequence comprises a sequence having 100% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78.
[0193] In some cases, the right recombinase sequence comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 91% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 92% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 93% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.In some cases, the right recombinase sequence comprises a sequence having at least about 94% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93. In some cases, the right recombinase sequence comprises a sequence having 100% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.
[0194] In some cases, the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence that comprises less than about 20 kilobases, less than about 15 kilobases, less than about 10 kilobases, or less than about 5 kilobases.
[0195] use The disclosed system can be used for a variety of applications, such as, for example, nucleic acid editing (e.g., gene editing) or binding (e.g., sequence-specific binding) to nucleic acid molecules. Such systems can be used, for example, to correct (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivated enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to render them unable to infect host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and / or to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors.
[0196] kit In some embodiments, the disclosure provides kits that include one or more nucleic acid constructs encoding various components of the fusion proteins or genome editing systems described herein, including, for example, nucleotide sequences encoding components of a fusion protein or genome editing system capable of modifying a target DNA sequence. In some embodiments, the nucleotide sequences include a heterologous promoter that drives expression of an RNA genome editing system component.
[0197] In some embodiments, the fusion protein or gene editing system comprising the class 2 type V effector, small prokaryotic ribosomal protein subunit S15, transposase, single gRNA, or any combination thereof disclosed herein is incorporated into a pharmaceutical, diagnostic, or research kit to facilitate its use in therapeutic, diagnostic, or research applications. The kit may include one or more containers housing any of the vectors disclosed herein, and instructions for use.
[0198] The kits can be designed to facilitate the use of the methods described herein by researchers and can take many forms. Each of the compositions of the kit can be provided in liquid form (e.g., in solution) or in solid form (e.g., dry powder), if applicable. In certain cases, some of the compositions may be configurable or otherwise processable (e.g., into an active form), for example, by the addition of suitable solvents or other species (e.g., water or cell culture medium), which may or may not be provided with the kit. As used herein, "instructions" defines an instructional and / or promotional component, and can typically involve written instructions on or associated with the packaging of the present disclosure. Instructions can also include any verbal or electronic instructions provided in any manner, such as audiovisual (e.g., videotape, DVD, etc.), internet, and / or web-based communication, etc., such that the user clearly recognizes that the instructions are related to the kit. The written instructions, in some embodiments, are in a form prescribed by a government agency that regulates the manufacture, use, or sale of pharmaceutical or biological products, and the instructions may also reflect approval by the agency of the manufacture, use, or sale for animal administration. EXAMPLES
[0199] The following examples are given for the purpose of illustrating various embodiments of the present disclosure, and are not intended to limit the present disclosure in any manner. The examples, together with the methods described herein, are representative of currently preferred embodiments, are exemplary, and are not intended as limitations on the scope of the present disclosure. Modifications therein and other uses encompassed within the spirit of the present disclosure as defined by the scope of the claims will occur to those skilled in the art.
[0200] Example 1 - (General Protocol) PAM Sequence Identification / Confirmation of the System Described Herein The putative endonucleases were expressed in an E. coli lysate-based expression system. The PAM sequences were determined by sequencing a plasmid containing randomly generated potential PAM sequences that could be cleaved by the putative nuclease. In this system, the E. coli codon-optimized nucleotide sequence encoding the putative nuclease was transcribed and translated in vitro from a PCR fragment under the control of a T7 promoter. A second PCR fragment carrying a T7 promoter followed by a minimal CRISPR array composed of a repeat-spacer-repeat sequence was transcribed in the same reaction. Successful expression of the endonuclease sequence and the repeat-spacer-repeat sequence in the in vitro expression system, followed by CRISPR array processing, provided an active in vitro CRISPR nuclease complex.
[0201] A library of target plasmids containing spacer sequences matching those in the minimal array preceded by 8N mixed bases (potential PAM sequences) was incubated with the products of the in vitro expression reaction. After 1–3 h, the reaction was stopped and DNA was recovered via a DNA clean-up kit. Adapter sequences were blunt-end ligated to DNA with active PAM sequences cleaved by the endonuclease, while uncleaved DNA was inaccessible for ligation. DNA segments containing active PAM sequences were then amplified by PCR using primers specific for the library and adapter sequences. PCR amplification products were resolved on a gel to identify amplicons corresponding to cleavage events. Amplified segments of the cleavage reaction were also used as templates for preparation of NGS libraries or as substrates for Sanger sequencing. Sequencing this resulting library, a subset of the starting 8N library, revealed sequences with PAM activity compatible with the CRISPR complex. For PAM testing with treated RNA constructs, the same procedure was repeated except that in vitro transcribed RNA was added along with the plasmid library and the minimal CRISPR array template was omitted.
[0202] Analysis of the intergenic regions surrounding the Cas effectors and CRISPR arrays identified potential repeat suppressor sequences that correspond to the double-stranded sequence of tracrRNA. TracrRNA and crRNA repeats were folded and trimmed, and a GAAA tetraloop sequence was added to maintain the stem-loop region of the crRNA-tracrRNA complex.
[0203] Example 2A - In vitro targeted integrase activity Integrase activity was assayed using previously identified PAMs, but could alternatively be performed with PAM library substrates of reduced efficiency. One arrangement of components for in vitro testing involved three plasmids other than the one containing the donor sequence: (1) an expression plasmid with an effector (or effectors) under a T7 promoter, (2) an expression plasmid with a transposase gene under a T7 promoter, sgRNA or crRNA and tracrRNA, (3) a target plasmid that contained a spacer site and an appropriate PAM, and (4) a donor plasmid containing the necessary left end (LE) and right end (RE) DNA sequences for transposition around a cargo gene (e.g., a selection marker such as a Tet resistance gene). An in vitro transcription / translation system (e.g., an E. coli lysate or reticulocyte lysate-based system) was used to express the effector and transposase genes. After expression, RNA, target DNA, and donor DNA were added and incubated to allow transposition to occur. Transcription was detected via PCR across the transposase site junction, with one primer on the target DNA and one on the donor DNA. The resulting PCR products were sequenced via NGS to determine the exact insertion topology relative to the sgRNA / crRNA target site. Primers were positioned downstream such that various insertion sites could be accommodated and detected. Primers were designed such that integration could be detected in either orientation of the cargo and on either side of the spacer, as integration direction has also not been previously documented.
[0204] Incorporation efficiency was measured via quantitative PCR (qPCR) of the experimental output of target DNA with incorporated cargo, normalized to the amount of unmodified target DNA, also measured via qPCR measurement.
[0205] This assay may also be performed with purified protein components rather than from lysate-based expression. In this case, proteins were expressed in E. coli protease-deficient B strain under a T7 inducible promoter, cells were lysed using sonication, and His-tagged proteins of interest were purified using Ni-NTA affinity chromatography on an FPLC. Purity was determined using SDS-PAGE and densitometry of resolved protein bands on Coomassie-stained acrylamide gels. Proteins were desalted in a storage buffer consisting of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol at pH 7.5 (or other buffer determined for maximum stability) and stored at -80°C. After purification, the effector and transposase were added to the sgRNA, target DNA, and donor DNA described above in reaction buffer, e.g., 26 mM HEPES at pH 7.5, 4.2 mM TRIS at pH 8, 50 μg / mL BSA, 2 mM ATP, 2.1 mM DTT, 0.05 mM EDTA, 0.2 mM MgCl2, 28 mM NaCl, 21 mM KCl, 1.35% glycerol (final pH 7.5), supplemented with 15 mM Mg(OAc)2.
[0206] Example 2B - In vitro activity Targeted Nucleases In situ expression and protein sequence analysis showed that several RNA-guided effectors were active nucleases: they contained predicted endonuclease-associated domains (matching the RuvC and HNH endonuclease domains) and / or predicted HNH and RuvC catalytic residues.
[0207] Candidate activity was tested with the engineered single guide RNA sequences using an in vitro expression system and in vitro transcribed RNA. Active proteins that successfully cleaved the library gave rise to a band of approximately 170 bp in the gel.
[0208] DNA integration and transposition A transposon is predicted to be active when the genomic sequence encoding the transposon contains one or more protein sequences with transposase and / or integrase functions within the left and right ends of the transposon. Tn7 transposons as defined herein may contain the catalytic transposase TnsB, but may also contain TnsA, TnsC, TnsD, TnsE, TniQ, and / or other transposases or integrases. The transposon ends contain predicted transposase binding sites that contain direct and / or inverted repeats of 15 bp to 150 bp in length flanking the transposase protein and other "cargo" genes. Protein sequence analysis shows that the transposases contain an integrase domain, a transposase domain, and / or transposase catalytic residues, suggesting that they are active (e.g., FIG. 4A).
[0209] Targeted DNA integration Putative CRISPR-associated transposons (CASTs) contain DNA and / or RNA-targeting CRISPR nucleases or effectors, as well as proteins with predicted transposase function in the vicinity of the CRISPR array. In some systems, the nuclease is predicted to be active based on the presence of an endonuclease-associated catalytic domain and / or catalytic residues.
[0210] In some systems, effectors are predicted to be inactive based on the absence of endonuclease domains and / or catalytic residues, although they have homology with documented CRISPR effector proteins.Transposases are predicted to associate with effectors when CRISPR loci (inactive CRISPR nucleases and arrays) and transposase proteins are located within the left and right ends of predicted transposons (Figure 4A).In this case, effectors are predicted to guide DNA integration to specific genomic locations based on guide RNAs.
[0211] CAST activity was tested using five types of components: (1) Cas effector proteins expressed by an in vitro expression system, (2) target DNA fragments or plasmids containing a target sequence and a PAM corresponding to the Cas enzyme, (3) donor DNA fragments containing markers or fragments of DNA flanking the LE and RE of the transposase system in the DNA fragment or plasmid, (4) any combination of transposase proteins expressed using an in vitro expression system, and (5) engineered in vitro transcribed single guide RNA sequences. Active systems with successful transposition of the donor fragments were assayed by PCR amplification of the donor-target junctions.
[0212] After performing the transposition reaction, PCR amplification of the junction showed that proper donor-target formation had occurred and that the transposition reaction was sg-dependent (Figure 6). PCR amplification of reactions #3 and #4 showed that both orientations of the donor to the target had been achieved, i.e., LE closer to the PAM and RE closer to the PAM. Although both transposition orientations were achieved, there was a preference for donor integration in the target with the LE closer to the PAM, as represented by the strong bands present for reactions #4 and #5.
[0213] Sanger sequencing of the preferred orientation products was performed. Of the integrations that occurred with the LE closer to the PAM, there was a clear degradation of the sequencing chromatogram signal from either the forward or reverse orientation across the target / donor junction. This indicated that of the products oriented with the LE closer to the PAM, integration occurred over a range of nucleotides, with the major product of the LE closer to the PAM as a 61 bp integration from the PAM (Figure 7a). Sequencing of the donor output across the donor-target junction defined the composition of the essential outer border of the LE and RE sequences (Figures 7A and 7B). Further investigation of the LE and RE domains will determine the inner limits of the LE and RE sequences that are essential for transposition. Sequencing of the RE on the product of the LE closer to the PAM showed a 3 bp overlap downstream of the donor RE (Figure 7B). This is due, in part, to a Tn7 transposase integration event that cut and ligated the donor fragment at staggered cut sites. The 3 bp overlap is smaller than the expected 5 bp overlap from other Tn7 transposases.
[0214] Sanger sequencing of PCR amplification products against the 8N library of target plasmids also revealed the PAM preference of the MG64-1 effector as nGTn / nGTt on the 5' end of the spacer (Figure 7C). NGS analysis of the PAM library targets confirmed the nGTn motif preference at the 5' end.
[0215] Example 3 - Predicted RNA folding The predicted RNA fold of the active single RNA sequence was calculated at 37° using the method of Andronescu 2007. All hairpin loop secondary structures were singly removed from the structure and compiled iteratively into a smaller single guide. In the second approach, the tracrRNA of MG64-1 was aligned to the documented Vk-type tracrRNA and the region of the unique insertion was mutated from the single guide and minimized by 57 bases. Figure 12A illustrates the predicted structure of the MG64-1 sgRNA. Figure 12B illustrates the predicted structure of the MG64-3 sgRNA. Figure 12C illustrates the predicted structure of the MG64-5 sgRNA.
[0216] Example 4 - Transposon end verification via gel shift Transposon ends were tested for TnsB binding via electrophoretic mobility shift assay (EMSA). In this case, potential LEs or REs were synthesized as DNA fragments (100-500 bp) and end-labeled with FAM via PCR using FAM-labeled primers. TnsB proteins were synthesized in an in vitro transcription / translation system. After synthesis, 1 μL of TnsB protein was added to 50 nM of labeled RE or LE in a 10 μL reaction in binding buffer (20 mM HEPES at pH 7.5, 2.5 mM Tris at pH 7.5, 10 mM NaCl, 0.0625 mM EDTA, 5 mM TCEP, 0.005% BSA, 1 ug / mL poly(dI-dC), and 5% glycerol). The ligation was incubated at 30° for 40 min, then 2 uL of 6X loading buffer (60 mM KCl, 10 mM Tris at pH 7.6, 50% glycerol) was added. Binding reactions were resolved and visualized on a 5% TBE gel. A shift in the LE or RE in the presence of TnsB was due to successful binding and indicated transposase activity (Figure 24).
[0217] Example 5 - Integrase activity in E. coli Transformation of E. coli with agents capable of inducing double-stranded breaks in the E. coli genome results in cell death, as E. coli lacks the ability to efficiently repair genomic double-stranded DNA breaks. Exploiting this phenomenon, endonuclease or effector-assisted integrase activity was tested in E. coli by recombinantly expressing either endonuclease or effector-assisted integrase and guide RNA (e.g., as determined in Example 3) in a target strain with spacer / target and PAM sequences integrated into its genomic DNA.
[0218] The engineered strains were then transformed with plasmids containing nucleases or effectors with single guide RNAs, plasmids expressing integrases and accessory genes, and plasmids containing a temperature-sensitive origin of replication with a selectable marker flanked by left-extremity (LE) and right-extremity (RE) transposon motifs for integration. Transformants induced for expression of these genes were then screened for transfer of the marker to the genomic target by selection at the restrictive temperature for plasmid replication, and marker integration within the genome was confirmed by PCR.
[0219] An unbiased approach was used to screen for off-target integration. Briefly, purified gDNA was fragmented with Tn5 transposase or sheared, and then the DNA of interest was PCR amplified using primers specific for the ligated adapter and selectable marker. The amplicons were then prepared for NGS sequencing. Analysis of the resulting sequences was trimmed from the transposon sequence and flanking sequences were mapped to the genome to determine the insertion location and to determine the off-target insertion rate.
[0220] Example 6 - Colony PCR screening for transposase activity For testing of nuclease or effector-assisted integrase activity in bacterial cells, strain MGB0032 was constructed from BL21(DE3) E. coli cells engineered to contain a target and corresponding PAM sequence specific for MG64_1. MGB0032 E. coli cells were then transformed with pJL56 (a plasmid expressing the MG64_1 effector and helper suite, ampicillin resistant) and pTCM 64_1 sg, a chloramphenicol resistant plasmid expressing a single guide RNA sequence of the engineered target of interest driven by a T7 promoter.
[0221] MGB0032 cultures containing both plasmids were then grown to saturation, diluted at least 1:10 into growth culture containing the appropriate antibiotic, and incubated at 37°C to an OD of approximately 1. Cells from this growth step were made electrocompetent and transformed with streamlined64_1 pDonor, a plasmid carrying a tetracycline resistance marker flanked by left-end (LE) and right-end (RE) transposon motifs for integration. Electroporated cells were then plated on LB-agar-ampicillin-chloramphenicol-tetracycline and allowed to recover for 2 hours on LB medium in the presence or absence of IPTG at a final concentration of 100 μM before being incubated at 37°C for 4 days. A sterile toothpick was used to sample each resulting CFU, which was mixed into water. To this solution was added Q5 High Fidelity PCR mastermix and primers LA155 (5'-GCTCTTCCGATCTNNNNNGATGAGCGCATTGTTAGATTTCAT-3') and oJL50 (5'-AAACCGACATCGCAGGCTTC-3'). These primers flank the predicted insertion junction. The predicted product size was 609 bp. DNA amplified PCR products were visualized on a 2% agarose gel. Sanger sequencing of the PCR products confirmed the transposition event.
[0222] Example 7 - Intracellular expression / in vitro assay To test the functionality of the NLS constructs in a physiologically relevant environment, constructs cloned with active NLS-tagged CAST components were integrated into K562 cells using lentiviral transduction. Briefly, constructs cloned into lentiviral transfer plasmids were transfected into 293T cells with envelope and packaging plasmids, and virus containing supernatants were harvested from the medium after 72 hours of incubation. The medium containing the virus was then incubated with K562 cell lines containing 8 μg / mL polybrene for 72 hours, and transfected cells were then selected for integration en masse using puromycin at 1 μg / mL for 4 days. The selected cell lines were harvested at the end of the 4 days and differentially lysed for nuclear and cytoplasmic fractions. Subsequent fractions were then tested for transposition competence using a complementary set of in vitro expression components.
[0223] Ten million cells were centrifuged and washed once with 1x PBS, pH 7.4. The supernatant wash was aspirated completely into the cell pellet and flash frozen at -80C for 16 hours. After thawing on ice, the cell pellet size was measured by mass and the proteins in the cell fraction were naturally extracted using appropriate extraction volumes of cell fraction and nuclear extraction reagent. Briefly, cytoplasmic extraction reagent was used at 1:10 mass of cells to volume of extraction reagent. The cell suspension was mixed by vortexing and lysed with non-ionic detergent. The cells were then centrifuged at 16,000xg for 5 minutes at 4°C. The cytoplasmic extraction supernatant was then decanted and saved for in vitro testing. Nuclear extraction reagent was then added at 1:2 original cell mass to nuclear extraction reagent and incubated on ice for 1 hour with intermittent vortexing. The nuclear suspension was then centrifuged at 16,000xg for 10 minutes at 4°C, and the supernatant nuclear extract was decanted and tested for in vitro transposition activity. In vitro transposition reactions were performed with complementary sets of in vitro expressed proteins, donor DNA, pTarget, and buffers, using 4μL of each cell and nuclear extract for each condition. Evidence of transposition activity was assayed by PCR amplification of the donor-target junction.
[0224] Example 8 - Activity in mammalian cells (predictive) To demonstrate targeting and cleavage activity in mammalian cells, nuclear localization sequences are fused to the C-terminus of each of the nuclease or effector protein and integrase protein, and the fusion protein is purified. A single guide RNA targeting the genomic locus of interest is synthesized and incubated with the nuclease / effector protein to form a ribonucleoprotein complex. Cells are transfected with a plasmid containing a selectable neomycin resistance marker (NeoR) or a fluorescent marker flanked by left-end (LE) and right-end (RE) motifs, allowed to recover for 4-6 hours, and then electroporated with the nuclease RNP and integrase protein. Integration of the plasmid into the genome is quantified by counting G418-resistant colonies or by fluorescence-activated cell cytometry. Genomic DNA is extracted 72 hours after electroporation and used for preparation of NGS libraries. Off-target frequency is assayed by shearing the genome and preparing amplicons of the transposon marker and flanking DNA for NGS library preparation. At least 40 different target sites are selected to test the activity of each targeting system.
[0225] Example 9 - Targeted Nuclease Activity In situ expression and protein sequence analysis suggested that several RNA-guided effectors were active nucleases: they contained predicted endonuclease-associated domains (matching RuvC and HNH endonuclease domains) and predicted HNH and RuvC catalytic residues (Figure 4A).
[0226] Candidate activity was tested with the engineered single guide RNA sequences using an in vitro expression system and in vitro transcribed RNA. Active proteins that successfully cleaved the library gave rise to a band of approximately 170 bp in the gel.
[0227] Example 10 - Identification of transposons A transposon is predicted to be active when it contains one or more protein sequences with transposase and / or integrase functions between the left and right ends of the transposon. Tn7 transposons as defined herein contain the catalytic transposase TnsB, but may also contain TnsA, TnsC, TnsD, TnsE, TniQ, and / or other transposases or integrases. The transposon ends contain predicted transposase binding sites that contain direct and / or inverted repeats of 15 bp to 150 bp in length flanking the transposase protein and other "cargo" genes. Protein sequence analysis shows that the transposases contain an integrase domain, a transposase domain, and / or transposase catalytic residues, suggesting that they are active (e.g., Figures 4A and 5A).
[0228] Example 11 - Identification of CRISPR-associated transposons Putative CRISPR-associated transposons (CASTs) contain DNA and / or RNA-targeting CRISPR effectors and proteins with predicted transposase function in the vicinity of CRISPR arrays. In some systems, effectors are predicted to have nuclease activity based on the presence of endonuclease-associated catalytic domains and / or catalytic residues (e.g., FIG. 4A). Transposases were predicted to associate with effectors when CRISPR loci (inactive CRISPR nucleases and arrays) and transposase proteins are located within the left and right ends of predicted transposons (e.g., FIG. 4B and FIG. 4C). In this case, effectors were predicted to direct DNA integration to specific genomic locations based on guide RNAs.
[0229] In some systems, effectors were predicted to be inactive based on the absence of endonuclease domains and / or catalytic residues, although they shared homology with documented CRISPR effector proteins (Figure 5A). Transposases were predicted to associate with effectors when the CRISPR locus (inactive CRISPR nucleases and arrays) and the transposase protein was located within the left and right ends of the predicted transposon (Figures 5A and 5B).
[0230] Example 12 - CAST Identification CRISPR-associated transposons (CASTs) are a transposon-containing system that has evolved to interact with the CRISPR system to promote targeted integration of DNA cargo.
[0231] CAST is a genomic sequence that encodes one or more protein sequences involved in DNA transposition within the characteristic left and right ends of the transposon. Tn7 transposons, as defined herein, contain the catalytic transposase TnsB, but may also contain the catalytic transposase TnsA, the loader proteins TnsC or TniB, and the target recognition proteins TnsD, TnsE, TniQ, and / or other transposon-associated components. The transposon ends contain predicted transposase binding sites, including direct and / or inverted repeats of 15 bp to 150 bp in length, flanking the transposon machinery and other "cargo" genes.
[0232] In addition, CAST also encodes DNA and / or RNA targeting CRISPR nuclease or effector near CRISPR array. In some systems, effector is predicted to be active nuclease based on the presence of endonuclease-associated catalytic domain and / or catalytic residue. In some systems, effector has sequence similarity with documented CRISPR effector protein, but is predicted to be inactive based on the absence of endonuclease domain and / or catalytic residue. Transposon is predicted to be associated with effector when CRISPR locus and transposon-associated protein are located within the left and right ends of predicted transposon. In this case, effector is predicted to guide DNA integration to specific genomic location based on guide RNA.
[0233] Example 13 - Class 2 Cas12K CAST The Cas12k CAST system encodes a nuclease-deficient CRISPR Cas12k effector, a CRISPR array, tracrRNA, and a Tn7-like transposition protein. Cas12k effectors are phylogenetically diverse, and features that confirm their association with CAST have been identified for some (Figure 8). For example, the left end of the transposon was identified downstream from the MG64-3 CRISPR locus, as indicated by a terminal inverted repeat and a self-matching spacer sequence (Figure 11A). The Cas12k CAST CRISPR repeat (crRNA) contains the conserved motif 5'-GNNGGNNTGAAAG-3' (Figure 9). Short repeat-repeat suppressors (RARs) within the crRNA motif aligned with distinct regions of the tracrRNA (Figures 9 and 10), and the RAR motifs appeared to define the start and end of the tracrRNA (e.g., for MG64-1, the 5' end of the tracrRNA contained RAR1 (TTTC) and the 3' end contained RAR2 (CCNNC) (Figures 10A-1 and 10A-2).
[0234] Example 14 - Transposon end prediction Transposon ends were inferred from the intergenic regions flanking the effector and transposon machinery. For example, for Cas12k CAST, the intergenic regions located directly upstream from TnsB and directly downstream from the CRISPR locus were predicted to contain the left and right ends (LE and RE) of the Tn7 transposon.
[0235] Direct and inverted repeats (DR / IR) of approximately 12 bp were predicted on the contigs with up to two mismatches. In addition, short (approximately 10-20 bp) DR / IRs flanking the CAST transposon were found using the Dotplot algorithm. Matching DR / IRs located in intergenic regions flanking the CAST effector and transposon genes are predicted to encode transposon binding sites. LEs and REs extracted from the intergenic regions, encoding putative transposon binding sites, were aligned to define transposon end boundaries. Putative transposon LE and RE ends are: a) regions located within 400 bp upstream and downstream from the first and last predicted transposon-encoded genes, b) regions sharing multiple short inverted repeats, and c) regions sharing >65% nucleotide identity.
[0236] Example 15 - Single Guide Design Analysis of the intergenic regions surrounding the Cas effectors and CRISPR arrays identified potential repeat suppressor sequences and a conserved "CYCC(n6)GGRG" stem-loop structure adjacent to the repeat suppressor sequence corresponding to the double-stranded sequence of tracrRNA (Figure 11B). TracrRNA and crRNA repeats were folded and trimmed, and a tetraloop sequence of GAAA was added to maintain the stem-loop region of the crRNA-tracrRNA complementary sequence.
[0237] Example 16 - In vitro integration activity using targeted nucleases In situ expression and protein sequence analysis showed that several RNA-guided effectors were active nucleases. They contained predicted endonuclease-associated domains (matching the RuvC and HNH endonuclease domains) and / or predicted HNH and RuvC catalytic residues. Candidate activity was tested using an in vitro expression system and in vitro transcribed RNA with engineered single guide RNA sequences. Active proteins that successfully cleaved the library gave rise to a band of approximately 170 bp in the gel.
[0238] Example 17 - Programmable DNA integration CAST activity was tested using five types of components: (1) Cas effector protein expressed by an in vitro expression system (SEQ ID NO:1), (2) target DNA fragment or plasmid (SEQ ID NO:31) containing a target sequence and a PAM corresponding to the Cas enzyme, (3) donor DNA fragment (SEQ ID NO:8-11) containing a marker or fragment of DNA flanking the LE and RE of the transposase system in the DNA fragment or plasmid, (4) any combination of transposase proteins expressed using an in vitro expression system (SEQ ID NO:2-4), and (5) an engineered in vitro transcribed single guide RNA sequence (SEQ ID NO:5). Active systems with successful transposition of the donor fragment were assayed by PCR amplification of the donor-target junction.
[0239] After performing the transposition reaction, PCR amplification of the junction showed that proper donor-target formation had occurred and that the transposition reaction was sg-dependent (Figure 9). PCR amplification of reactions #3 and #4 showed that both orientations of the donor to the target had been achieved, i.e., LE closer to the PAM and RE closer to the PAM. Although both transposition orientations occurred, there appeared to be a preference for donor integration in the target with the LE closer to the PAM, as represented by the strong bands present for reactions #4 and #5.
[0240] Sanger sequencing of the preferred orientation products was performed. Of the integrations that occurred with the LE closer to the PAM, there was a clear degradation of the sequencing chromatogram signal from either the forward or reverse orientation across the target / donor junction. This indicated that of the products oriented with the LE closer to the PAM, integration occurred over a range of nucleotides, with the major product of the LE closer to the PAM as a 61 bp integration from the PAM (Figure 10a). Sequencing of the donor output across the donor-target junction defined the composition of the essential outer boundary of the LE and RE sequences (Figure 10a, Figure 10b). Sequencing of the RE on the product of the LE closer to the PAM showed a 3 bp overlap downstream of the donor RE (Figure 10b). This is due in part to a Tn7 transposase integration event that cut and ligated the donor fragment at the staggered cut site. The 3 bp overlap is smaller than the expected 5 bp overlap from other Tn7 transposases.
[0241] Sanger sequencing of PCR amplification products against the 8N library of target plasmids also showed the PAM preference of the MG64-1 effector as nGTn / nGTt on the 5' end of the spacer (Figure 10c). NGS analysis of the PAM library targets confirmed the nGTn motif preference at the 5' end.
[0242] Further development of single-guide studies confirmed the activity of MG64-1 with the new sgRNA scaffold (Figure 13).
[0243] Example 18 - Determining the Embedded Window The PCR junctions of the amplified PAM were indexed and sequenced for the NGS libraries. Reads were mapped and quantified using CRISPResso using the amplicon sequence of the putative transposition sequence with an integration distance of 60 bp from the PAM (guideseq=20 bp 3' end of LE or RE, window center=0, window size=20). Indel histograms were normalized to the total indel reads detected and the frequency was plotted against the 60 bp reference sequence (Figure 14).
[0244] Both PCR reaction 5 (LE proximal to the PAM, top panel of FIG. 14) and PCR 4 (RE distal to the PAM, bottom panel of FIG. 14) were plotted over sequence and distance from the PAM for MG64-1. Analysis of the integration window indicates that 95% of integrations that occurred at the spacer PAM site were within a 10 bp window 58-68 nucleotides away from the PAM. The difference in integration distance between distal and proximal frequencies reflected integration site overlap, i.e., an overlap of 3-5 base pairs as a result of the staggered nuclease activity of the transposase upon integration.
[0245] Example 19 - Colony PCR screening for transposase activity Transposition activity was assayed via colony PCR screening. After transformation with the pDonor plasmid, E. coli were plated on LB-agar containing ampicillin, chloramphenicol, and tetracycline. Picked CFU were added to a solution containing PCR reagents and primers flanking the selected insertion junction. PCR reactions of integration products were visible on a gel (Figure 15). Sequencing results of picked colony PCR products confirmed that they represented transposition events, as they spanned the junction between the LE and PAM at the engineered target site within the lacZ gene (Figure 16).
[0246] Example 20 - Single guide operation Using the method of Andronescu 2007, the predicted RNA folding of the active single RNA sequence was calculated at 37°. All hairpin loop secondary structures were singly removed from the construct and compiled iteratively into smaller single guides. Engineered single guides (esg) 4, 6, 7, 8, 9 were active for donor transposition (Figure 17C and Figure 17D), and engineered sgRNAs 8 and 9 were weaker single guides and transposed with PCR5 (Figure 17D). Engineered guide 5 was able to transpose, but engineered sgRNA 10 transposed weakly with PCR5 (Figure 17E and Figure 17F). Esg17 is a combination of deletions in esg6 and esg7, and esg18 is a combination of esg4 and esg5. Both were able to transpose strongly across both PCRs 4 and 5 (Figure 17G and 17H), but the combined addition of esg6 and esg18 to create esg19 resulted in weaker transposition in PCR 5, and the addition of esg7 to esg19 to create esg20 resulted in a very weak junction of transposition for PCR 5 (Figure 8G and 8H). In a second approach, the tracrRNA of MG64-1 was aligned to the documented Vk-type tracrRNA and the region of the unique insertion was mutated from the single guide. The sgRNA was minimized by truncation of the insertion sequence of the MG64-1 sgRNA (Figure 14). Two subsequent deletions, esg2 and esg3, were also tested (Figure 17A and 17B), but neither esg2 nor esg3 resulted in significant transposition, and thus the single guide was minimized by 57 bases.
[0247] Example 21 - LE-RE Minimization Sequencing of the target-transposition junction assisted in the identification of terminal inverted repeats by identifying the outermost sequence from the donor plasmid incorporated in the targeting reaction. By performing a 14 bp repeat analysis with 10% variability, short repeats contained within the termini were identified and truncations of these minimal termini were designed to retain the repeats while deleting the extra sequence. Multiple rounds of prediction and cloning were performed and each interaction was tested by in vitro transposition. Initial LE and RE deletions were designed alone and cloned at 68 bp, 86 bp, and 105 bp for the LE and 178 bp, 196 bp, and 242 bp for the RE. The RE of 64-1 also had a significant stretch of sequence that was free of repeats, so both 50 bp and 81 bp internal deletions were designed and cloned. Transpositions between all single deletions were robust for both PCR4 and PCR5 (Figures 18A and 18B), and then the 81 bp internal deletion was pursued along with the combined deletion of the RE. The previous 178, 196, and 212 bp trimmed ends were cloned onto the 81 bp internal deletion and transposition was tested. Transposition was active for all designed constructs. In combination with a 68 bp LE, transposition was demonstrated to be active up to a 68 bp LE region combined with a 96 bp RE region (Figures 18E and 18F).
[0248] Example 22 - Overhang effect of dislocation To test whether extra sequences outside the TnsB binding motif are required for transposition, oligos designed for the TGTACA motif in both the LE and RE were designed and synthesized with 0, 1, 2, 3, 5, and 10 bp of extra base pairs. These synthetic oligos were used to generate donor PCR fragments with overhangs and tested for their ability to transpose into the target site. Most notably, PCR6 was rarely detected from in vitro reactions (Figure 18G, lanes 1, 2), but efficient incorporation with PCR6 was detected with small 0-3 bp overhangs, reflecting the orientation of the RE proximal to the PAM that is not detected with larger flanking sequences.
[0249] Example 23 - CAST NLS Design Eukaryotic genome editing for therapeutic purposes relies primarily on the import of editing enzymes into the nucleus. Small polypeptide extensions of larger proteins signal cellular components for protein import across the nuclear membrane. The placement of these tags is not straightforward as these NLS tags must provide import function while maintaining the function of the protein to which it is fused. To test the functional orientation of the NLS to each of the components of the CAST complex, constructs were designed and synthesized that fuse a nucleoplasmic NLS to the N-terminus and an SV40 NLS to the C-terminus of each of the components of MG CAST. Proteins of these constructs were expressed in cell-free in vitro transcription / translation reactions and tested for in vitro transposition activity with a complementary set of untagged components. NLS-tagged constructs were assessed for maintenance of activity by PCR of the donor-target junction using PCR4 (to assess RE-distal transposition) and the cognate transposition event PCR5 (LE-to-proximal transposition).
[0250] Most constructs yielded a single NLS orientation that maintained activity. TnsB was a CAST construct that was active with both N- and C-terminal NLSs by both PCR4 and PCR5 (Figure 19A and Figure 19B). TniQ was active with an N-terminal NLS tag (Figure 19C and Figure 19D). And Cas12k constructs were active with a C-terminal tagged NLS (Figure 19E and Figure 19F, lanes 5, 6). Further development of Cas12k containing both nucleoplasmic and SV40 NLS tags was tested and found to be active (Figure 19I and Figure 19J, lane 4). TnsC was weakly active with an N-terminal NLS (Figure 19E and Figure 19F, lane 7), but further exploration of TnsC tagging identified new functional NLS-HA-TnsC and NLS-FLAG-TnsC constructs (Figure 19G and Figure 19H, lanes 3 and 7, respectively). The end result was a set of complete NLS-tagged constructs that were active in vitro in both NLS-TnsB and TnsB-NLS orientations (Figures 20A and 20B, lanes 5, 6).
[0251] Example 24 - Design and testing of Cas12k and TniQ protein fusion constructs In an attempt to simplify the expression of protein components and minimize the delivery of these components into cells, fusion constructs between Cas12k effectors and TniQ proteins were designed, synthesized, and tested. Both orientations of TniQ fused to Cas12k were designed and a C-terminal fusion, Cas-TniQ, and an N-terminal fusion, TniQ-Cas, were synthesized. Both constructs were weakly active with PCR4 (Figure 21A), but when expressed in vitro and assayed for transposition ability, the PCR5 junction was robustly formed by the TniQ-Cas fusion protein (Figure 21B). Transposition length was assayed with variable linker domains including original (20 amino acid linker), 48, 68, 72, and 77 (Figures 21C-F). NLS tags were then ligated to the N-terminus of TniQ and the C-terminus of Cas12k and found to still be active by PCR5 (Figures 20E and 20F).
[0252] Two other linkers were used to fuse the effector gene and the TniQ gene: the self-terminating translation sequence P2A was active in the Cas-NLS-P2A-NLS-TniQ construct (Figures 21G and 21H, lane 6), and the MCV internal ribosome entry sequence (IRES) mRNA-based linker allowed independent translation of the two components in cells (Figures 23F and 23G).
[0253] Example 25 - Intracellular expression-coupled in vitro translocation assay To test the functionality of the NLS constructs in a physiologically relevant environment, constructs cloned with active NLS-tagged CAST components were integrated into K562 cells using lentiviral transduction. Briefly, constructs cloned into lentiviral transfer plasmids were transfected into 293T cells with envelope and packaging plasmids, and virus containing supernatants were harvested from the medium after 72 hours of incubation. The medium containing the virus was then incubated with K562 cell lines containing 8 μg / mL polybrene for 72 hours, and transfected cells were then selected for integration en masse using puromycin at 1 μg / mL for 4 days. The selected cell lines were harvested at the end of the 4 days and differentially lysed for nuclear and cytoplasmic fractions. Subsequent fractions were then tested for transposition competence using a complementary set of in vitro expression components.
[0254] Both NLS-TnsB and TnsB-NLS were tested by cell fractionation and in vitro translocation, and translocation was detected across both cytoplasmic and nuclear fractions, with NLS-TniQ having detectable activity in the cytoplasm (Figures 22A and 22B). NLS-HA-TnsC and NLS-FLAG-TnsC were active in both cytoplasmic and nuclear fractions when expressed (Figure 22D), whereas PCR4 was formed in the nuclear fraction for both TnsC constructs (Figure 22C).
[0255] When both NLS-TnsB or TnsB-NLS were linked to NLS-FLAG-TnsC by using IRES, NLS-TnsB-IRES-NLS-FLAG-TnsC was mostly active in the nuclear fraction, while TnsB-NLS-IRES-NLS-FLAG-TnsC was active in both the cytoplasmic and nuclear fractions, indicating that NLS-TnsB has a higher ability to transport to the nucleus (Figure 21E, Figure 21F).
[0256] Cas12k fusions in cells were similarly fractionated and tested for translocation. Cas-NLS Cas-NLS-P2A-NLS-TniQ was transduced into cells, fractionated, and tested in vitro for intracellular activity. Cas-NLS-P2A-NLS-TniQ could be translocated into the cytoplasm by adding a single guide to the reaction (Figure 23A). The Cas-NLS-P2A-NLS-TniQ construct in the nuclear fraction was complemented by supplementing the holo-Cas protein (+sgRNA) or additional TniQ with sgRNA. This indicates that both Cas-NLS and NLS-TniQ are making their way to the nucleus (Figures 23B and 23C). The NLS-TniQ-Cas-NLS fusion protein had similar results but required additional recruitment of TniQ (Figures 23D and 23E), and Cas-NLS-IRES-NLS-TniQ required recruitment from holo-Cas-NLS only (Figures 23F and 23G). Overall, this indicates that all components of CAST were able to be delivered to the nuclear fraction of cells.
[0257] Example 26 - Transposon end verification via gel shift To verify the activity of TnsB on the predicted transposon end sequences, FAM-labeled oligos were used to amplify the LE of MG64-1. Using a cell-free transcription / translation system, MG64-1 TnsB protein was expressed and incubated with the LE FAM-labeled product. After 30 min of incubation, binding was observed on a native 5% TBE gel (Figure 24). Multiple bands of fluorescent product in the co-incubated lanes (Figure 24, lane 3) indicated a minimum of two TnsB binding sites.
[0258] The disclosed system can be used for a variety of applications, such as, for example, nucleic acid editing (e.g., gene editing) or binding (e.g., sequence-specific binding) to nucleic acid molecules. Such systems can be used, for example, to correct (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences encoding disease-causing mutations), as inactivated enzymes combined with probes to target and detect specific nucleotide sequences (e.g., sequences encoding antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to render them unable to infect host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and / or to detect cellular perturbations by exogenous small molecules and nucleotides as biosensors.
[0259] Example 27 - Defined Domains Functional domains (FDs) are small proteins that can facilitate protein interaction with DNA, such as DNA binding domains (DBDs) and chromatin modulating domains (CMDs). By using binding domains that are non-specific to DNA sequences, the affinity of the functional protein was increased without adversely affecting function. Four functional domains were selected for their ability to non-specifically bind DNA and DNA-associated proteins: human histone 1 central globular domain (aa.22-101) (SEQ ID NO:112), HMGN1 (1-100) (SEQ ID NO:111), human Cbx5 (1006-1274) (SEQ ID NO:110), and Saccharolobus solfataricus sso7d (1-64) (SEQ ID NO:109).
[0260] Example 28 - Cloning Using the functional domains of Example 27, (a) CAST-derived Cas-FD fusions and (b) CAST-derived TniQ-FD fusions were constructed to determine whether the functional domains increased the activity of these CAST system components in cells. The DNA binding domains were codon optimized for human expression and were synthesized or assembled using PCR stitching of oligos. To construct Cas12k and TniQ DBD fusion proteins, the DBD proteins were amplified using primers and assembled with Cas12k-NLS and NLS-TniQ. The DNA sequences of the cloned fusion genes were confirmed by Sanger sequencing.
[0261] Example 29 - In vitro testing of fusions to functional domains CAST activity will be tested using five types of components: (1) Cas-NLS effector or Cas12k-FD-NLS protein expressed by an in vitro expression system, (2) target DNA fragment or plasmid containing a target sequence and PAM corresponding to the Cas enzyme, (3) donor DNA fragment containing a marker or fragment of DNA flanking the LE and RE of the transposase system in the DNA fragment or plasmid, (4) any combination of transposase protein, transposase NLS protein, or transposase-FD-NLS protein expressed using an in vitro expression system, and (5) engineered in vitro transcribed single guide RNA sequence. Active systems resulting in successful transposition of the donor fragment will be assayed by PCR amplification of the donor-target junction.
[0262] CAST NLS or CAST-FD-NLS fusion proteins were expressed in vitro and tested for functionality in transposition reactions by exchanging non-FD fusion components with the fusion proteins (e.g., exchanging Cas12k-NLS with Cas12k-FD-NLS). When tested individually, sso7d, HMGN1, and Cbx5 fusions with Cas12k were active for transposition (Figures 25A and 25B, lane 6; Figure 25C, lane 7; and Figure 25D, lane 6). TniQ fusions were active for sso7d, HMGN1, H1core, and Cbx5 fusions (Figure 25C, lanes 3-5, and Figure 25D, lane 4).
[0263] Example 30 - Nuclear function of Cas-sso7d The lentiviral cargo vector containing the active Cas12k-DBD-NLS fusion protein was transfected into 293w cells containing the envelope and packaging plasmids. After 72 hours at 37°C, the supernatant containing the active lentiviral particles was incubated with K562 cells for viral transduction. The cells were selected for lentiviral integration by selection on 2μg / mL puromycin for 4 days at 37°C. After selection, the cell nuclei were extracted and tested for nuclear activity of the Cas12k-DBD-NLS fusion protein.
[0264] Cas12k-Cbx5 was poorly active and could translocate into the nucleus, but not efficiently enough to have active translocation in the extracted nuclear reaction. However, when additional in vitro expressed TniQ was added to the nuclear fraction, translocation was detectable by PCR. This implies the presence of active Cas12k protein in the nucleus, but stoichiometrically requires additional TniQ to have active translocation.
[0265] To test Cas12k-sso7d, lentiviral transduced and puromycin selected cells were extracted for nuclear fractionation. When nuclear extracts were used to test the transposition junction, PCR of LE to the target junction had a faint transposition band in the nuclear fraction without the addition of in vitro Cas12k (Figure 26A lane 6). The faint signal was Sanger sequenced and aligned with the predicted junction PCR to verify transposition (Figure 26B). From the sequence signal, it was concluded that Cas12k-sso7d was capable of transposition, albeit at a low efficiency.
[0266] Example 31 - TniQ with DBD in nuclear extracts To establish a functional version of TniQ fused to the DBD, sso7d, Cbx5, HMGN1, and H1core fusion constructs were tested with TniQ. These constructs, along with WT TniQ, were expressed in vitro and supplemented with nuclear extracts containing all four components of the CAST system. WT TniQ had a moderate ability to transpose the donor to fragments, but the use of sso7d, cbx5, HMGN1, and H1core improved the transposition capacity of the nuclear extracts, indicating the ability of all four constructs to improve transposition function in vitro (Figure 26C).
[0267] Example 32 - Nuclear extract and full suite expression of Cas-sso7d and TniQ co-expression Of the four TniQ-DBD fusions tested in nuclear extracts that were more active than WT TniQ, co-expression constructs were prepared that co-expressed the TniQ DBD format with Cas12k and Cas12k-sso7d in a combinatorial fashion using an IRES. Upon transduction and selection on puromycin, cells expressing both versions of Cas12k and TniQ were extracted for nuclear fractions. Exogenous TnsB and TnsC were added to the in vitro reaction. Of the eight constructs tested, only two versions of the Cas12k and TniQ fusion proteins were active, with the cells co-expressing Cas12k-sso7d and HMGN1-TniQ, and Cas12k-sso7d and H1core-TniQ. These nuclear fraction translocations were detectable, and the sequences were verified using Sanger sequencing of potential PCR junctions (Figures 27A and 27B).
[0268] To test the activity of all components in cells, co-expression constructs of TnsB and TnsC were transduced into cells known to have high nuclear fractionation activity. Both Cas12k-sso7d and HMGN1-TniQ, and Cas12k-sso7d and H1core-TniQ populations were transduced with TnsB and TnsC lentiviral constructs. Nuclear extracts of these cell populations containing all four components confirm the function of all protein CAST components in the nuclear extracts (Figure 27C and Figure 27D).
[0269] Example 33 - Exploration of target DNA dilution in in vitro transposition assays In vitro targeted integrase activity experiments Integrase activity was assayed with a target plasmid (pTarget) containing a PAM flanked by protospacer sequences (Figure 28A). A T7 promoter-leader sequence was introduced by PCR amplification of all transposase, single-stranded RNA (sgRNA), and effector components and expressed independently in an in vitro transcription / translation system (Figure 28A). Purified in vitro transcribed single guide RNAs were refolded in duplex buffer (10 mM Tris, 150 mM NaCl, 1 mM MgCl2 at pH 7.0) and normalized to 1 μM. Donor fragments were PCR amplified from plasmid pDonor containing kanamycin or tetracycline resistance markers flanked by MG64-1 left end (LE) and right end (RE) transposon motifs and normalized to 50 ng / μL.
[0270] After expression, 1 μL of the Cas12k in vitro expression reaction was added to 0.5 pmol of sgRNA and incubated at 25° C. for 20 minutes. Then, the individually expressed transposase protein was added volumetrically at 1 μL per expression. The target DNA and 50 ng of donor DNA were then added to the transposition reaction in a reaction buffer with a final concentration of 26 mM HEPES at pH 7.5, 4.2 mM Tris at pH 8, 50 μg / mL BSA, 2 mM ATP, 2.1 mM TCEP, 0.05 mM EDTA, 0.2 mM MgCl2, 28 mM NaCl, 21 mM KCl, 1.35% glycerol (final pH 7.5), and 15 mM Mg(OAc)2. The in vitro transposition reaction was carried out at 37° C. for 2 h, and the transposition reaction was diluted 10-fold in water and then used as a template for junction PCR analysis.
[0271] Junction PCR analysis Junction PCR reactions were performed with Q5 polymerase and amplified with flanking primers: Rxn#1 (target), Rxn#2 (donor), Rxn#3 (reverse LE), Rxn#4 (forward RE), Rxn#5 (forward LE), and Rxn#6 (reverse RE) (Figure 28B). PCR fragments were run on a 2% agarose gel in 1xTAE and analyzed for size discrimination. Bands of appropriate size for each PCR junction were gel excised and PCR fragments were recovered through purification and Sanger sequenced using both amplification primers. The resulting Sanger sequences were mapped to the donor and target sequences, confirming integration approximately 60 bp away from the PAM.
[0272] Transposition reaction with dilution of target plasmid DNA Due to the known low availability of single copy target sites in the human genome (e.g., Moreb et al., 2020), it was predicted that low availability targets in in vitro experiments could mimic the search space in human genomic DNA by MG64-1 CAST. To determine the minimum amount of target plasmid (pTarget) required to detect targeting by MG64-1 in vitro, CAST was used in serial dilutions of total target DNA in transposition reactions. The target plasmid amount was serially diluted 10-fold from 50 ng DNA (mol) to 0.00005 ng (mol) per reaction and then added to a transposition reaction containing MG64-1 suite, sgRNA, and donor plasmid (Figure 28A). PCR amplification of the target-donor junction across the transposition product was then analyzed by gel electrophoresis (Figure 28B).
[0273] Results: Serial dilutions of target plasmid DNA In vitro transposition experiments with dilution of the target plasmid showed that a single guide RNA (sgRNA) was required for targeted integration of the donor onto the plasmid at 50 ng of target DNA (Figure 28C, lanes 1 and 2). As the target was serially diluted, the target DNA was still detectable at 0.00005 ng / reaction, albeit at lower intensity (Figure 28C, "target" band on lane 8). Both the reverse LE integration product (Rxn#3) and the reverse RE product (Rxn#6) decreased with 10-fold dilution, resulting in the detectable presence of Rxn#3 down to 0.5 ng of target DNA (Figure 28C). Both transposition reactions Rxn#4 and #5 were equally robust and detectable down to 0.05 ng of target DNA.
[0274] Based on the presence of three of the four expected transposition reaction products, the 0.5 ng target plasmid condition was chosen to test whether transposition would be detected when increasing the complexity of the DNA search space. Increasing amounts of exogenous human genomic DNA (gDNA) were added to reactions with 0.5 ng of immobilized target plasmid, MG64-1 CAST, and sgRNA (Figure 28A and Figure 28D). When no gDNA was added, the transposition experiment confirmed that sgRNA was necessary for targeted integration of the donor onto the target plasmid at 0.5 ng of DNA (Figure 28D, lanes 1 and 2). When gDNA was added to the reactions in increasing amounts (0 to 2000 ng of DNA per reaction), the transposition products were also diluted, as given by a faint band compared to the no gDNA control (Figure 28D). For example, the reverse LE integration product (Rxn#3) was visible when only 50 ng of gDNA was added at most. In addition, robust transposition products were detected for forward RE (Rxn#4) and reverse LE (Rxn#5) products at 125-1000 ng of gDNA with a fixed amount of 0.5 ng of target plasmid (Figure 28D, lanes 4-8).
[0275] Example 34 - In vitro transposition into human gDNA using MG64-1 CAST In vitro targeting of high copy elements across the human genome This example evaluates a dilution series of natural targets in the genome for transposition as a function of target frequency.
[0276] Results: Target sites identified within high copy regions of the human genome High copy targets were identified in the human genome using the Cas off-target finder. Using 200-300 bp target sequences within the most conserved space of each duplicated element, 15 target sites for each line 1 3' and HERV were identified, and 7 target sites were designed for SVA elements with various GC content, orientation, and substitutions of the MG64-1 rGTN PAM (Figure 29A). Target (spacer) sequences were synthesized as oligos and PCR amplified onto the MG64-1 sgRNA template with a T7 promoter upstream of the single guide backbone. The PCR reaction of the MG64-1 sgRNA was then purified and in vitro transcribed. The NLS-tagged MG64-1 protein components were used to assemble in vitro transposition reactions as described above, using purified HEK293T gDNA at 1 μg / reaction as the target DNA.
[0277] Results: In vitro targeted transposition into high copy regions of the human genome. Successful transposition by MG64-1 was assessed by the resulting Fwd PCR and Rev PCR junction products at each of the 15 target sites in the high copy element (Figure 29B). MG64-1 promoted transposition into line 1 targets 3, 5, 6, 7, 10, and 13 in the forward orientation, while targets 1, 2, 8, 9, 12, 14, and 15 were reactive for transposition in both forward and reverse orientations (Figure 29B). In addition, active transposition into SVA target 3 and HERV target 5 was specific for both forward and reverse orientations (Figures 29C-D). Sanger sequencing of line 1, SVA, and HERV transposition reaction products confirmed that in vitro transposition was specific and RNA-guided (Figures 29E-H).
[0278] Example 35 - NLS functional domain fusions with MG64-1 CAST can be targeted to high copy elements Functional domains fused to CAST components for in vitro targeted transposition to high-copy elements in the human genome Using experiments showing that high copy elements are efficient targets for integration in vitro, functional domains were tested for their effect on CAST targeting at these sites. Functional domain fusions of Cas12k-sso7d were challenged with H1core-TniQ, or Cas12k-sso7d with HMGN1-TniQ, transposing donor DNA when reacting with NLS-TnsB and NLS-TnsC fusions at high copy elements line 1 target 12 and target 15. Transduction reactions were assembled using either Cas12k alone or fused Cas12k-sso7d where indicated, with no sgRNA (-sg) conditions as a negative control for transposition, and with sgRNAs for target 12 and target 15 of the line 1 3' element where indicated (Figure 30). In addition, the transposition reaction was supplemented with translated NLS-TniQ, NLS-H1core-TniQ or NLS-HMGN1-TniQ, NLS-TnsB, NLS-TnsC, pDonor, buffer, and human gDNA for targeting.
[0279] result In vitro transposition assays using fusion domains targeting line 1 target 12 and target 15 showed successful integration with no MG64-1 fusion, the positive control (Figure 30, lanes 3 and 7), and with the fusion Cas12k-sso7d with HMGN1-TniQ (Figure 30, lanes 5 and 9). No targeting occurred when sgRNA was not added to the reaction (Figure 30, lanes 2 and 6).
[0280] Example 36 - NLS fusion to S15 for targeted transposition An NLS fusion to the small ribosomal protein subunit S15 is required for correct orientation of the tag Recently, the small prokaryotic ribosomal protein subunit S15 was deemed necessary for targeted transposition by Cas12k CAST in vitro (Schmidt et al., 2022; Park et al., 2022). Therefore, we evaluated the necessity of S15 with or without an NLS tag in transposition experiments using MG64-1. Because the in vitro expression reagent already contains S15 (S15 is part of the prokaryotic ribosomal complex required for protein expression), in vitro transposition experiments using CAST protein expressed in an in vitro expression system likely carried over S15 from the expression step and was subsequently recruited for CAST targeting.
[0281] Results: S15 increases transposition efficiency in vitro Wheat germ extract was used in an S15-free eukaryotic transcription / translation system to express the MG64-1 CAST components. The CAST templates were amplified to contain a T7 promoter and a 40 bp polyA tail for transcriptional stability of the mRNA template. Proteins were expressed from dsDNA templates via transcription / translation reactions, which were then used in in vitro transposition reactions as described above. The results show that S15 addition increased the targeted transposition efficiency, as indicated by the intensity of the band from the junction PCR product (Rxn#5) (Figure 31A, lanes 4-5).
[0282] Results: The S15-NLS fusion is the preferred orientation for in vitro transposition In eukaryotic conditions, protein translation is carried out exclusively in the cytoplasm, while the translocation reaction mediated by CAST most likely occurs in the nucleus. The necessity of the NLS tag for S15 nuclear localization was evaluated. NLS tags were fused to both the N-terminus and C-terminus of S15 and tested in eukaryotic in vitro transcription / translation reactions and in vitro translocation experiments (Figure 31A, lane 5, and Figure 31B, lanes 4 and 5). The results show that S15-NLS was more efficient for translocation than the other tested conditions (Figure 31A, lane 5).
[0283] Example 37 - S15 is required for cellular translation of CAST Design of CAST vectors MG64-1 CAST protein was expressed on two high expression plasmids for transposition experiments in human cells. One plasmid expresses the protein targeting complex under the control of the pCAG promoter. Two versions of the protein targeting complex were designed. One version contains a Cas12k-sso7d functional domain fusion with S15-NLS, IRES, and a 2A polypeptide fused to NLS-H1core-TniQ (Figure 32A, top left). The second version contains Cas12k-sso7d-2A-S15-NLS with an NLS-HMGN1-TniQ fusion (Figure 32A, bottom left). The targeting plasmid also contained a pU6 PolIII promoter driving transcription of humanized MG64-1 sgRNA for targeting one of line 1 targets 8, 12, and 15, as well as SVA target 3. The second plasmid transfected into the cells was a donor plasmid containing NLS-TnsB and NLS-TnsC separated by an IRES under the expression of the pCAG promoter. On this plasmid, a 2.5 kb DNA cargo was contained between the LE and RE terminal inverted repeats (Figure 32A, right).
[0284] HEK293T lipofection 2.5 million HEK293T cells were seeded 24 hours prior to lipid-based transfection of the two-plasmid system at 9 μg:9 μg of targeting:donor plasmid. Cells were incubated at 37°C for 72 hours and then harvested by resuspension in 4 mL of 1x PBS at pH 7.2. 2 mL of resuspended cells were harvested for gDNA and eluted in 200 μL of elution buffer. 5 μL of extracted gDNA was assayed for transposition in a 100 μL Q5 PCR reaction with primers specific for the high copy element target. Forward transposition was determined by amplification with primers specific for Fwd PCR and reverse transposition was determined by amplification with primers specific for Rev PCR. Amplified PCR reactions were visualized on a 2% agarose gel. Transposition was predicted to transpose 60 nt away from the PAM as observed in the in vitro transposition experiments and was determined to be active by the presence of a single band for junction PCR amplification at the predicted size. PCR amplicons were Sanger sequenced and NGS sequenced for rearrangement profile analysis.
[0285] result Cells transfected with both versions of the targeting complex plasmid with H1core-TniQ or HMGN1-TniQ were analyzed for transposition (Figure 32B). Both versions of the targeting complex plasmid promoted transposition at all four target sites (Figure 32B, arrows). Line 1 targets 8 and 15 were only detectable in the LE to 5' target orientation, while line 1 target 12 was only detectable in the LE to 3' orientation (Figure 32B). The results indicate that targeted integration in human cells has a strong preference for directionality from the PAM.
[0286] Sanger sequencing of the PCR junctions confirmed integration at line 1 targets 8, 12, and 15, 59, 62, and 60 nt away from the PAM, respectively (Figures 32C-H). Sequencing signal degradation was observed at the translocation junctions, resulting from a mixture of population events. PCR amplicons were sequenced via NGS to determine the single molecule profile of each integration event. Reads resulting from NGS sequencing confirmed targeted integration at line 1 targets 8, 12, and 15, as well as SVA target 3. Variation in the targeted regions indic...
Claims
1. 1. A system for translocating a cargo nucleotide sequence into a target nucleic acid site in a target nucleic acid, comprising: a) a Cas effector complex comprising a class 2 type V Cas effector, a small prokaryotic ribosomal protein subunit S15, and an engineered guide polynucleotide configured to hybridize to the target nucleic acid site; b) a Tn7-type transposase complex configured to bind to the Cas effector complex, the Tn7-type transposase complex comprising a TnsB component, a TnsC component, and a TniQ component; c) a double-stranded nucleic acid configured to interact with the Tn7-type transposase complex and comprising the cargo nucleotide sequence; d) a functional domain comprising a DNA binding domain (DBD) or a chromatin modulation domain (CMD).
2. 2. The system of claim 1, wherein the Cas effector complex is (a) non-covalently bound to the Tn7-type transposase complex, (b) covalently bound to the Tn7-type transposase complex, or (c) fused to the Tn7-type transposase complex.
3. The system described in claim 1, wherein the cargo nucleotide sequence is adjacent to a left transposase recognition sequence and a right transposase recognition sequence recognized by the Tn7-type transposase complex.
4. The system described in claim 3, wherein the left recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 9, 11, 36-38, 76, and 78, and the right recombinase sequence comprises a sequence having at least 80% identity to any one of SEQ ID NOs: 8, 10, 39-44, 77, 79, and 93.
5. The system described in claim 1, wherein the target nucleic acid comprises a PAM sequence compatible with the Cas effector complex, and the PAM sequence comprises sequence number 31.
6. The system described in claim 5, wherein the PAM sequence is located approximately 50 to approximately 70 base pairs from the target nucleic acid site.
7. The system described in claim 1, wherein the class 2 V-type Cas effector is a Cas12k effector comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 1, 12, 16, 20-30, 64, 80-85, and 200.
8. The system described in claim 1, wherein the TnsB component comprises a polypeptide having an sequence having at least 80% identity to any one of SEQ ID NOs: 2, 13, 17, and 65.
9. The system described in claim 1, wherein the Tn7 transposase complex comprises at least a first polypeptide and a second polypeptide, each independently comprising a sequence having at least 80% identity to any one of SEQ ID NOs: 3-4, 14-15, 18-19, and 66-67.
10. The system described in claim 1, wherein the engineered guide polynucleotide comprises a sequence comprising at least about 46 to 80 consecutive nucleotides having at least 80% identity to any one of SEQ ID NOs: 5-6, 32-33, 94-95, 104-105, and 202.
11. The system described in claim 1, wherein the engineered guide polynucleotide comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 106, 107, 108, 5, 45-63, 68-75, 96-103, and 165.
12. The system described in claim 1, wherein the functional domain is derived from human histone 1 central globular domain, HMGN1, cbx5, or Saccharolobus solfataricus sso7d.
13. The system described in claim 1, wherein the class 2 V-type Cas effector is fused to the functional domain to form a fusion protein having at least 80% identity to any one of SEQ ID NOs: 113 to 116.
14. The system of claim 1, wherein the Tn7 transposase complex comprises a TniQ protein fused to the functional domain to form a fusion protein.
15. The system described in claim 14, wherein the TniQ protein comprises a sequence having at least 80% sequence identity with any one of the TniQ domains of SEQ ID NOs: 117 to 120.
16. The system described in claim 1, wherein the small prokaryotic ribosomal protein subunit S15 comprises a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 167-169.
17. The system described in claim 1, wherein the small prokaryotic ribosomal protein subunit S15 is encoded by a sequence having at least 80% sequence identity with any one of SEQ ID NOs: 161 to 163.
18. The system described in claim 1, wherein the class 2 V-type Cas effector and the Tn7-type transposase complex are encoded by a polynucleotide sequence comprising less than about 10 kilobases.
19. A method for translocating a cargo nucleotide sequence within a target nucleic acid site, comprising introducing into a cell a system described in any one of claims 1 to 18.
20. A cell comprising a system described in any one of claims 1 to 18.
21. The cell described in claim 20, wherein the cell is a eukaryotic cell, a mammalian cell, an immortalized cell, an insect cell, a yeast cell, a plant cell, a fungal cell, or a prokaryotic cell.
22. The cell of claim 20, wherein the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.