RNA-guided genome recombineering at the kilobase scale
Patent Information
- Application Number
- JP2024547600
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-01
- Filing Date
- 2023-02-10
- Publication Date
- 2026-02-10
AI Technical Summary
【0008】 概要 高精度でありオフターゲットエラーが少ない、大規模な核酸編集を可能にする方式で核酸編集を促進する系および方法が本明細書中で提供される。これらの系および方法は、組換えタンパク質成分および必要に応じてCRISPR成分を用いる。
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Patent Application Nos. 63 / 308,830, 63 / 308,834, and 63 / 308,837, filed February 10, 2022, respectively, and PCT / US2022 / 075850, filed September 1, 2022, each of which is incorporated by reference in its entirety.
[0002] See U.S. Patent Application No. 62 / 984,618, filed March 3, 2020, U.S. Patent Application No. 63 / 146,447, filed February 5, 2021, and PCT / US2021 / 020513, filed March 2, 2021.
[0003] The aforementioned applications, and all documents cited therein or during prosecution thereof ("Application Citations"), and all documents cited or referenced in the Application Citations, and all documents cited or referenced herein ("Prescription Citations"), and all documents cited or referenced in the Prescription Citations, together with manufacturer's directions, descriptions, product specifications, and product sheets for any products mentioned herein or any documents incorporated by reference herein, are hereby incorporated by reference and may be used in the practice of this invention. More specifically, all references are incorporated by reference to the same extent as if each individual document was specifically and individually indicated to be incorporated by reference.
[0004] FIELD OF THE INVENTION The present invention relates to an RNA-guided recombineering editing system using phage recombinase, as well as methods, vectors, nucleic acid compositions, and kits therefor.
[0005] Sequence Listing The contents of the Electronic Sequence Listing entitled STDU2-41722_601_SQL.xml (Size: 758,685 bytes; Created February 10, 2023) are incorporated herein by reference in their entirety. [Background technology]
[0006] background The clustered regularly interspaced short palindromic repeats (CRISPR) system, originally found in bacteria and archaea as part of the immune system to defend against invading viruses, forms the basis for genome editing technologies that can be programmed to target specific stretches of genome or other DNA for precise editing. While various CRISPR-based tools are available, most are adapted to editing short sequences. There is a strong demand for editing longer sequences in the engineering of model systems, the generation of therapeutic cells, and gene therapy. Previous studies have developed techniques to improve Cas9-mediated homology-5 directed repair (HDR) (K.S. Pawelczak, et al., ACS Chem. Biol. 13, 389-396 (2018)). Tools utilizing nucleic acid-modifying enzymes in conjunction with Cas9, such as prime editing (A.V. Anzalone, et al., Nature. 576, 149-157 (2019)), have demonstrated edits up to 80 base pairs (bp) in length. Despite these advances, there remains a continuing need for large-scale mammalian genome engineering with high efficiency and fidelity. Citation or identification of any document in this application is not an admission that such document is available as prior art to the present invention. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] KS Pawelczak, et al., ACS Chem. Biol. 13, 389-396 (2018) [Non-patent document 2] AV Anzalone, et al., Nature. 576, 149-157 (2019) Summary of the Invention [Means for solving the problem]
[0008] overview Provided herein are systems and methods that facilitate nucleic acid editing in a manner that allows for large-scale nucleic acid editing with high precision and reduced off-target errors.These systems and methods use recombinant protein components and, optionally, CRISPR components.
[0009] For example, disclosed herein is a system comprising a binding protein, a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence, and a recombination protein. The recombination protein can be a microbial recombination protein, such as a single-stranded DNA annealing protein (SSAP), including, but not limited to, RecE, RecT, lambda exonuclease (Exo), Bet protein (betA, redB), exonuclease gp6, single-stranded DNA binding protein gp2.5, or a derivative or variant thereof. In some embodiments, the system further comprises donor DNA. In some embodiments, the target DNA sequence is a genomic DNA sequence in a host cell. In certain embodiments, CRISPR components are absent. In certain embodiments, the system comprises a mobilization system that mobilizes the recombination protein and a nucleic acid that directs the recombination protein to a target. In certain embodiments, the mobilization system mobilizes the recombination protein, the nucleic acid that directs the recombination protein, and the CRISPR components.
[0010] In one aspect, the present invention provides a recombinant system that includes SSAP and lacks CRISPR components. In certain embodiments, the present invention provides a system or composition that includes: (i) a nucleic acid molecule that includes a guide RNA sequence that is complementary to a target DNA sequence; and (ii) a recombinant protein, wherein the recombinant protein includes an exonuclease, a single-stranded DNA annealing protein (SSAP), or a single-stranded DNA binding protein (SSB), or a combination of two or more thereof; or (iii) a nucleic acid molecule that encodes or delivers (i) and / or (ii) for in vivo expression in cells; or (iv) a vector that contains the nucleic acid molecule described in (iii) for in vivo expression in cells. In certain embodiments, the system or composition does not include a CRISPR protein, or does not include a Cas protein, or does not include a Cas9 protein, or does not include a Cas12a protein.
[0011] In certain embodiments, the system or composition comprises a recruitment system for recruiting a guide nucleic acid and a recombinant protein. In certain embodiments, the recruitment system comprises at least one aptamer sequence and an aptamer-binding protein operably linked to the recombinant protein as part of a fusion protein. In certain embodiments, the at least one aptamer sequence is an RNA aptamer sequence or a peptide aptamer sequence. In certain embodiments, the nucleic acid molecule or multiple nucleic acid molecules additionally comprise at least one RNA aptamer sequence, or comprise one, two, three, or more RNA aptamer sequences. In certain embodiments, the two aptamer sequences comprise the same sequence or comprise sequences that bind to the same aptamer-binding protein.
[0012] In certain embodiments, the aptamer-binding protein comprises an MS2 coat protein, or a functional derivative or variant thereof. In certain embodiments, the aptamer-binding protein comprises a phage N-peptide, or a functional derivative or variant thereof. In certain embodiments, at least one peptide aptamer sequence is conjugated to a guide RNA. In certain embodiments, the at least one peptide aptamer sequence comprises 1 to 24 peptide aptamer sequences. In certain embodiments, two or more aptamer sequences comprise the same sequence. In certain embodiments, the aptamer sequence comprises a GCN4 peptide sequence.
[0013] In certain embodiments, the N-terminus of the recombinant protein is linked to the C-terminus of the aptamer-binding protein. In certain embodiments, the recombinant protein and the aptamer-binding protein are operably linked by a linker.
[0014] In certain embodiments, the recombinant system or composition comprises at least one nuclear localization sequence (NLS), and optionally, the NLS(s) are linked to the recombinant protein. In certain embodiments, the NLS is located at the C-terminus of the recombinant protein or at the N-terminus of the recombinant protein.
[0015] In certain embodiments, the recombinant protein comprises a microbial recombinant protein or active portion thereof, a mitochondrial recombinant protein or active portion thereof, a viral recombinant protein or active portion thereof, or a eukaryotic recombinant protein or active portion thereof, including, but not limited to, a recombinant protein set forth in Table 12, or a derivative or variant or functional portion thereof. In certain embodiments, the recombinant protein comprises an amino acid sequence having at least 70% identity, or at least 75% identity, or at least 80% identity, or at least 85% identity, or at least 90% identity, or at least 92% identity, or at least 95% identity, or at least 96% identity, or at least 97% identity, or at least 98% identity, or at least 99% identity to a recombinant protein set forth in Table 12, or a derivative or variant or functional portion thereof.
[0016] In certain embodiments, the system or composition comprises a donor nucleic acid. In certain embodiments, the donor nucleic acid comprises a homology arm.
[0017] In certain embodiments, the recombination system is comprised in a cell, eg, a eukaryotic cell, a mammalian cell, an animal cell, a human cell, or a plant cell.
[0018] The mobilization system is adaptable to numerous combinations and configurations of recombination proteins. For example, by selecting and incorporating multiple nucleic acid aptamers, the system can include multiple recombination proteins, which may be the same or different, in various ratios. In certain embodiments, the system includes an exonuclease. In certain embodiments, the system includes an SSAP. In certain embodiments, the system includes an SSB. In certain embodiments, the system includes an exonuclease and an SSAP. In certain embodiments, the system includes an exonuclease and an SSB. In certain embodiments, the system includes an SSAP and an SSB. In certain embodiments, the system includes an exonuclease and an SSAP, but not an SSB. In certain embodiments, the system includes an exonuclease and an SSB, but not an SSAP. In certain embodiments, the system includes an SSAP and an SSB, but not an exonuclease. In certain embodiments, the system includes an exonuclease, an SSAP, and an SSB.
[0019] In one aspect, the present invention provides a recombination system comprising an SSAP and a reverse transcriptase (RT). In certain embodiments, the present invention provides a system or composition comprising: (i) a reverse transcriptase(s) (RT); (ii) a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence and an RNA for reverse transcription, or a plurality of nucleic acid molecules comprising a guide RNA sequence complementary to a target DNA sequence and an RNA for reverse transcription; (iii) a recombination protein comprising an exonuclease, a single-stranded DNA annealing protein (SSAP), or a single-stranded DNA binding protein (SSB), or a combination of two or more thereof; or (iv) a nucleic acid molecule(s) encoding or delivering (i) and / or (ii) and / or (iii) for in vivo expression in a cell; or (v) a vector containing the nucleic acid molecule(s) of (iv) for in vivo expression in a cell.
[0020] In certain embodiments, the system or composition further comprises a Cas protein, or (iv) comprises nucleic acid molecule(s) encoding or delivering (i) and / or (ii) and / or (iii) and / or the Cas protein for in vivo expression in a cell, or the vector(s) of (v) additionally contain nucleic acid molecule(s) encoding the Cas protein.
[0021] In certain embodiments, one or more of the components are provided as a complex. For example, the protein or fusion protein and nucleic acid are provided as a ribonucleoprotein (RNP). Non-limiting examples of RNPs include CRISPR guide RNA complexes and SSAP guide RNA complexes. In certain embodiments, a fusion protein comprises one or more components. Non-limiting examples include Cas9-SSAP fusions, Cas9-RT fusions, and SSAP-RT fusions.
[0022] In certain embodiments, the system or composition comprises a recruitment system for recruiting a guide nucleic acid and a recombinant protein. In certain embodiments, the recruitment system comprises at least one aptamer sequence and an aptamer-binding protein operably linked to the recombinant protein as part of a fusion protein. In certain embodiments, the at least one aptamer sequence is an RNA aptamer sequence or a peptide aptamer sequence. In certain embodiments, the nucleic acid molecule or multiple nucleic acid molecules additionally comprise at least one RNA aptamer sequence, or comprise one, two, three, or more RNA aptamer sequences. In certain embodiments, the two aptamer sequences comprise the same sequence or comprise sequences that bind to the same aptamer-binding protein.
[0023] In certain embodiments, the aptamer-binding protein comprises an MS2 coat protein, or a functional derivative or variant thereof. In certain embodiments, the aptamer-binding protein comprises a phage N-peptide, or a functional derivative or variant thereof. In certain embodiments, at least one peptide aptamer sequence is conjugated to a guide RNA. In certain embodiments, the at least one peptide aptamer sequence comprises 1 to 24 peptide aptamer sequences. In certain embodiments, two or more aptamer sequences comprise the same sequence. In certain embodiments, the aptamer sequence comprises a GCN4 peptide sequence.
[0024] In certain embodiments, the N-terminus of the recombinant protein is linked to the C-terminus of the aptamer-binding protein. In certain embodiments, the recombinant protein and the aptamer-binding protein are operably linked by a linker.
[0025] In certain embodiments, the recombinant system or composition comprises at least one nuclear localization sequence (NLS), and optionally, the NLS(s) are linked to the recombinant protein. In certain embodiments, the NLS is located at the C-terminus of the recombinant protein or at the N-terminus of the recombinant protein.
[0026] In certain embodiments, the recombinant protein comprises a microbial recombinant protein or active portion thereof, a mitochondrial recombinant protein or active portion thereof, a viral recombinant protein or active portion thereof, or a eukaryotic recombinant protein or active portion thereof, including, but not limited to, a recombinant protein set forth in Table 12, or a derivative or variant or functional portion thereof. In certain embodiments, the recombinant protein comprises an amino acid sequence having at least 70% identity, or at least 75% identity, or at least 80% identity, or at least 85% identity, or at least 90% identity, or at least 92% identity, or at least 95% identity, or at least 96% identity, or at least 97% identity, or at least 98% identity, or at least 99% identity to a recombinant protein set forth in Table 12, or a derivative or variant or functional portion thereof.
[0027] In certain embodiments, the system or composition comprises a donor nucleic acid. In certain embodiments, the donor nucleic acid comprises a homology arm.
[0028] In certain embodiments, the recombination system is comprised in a cell, eg, a eukaryotic cell, a mammalian cell, an animal cell, a human cell, or a plant cell.
[0029] In one aspect, the invention provides a method of recombination, the method comprising providing in a cell a system or composition: (i) a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence, wherein the target DNA sequence comprises a genomic DNA sequence in the cell; and (ii) a recombination protein, wherein the recombination protein comprises an exonuclease, a single-stranded DNA annealing protein (SSAP), or a single-stranded DNA binding protein (SSB), or a combination of two or more thereof; or (iii) nucleic acid molecule(s) encoding or delivering (i) and / or (ii) for in vivo expression in the cell; or (iv) a vector containing the nucleic acid molecule(s) of (iii) for in vivo expression in the cell.
[0030] In certain embodiments, (i) and (ii) further comprise a Cas protein or a nucleic acid polymerase (including, but not limited to, a reverse transcriptase (RT) or a naturally occurring or engineered polymerase with reverse transcriptase activity such as a Cas protein and an RT), or (iii) comprise a nucleic acid molecule encoding or delivering (i) and / or (ii) and / or the Cas protein and / or RT for in vivo expression in a cell, or the vector(s) of (iv) additionally contain a nucleic acid molecule(s) encoding a Cas protein and / or RT.
[0031] In certain embodiments, one or more of the components are provided as a complex. For example, the protein or fusion protein and nucleic acid are provided as a ribonucleoprotein (RNP). Non-limiting examples of RNPs include CRISPR guide RNA complexes and SSAP guide RNA complexes. In certain embodiments, a fusion protein comprises one or more components. Non-limiting examples include Cas9-SSAP fusions, Cas9-RT fusions, and SSAP-RT fusions.
[0032] In certain embodiments, the target DNA sequence comprises the genomic sequence of albumin (ALB), AAVS1, HSP90AA1, DYNLT1, ACTB, BCAP31, HIST1H2BK, CLTA, or RAB11A.
[0033] In certain embodiments, the system or composition comprises a recruitment system for recruiting a guide nucleic acid and a recombinant protein. In certain embodiments, the recruitment system comprises at least one aptamer sequence and an aptamer-binding protein operably linked to the recombinant protein as part of a fusion protein. In certain embodiments, the at least one aptamer sequence is an RNA aptamer sequence or a peptide aptamer sequence. In certain embodiments, the nucleic acid molecule or multiple nucleic acid molecules additionally comprise at least one RNA aptamer sequence, or comprise one, two, three, or more RNA aptamer sequences. In certain embodiments, the two aptamer sequences comprise the same sequence or comprise sequences that bind to the same aptamer-binding protein.
[0034] In certain embodiments, the aptamer-binding protein comprises an MS2 coat protein, or a functional derivative or variant thereof. In certain embodiments, the aptamer-binding protein comprises a phage N-peptide, or a functional derivative or variant thereof. In certain embodiments, at least one peptide aptamer sequence is conjugated to a guide RNA. In certain embodiments, the at least one peptide aptamer sequence comprises 1 to 24 peptide aptamer sequences. In certain embodiments, two or more aptamer sequences comprise the same sequence. In certain embodiments, the aptamer sequence comprises a GCN4 peptide sequence.
[0035] In certain embodiments, the N-terminus of the recombinant protein is linked to the C-terminus of the aptamer-binding protein. In certain embodiments, the recombinant protein and the aptamer-binding protein are operably linked by a linker. In certain embodiments, the linker comprises 39115.
[0036] In certain embodiments, the recombinant system or composition comprises at least one nuclear localization sequence (NLS), and optionally, the NLS(s) are linked to the recombinant protein. In certain embodiments, the NLS comprises the amino acid sequence of SEQ ID NO: 16. In certain embodiments, the NLS is located at the C-terminus of the recombinant protein or the N-terminus of the recombinant protein.
[0037] In certain embodiments, the recombinant protein comprises a microbial recombinant protein or active portion thereof, a mitochondrial recombinant protein or active portion thereof, a viral recombinant protein or active portion thereof, or a eukaryotic recombinant protein or active portion thereof, including, but not limited to, a recombinant protein set forth in Table 12, or a derivative or variant or functional portion thereof. In certain embodiments, the recombinant protein comprises an amino acid sequence having at least 70% identity, or at least 75% identity, or at least 80% identity, or at least 85% identity, or at least 90% identity, or at least 92% identity, or at least 95% identity, or at least 96% identity, or at least 97% identity, or at least 98% identity, or at least 99% identity to a recombinant protein set forth in Table 12, or a derivative or variant or functional portion thereof.
[0038] In certain embodiments, the system or composition comprises a donor nucleic acid. In certain embodiments, the donor nucleic acid comprises a homology arm.
[0039] In certain embodiments, the recombination system is comprised in a cell, eg, a eukaryotic cell, a mammalian cell, an animal cell, a human cell, or a plant cell.
[0040] In some embodiments, the Cas protein is Cas9 or Cas12a. In some embodiments, the Cas protein is catalytically inactive. In some embodiments, the Cas9 protein is wild-type Streptococcus pyogenes Cas9 or wild-type Staphylococcus aureus Cas9. In some embodiments, the Cas9 protein is Cas9 nickase (e.g., wild-type Streptococcus pyogenes Cas9 with an amino acid substitution of D10A at position 10).
[0041] Also disclosed are eukaryotic cells comprising the systems or vectors disclosed herein.
[0042] Further disclosed herein is a method for changing a target genomic DNA sequence in a host cell.The method comprises contacting a system, composition or vector described herein with a target DNA sequence (for example, introducing a system, composition or vector described herein into a host cell that contains a target genomic DNA sequence).Also disclosed herein is a kit that contains one or more reagents or other components that are useful, necessary or sufficient for carrying out any of the above methods.
[0043] In some embodiments, the invention provides a system or composition comprising: (i) a nucleic acid polymerase(s), such as a reverse transcriptase(s) (RT); (ii) a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence and an RNA for reverse transcription, or a plurality of nucleic acid molecules comprising a guide RNA sequence complementary to a target DNA sequence and an RNA for reverse transcription; (iii) a recombinant protein, wherein the recombinant protein comprises an exonuclease, a single-stranded DNA annealing protein (SSAP), or a single-stranded DNA binding protein (SSB), or a combination of two or more thereof; or (iv) a nucleic acid molecule(s) encoding or delivering (i), (ii), and (iii) for in vivo expression in a cell; or (v) a vector(s) containing the nucleic acid molecule(s) of (iv) for in vivo expression in a cell. In this system or composition involving RT ("RT system or composition"), (iv) can include (i) being an enzyme, (ii) being a nucleic acid molecule(s), and (iii) being a plurality of nucleic acid molecules, or (i) being a nucleic acid molecule(s) encoding an enzyme(s), (ii) being a nucleic acid molecule(s), and (iii) being a protein, or all of (i), (ii), and (iii) being a plurality of nucleic acid molecules. In some embodiments, the RT system or composition can include more than one reverse transcriptase. When more than one reverse transcriptase is present, there can be more than one RNA for reverse transcription. In some embodiments, in the RT system or composition, (i), (ii), and (iii) further comprise a Cas protein, or (iv) further comprises a nucleic acid molecule(s) encoding a Cas protein, e.g., (iv) comprises a nucleic acid molecule(s) encoding or delivering (i) and / or (ii) and / or (iii) and / or a Cas protein for in vivo expression in a cell, or the vector(s) of (v) additionally contain a nucleic acid molecule(s) encoding a Cas protein.
[0044] Reverse transcriptases that can be used in accordance with the present invention include, but are not limited to, reverse transcriptase, retrotransposon reverse transcriptase, retron reverse transcriptase, LINE-1 reverse transcriptase, Ec86 reverse transcriptase, human immunodeficiency virus (HIV) RT, avian myoblastosis virus (AMV) RT, Moloney murine leukemia virus (M-MLV) RT, group II intron RT, group II intron-like RT, chimeric RT, Ma Luoni murine leukemia virus (M-MLV) transcriptase, Rous sarcoma virus (Rous sarcoma virus, RSV), avian myeloblastosis virus (AMV) reverse transcriptase, Lao Sishi-related virus (RAV) reverse transcriptase and myeloblast tumor-related virus (MAV) reverse transcriptase or other avian sarcoma leukovirus (avian sarcoma leukosis virus, ASLV) reverse transcriptase, as well as other naturally occurring and engineered nucleic acid polymerases. Such engineered polymerases include, but are not limited to, human DNA polymerase η, which has reverse transcriptase activity in a cellular environment (Su et al. 2019, J. Biol. Chem. 294(15):6073-81), and Taq DNA polymerase, which has been engineered to enhance reverse transcription and strand displacement (Barnes et al., Front. Bioeng. Biotechnol., 14 January 2021, doi.org / 10.3389 / fbioe.2020.553474).
[0045] In some embodiments, the RT system or composition further comprises a recruitment system comprising at least one aptamer sequence and an aptamer-binding protein operably linked to the recombinant protein as part of a fusion protein. In some embodiments, in the RT system or composition or composition comprising the recruitment system, the at least one aptamer sequence is an RNA aptamer sequence or a peptide aptamer sequence. In some embodiments, the RT system or composition or composition comprising the recruitment system comprises a nucleic acid molecule or multiple nucleic acid molecules additionally comprising at least one RNA aptamer sequence, e.g., the nucleic acid molecule or multiple nucleic acid molecules comprise two RNA aptamer sequences, e.g., in this case, the two RNA aptamer sequences comprise the same sequence. In some embodiments, the RT system or composition or composition comprising a recruitment system has an aptamer-binding protein comprising an MS2 coat protein or a functional derivative or variant thereof, and / or the aptamer-binding protein comprises a phage N-peptide or a functional derivative or variant thereof, and / or the at least one peptide aptamer sequence is conjugated to a Cas protein, and / or the at least one peptide aptamer sequence comprises 1 to 24 peptide aptamer sequences, and / or the aptamer sequences comprise the same sequence. In some embodiments, the RT system or composition or composition comprising a recruitment system has an aptamer sequence comprising a GCN4 peptide sequence.
[0046] In some embodiments of the RT system or composition, the N-terminus of the recombinant protein is linked to the C-terminus of the aptamer-binding protein, and in some embodiments, the RT system or composition further comprises a linker between the recombinant protein and the aptamer-binding protein, for example, in some embodiments, the linker comprises the amino acid sequence of SEQ ID NO: 15.
[0047] In some embodiments of the RT system or composition, the system or composition comprises at least one nuclear localization sequence (NLS), optionally linked to the recombinant protein or the Cas protein or the reverse transcriptase, or to at least one NLS on each of at least two or three of the recombinant proteins, the reverse transcriptase, or the Cas protein, e.g., in some embodiments, the nuclear localization sequence comprises the amino acid sequence of SEQ ID NO: 16. In some embodiments of the RT system or composition, the nuclear localization sequence is on the recombinant protein or the Cas protein at the C-terminus of the recombinant protein.
[0048] In some embodiments of the RT system or composition, the recombinant protein comprises a recombinant protein or an active portion thereof. In some embodiments of the RT system or composition, the recombinant protein comprises a mitochondrial recombinant protein or an active portion thereof. In some embodiments of the RT system or composition, the recombinant protein comprises a viral recombinant protein or an active portion thereof. In some embodiments of the RT system or composition, the recombinant protein comprises a eukaryotic recombinant protein or an active portion thereof. In some embodiments of the RT system or composition, the recombinant protein comprises RecE or RecT or RecE and RecT, or a derivative or variant or functional portion thereof. In some embodiments of the RT system or composition, RecE or a derivative or variant thereof comprises an amino acid sequence having at least 70% (or any integer between 70 and 100%, e.g., at least 71%, 72%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) similarity or identity or homology to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-8. In some embodiments of the RT system or composition, the fusion protein comprises RecT or a derivative or variant thereof. In some embodiments of the RT system or composition, RecT or a derivative or variant thereof comprises an amino acid sequence having at least 70% (or any integer between 70 and 100%, e.g., at least 71%, 72%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%) similarity or identity or homology to an amino acid sequence selected from the group consisting of SEQ ID NOs:9-14.
[0049] In some embodiments of the RT system or composition, the Cas protein is catalytically inactive (less than 5% nuclease activity compared to a wild-type or non-mutated form of the Cas protein) or catalytically dead. In some embodiments of the RT system or composition, the Cas protein comprises Cas9 or Cas12a. In some embodiments of the RT system or composition, the Cas9 protein comprises wild-type Streptococcus pyogenes Cas9 or wild-type Staphylococcus aureus Cas9. In some embodiments of the RT system or composition, the Cas protein comprises a nickase. In some embodiments of the RT system or composition, the nickase comprises wild-type Streptococcus pyogenes Cas9 with an amino acid substitution at position 10, D10A.
[0050] In some embodiments of the RT system or composition, the target DNA sequence is a genomic DNA sequence in the host cell.
[0051] In some embodiments of the RT system or composition, the RT and the recombinant protein are operably linked to each other to form a fusion protein. In some embodiments of the RT system or composition, the aptamer-binding protein and the recombinant protein are operably linked to each other to form a fusion protein. In some embodiments of the RT system or composition, the RT and the Cas protein are operably linked to each other to form a fusion protein. In some embodiments of the RT system or composition, the recombinant protein and the Cas protein are operably linked to each other to form a fusion protein. In some embodiments of the RT system or composition, the RT, the Cas protein, and the recombinant protein are operably linked to each other to form a fusion protein.
[0052] With respect to RT, and linkers or methods for operably linking components of embodiments of a RT system or composition (as well as with respect to linkers or methods for operably linking components of systems or compositions discussed herein that do not include RT), prime editing and twin prime editing are examples. Reference is made to International Publication Nos. WO 2020 / 191241, WO 2020 / 191153, WO 2020 / 191245, WO 2020 / 191243, WO 2020 / 191233, WO 2020 / 191246, WO 2020 / 191249, WO 2020 / 191239, WO 2020 / 191234, WO 2020 / 191242, WO 2020 / 191248, WO 2020191171 and WO 2021 / 226558, including those known as the International Publication Nos. WO 2020 / 191246, WO 2020 / 191249, WO 2020 / 191239, WO 2020 / 191234, WO 2020 / 191242, WO 2020 / 191248, WO 2020191171 and WO 2021 / 226558. Each of International Publication No. 2020 / 191241, International Publication No. 2020 / 191153, International Publication No. 2020 / 191245, International Publication No. 2020 / 191243, International Publication No. 2020 / 191233, International Publication No. 2020 / 191246, International Publication No. 2020 / 191249, International Publication No. 2020 / 191239, International Publication No. 2020 / 191234, International Publication No. 2020 / 191242, International Publication No. 2020 / 191248, International Publication No. 2020191171, and International Publication No. 2021 / 226558 is incorporated herein by reference. The RTs of International Publication No. 2020 / 191241, International Publication No. 2020 / 191153, International Publication No. 2020 / 191245, International Publication No. 2020 / 191243, International Publication No. 2020 / 191233, International Publication No. 2020 / 191246, International Publication No. 2020 / 191249, International Publication No. 2020 / 191239, International Publication No. 2020 / 191234, International Publication No. 2020 / 191242, International Publication No. 2020 / 191248, International Publication No. 2020191171, and International Publication No. 2021 / 226558 can be used in the practice of the present invention.The linkers or operably linking methods of WO 2020 / 191241, WO 2020 / 191153, WO 2020 / 191245, WO 2020 / 191243, WO 2020 / 191233, WO 2020 / 191246, WO 2020 / 191249, WO 2020 / 191239, WO 2020 / 191234, WO 2020 / 191242, WO 2020 / 191248, WO 2020191171, and WO 2021 / 226558 can be used in the practice of the invention.
[0053] In some embodiments, the invention encompasses cells or eukaryotic cells comprising any RT system or composition described or discussed herein. In some embodiments, the invention encompasses methods of modifying a target genomic DNA sequence in a cell comprising the target genomic DNA sequence, the method comprising introducing any RT system or composition discussed or described herein. In some embodiments, the cell or eukaryotic cell is a mammalian cell, or in the methods, the cell or eukaryotic cell is a mammalian cell, e.g., a human cell, e.g., a stem cell. In some embodiments, the method comprises a target genomic DNA sequence encoding a gene product. In some embodiments, the method, introducing into the cell comprises administering to a subject. In some embodiments, the method comprises the subject being a mammalian non-human animal (e.g., an experimental animal such as a rodent, rat, mouse, rabbit, or a domestic animal such as a horse, dog or canine, cat or feline, or a zoo animal (a non-domesticated animal under human captive care), or a production animal such as a cow or pig), or a human. In some embodiments, in the methods, administering comprises in vivo administration. In some embodiments, in the method, the cell or eukaryotic cell or mammalian cell is an ex vivo cell or an in vitro cell. In some embodiments, the method includes, after the introducing step, administering the ex vivo cell or the in vitro cell to a subject, and in such embodiments, the subject is a mammalian non-human animal or a human.
[0054] In some embodiments, the invention involves the use of an RT system or composition for the modification of a target DNA sequence in a cell.
[0055] In systems or compositions discussed herein that do not include RT, aspects of the RT system or RT system (e.g., linker) that are not related to RT can be applied to any system or composition discussed herein that does not include RT.
[0056] The linker can be a peptide of 5 to 30, 10 to 30, 10 to 20, or 15 amino acid residues. The linker can be -(Gly-Gly-Gly-Gly-Ser)- (SEQ ID NO: 560), -(Gly-Gly-Gly-Gly-Ser)- (SEQ ID NO: 561), or -(Gly-Gly-Gly-Gly-Ser)- (SEQ ID NO: 562). In certain embodiments, the linker is -(Gly-Gly-Gly-Gly-Ser)- (SEQ ID NO: 561). The amino acid sequence of SEQ ID NO: 561 can be encoded by the nucleic acid sequence of SEQ ID NO: 563.
[0057] In certain embodiments, the linker is composed of a majority of amino acids that are sterically unhindered, such as glycine and alanine. Exemplary linkers are polyglycines (particularly poly(Gly-Ser), poly(Gly-Ala), and polyalanines). One exemplary suitable linker shown in the examples below is (Gly-Ser), e.g., -(Gly-Gly-Gly-Gly-Ser)- (SEQ ID NO: 560), -(Gly-Gly-Gly-Gly-Ser)- (SEQ ID NO: 561), or -(Gly-Gly-Gly-Gly-Ser)- (SEQ ID NO: 562).
[0058] The linker may also be a non-peptide linker. For example, alkyl linkers such as -NH-, -(CH2)sC(O)-, etc., where s=2 to 20, may be used. These alkyl linkers are preferably lower alkyl (e.g., C 1~4) It may be further substituted with any sterically unhindered group such as lower acyl, halogen (e.g., Cl, Br), CN, NH2, phenyl, etc. [Table 13]
[0059] Accordingly, it is the intent of the present invention not to encompass within the present invention any previously known products, processes for making a product, or methods for using a product, and consequently, the applicant reserves the right, and hereby discloses, a disclaimer of any previously known products, processes, or methods. Furthermore, it is noted that the present invention is not intended to encompass within its scope any products, processes, or methods for making a product or using a product that do not meet the description and enablement requirements of the USPTO (35 USC §112, first paragraph) or the EPO (Article 83 EPC), and consequently, the applicant reserves the right, and hereby discloses, a disclaimer of any previously described products, processes for making a product, or methods for using a product. In practicing the present invention, it may be advantageous to comply with Article 53(c) of the EPC and Articles 28(b) and (c) of the EPC Rules. All rights are expressly reserved to expressly disclaim any embodiments that are the subject of any granted patent(s) of the applicant in this application line or in any other line or in any previously filed application of any third party. Nothing herein should be construed as a commitment.
[0060] It should be noted that in this disclosure, particularly in the claims and / or paragraphs, terms such as "comprises," "comprised," "comprising," and the like can have the meaning ascribed to them in U.S. patent law. For example, they can mean "includes," "included," "including," etc. Terms such as "consisting essentially of" and "consists essentially of" have the meaning ascribed to them in U.S. patent law, for example, they allow for elements not expressly recited, but exclude elements found in the prior art or that affect a basic or novel characteristic of the invention.
[0061] These and other embodiments are disclosed or are obvious from and encompassed by the following detailed description.
[0062] Other aspects and embodiments of the present invention will become apparent in light of the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0063] [Figure 1] Figures 1A and 1B show the reconstructed phylogenetic trees of eukaryotic recombinases RecE (Figure 1A) and RecT (Figure 1B) from yeast and human.
[0064] [Figure 2-1] Figure 2A is a phylogenetic tree and length distribution of RecE / RecT homologs. Figure 2B is a metagenomics distribution of RecE / T. Figure 2C is a schematic diagram showing the central model disclosed herein. Figure 2D is a graph of genome knock-in efficiency of RecE / T homologs. [Figure 2-2] Same as above. [Figure 2-3] Same as above.
[0065] [Figure 3-1] Figures 3A and 3B are graphs of high-throughput sequencing (HTS) reads of homologous recombination repair (HDR) at the EMX1 (Figure 3A) and VEGFA (Figure 3B) loci. Figures 3C-3E are graphs of mKate knock-in efficiency at the HSP90AA1 (Figure 3C), DYNLT1 (Figure 3D), and AAVS1 (Figure 3E) loci in HEK293T cells. Figure 3F is an image of mKate knock-in efficiency in HEK293T cells using RecT. Figure 3G is a schematic of chromatogram traces from an exemplary AAVS1 knock-in strategy and a RecT knock-in group. Figure 3H is a schematic and graph of a mobilization control experiment and the corresponding knock-in efficiency. All results are normalized to NR (NC, no cleavage; NR, no recombinator). [Figure 3-2] Same as above. [Figure 3-3] Same as above.
[0066] [Figure 4-1] Figures 4A-4C are graphs of the relative mKate knock-in efficiency at the HSP90AA1 (Figure 4A), DYNLT1 (Figure 4B), and AAVS1 (Figure 4C) loci in HEK293T cells compared to the NE group. (NC, no cleavage control; NR, no recombinator control). Figure 4D is an example agarose gel image of junction PCR verifying mKate knock-in at the AAVS1 locus. Figures 4E and 4F are graphs of the absolute (Figure 4E) and relative (Figure 4F) LOV knock-in efficiency at the AAVS1 locus. Figure 4G is the Sanger sequencing result of the junction PCR product of an example mKate knock-in at the AAVS1 locus. [Figure 4-2] Same as above.
[0067] [Figure 5-1]Figures 5A-5D are graphs of genome knock-in efficiency at different loci across cell lines A549 (Figure 5A), HepG2 (Figure 5B), HeLa (Figure 5C), and hESC (H9) (Figure 5D). Figure 5E is an image of mKate knock-in in hESC. Figures 5F and 5G are genome-wide off-target site (OTS) counts (Figure 5F) and OTS chromosomal distribution (Figure 5G) from the REDITv1 tool. [Figure 5-2] Same as above. [Figure 5-3] Same as above.
[0068] [Figure 6-1] Figures 6A-6D show graphs of the relative mKate knock-in efficiency at the AAVS1 and DYNT1 loci in the A549 cell line (Figure 6A), the DYNLT1 and HSP90AA1 loci in the HepG2 cell line (Figure 6B), the DYNLT1 and HSP90AA1 loci in the HeLa cell line (Figure 6C), and the HSP90AA1 and OCT4 loci in the hES-H9 cell line (Figure 6D). (NC, no cleavage control; NR, no recombinator control. All data are normalized to the NR control.) Figure 6E shows representative FACS results of HSP90AA1 mKate knock-in in hES-H9 cells. [Figure 6-2] Same as above.
[0069] [Figure 7-1] Figures 7A-7D are graphs of absolute mKate knock-in efficiency of different homology arm lengths with and without recombinator controls for the DYNLT1 (Figure 7A) and HSP90AA1 (Figure 7B) loci and DYNLT1 (Figure 7C) and HSP90AA1 (Figure 7D). [Figure 7-2] Same as above.
[0070] [Figure 8]Figures 8A-8E are graphs of indel rates for the top three predicted off-target loci associated with sgEMX1 (Figures 8A-8C) or sgVEGFA (Figures 8D-8E) in the REDITv1 system.
[0071] [Figure 9-1] Figure 9A is a schematic diagram of a selected embodiment of REDITv2N and corresponding knock-in efficiency in HEK293T cells. Figures 9B and 9C are graphs of genome-wide off-target site (OTS) counts (Figure 9B) and OTS chromosomal distribution (Figure 9C) comparing REDITv2N with REDITv1. Figure 9D is a schematic diagram of a selected embodiment of REDITv2D and corresponding knock-in efficiency. Figure 9E is a graph of the editing efficiencies of REDITv1, REDITv2N, and REDITv2D under serum starvation conditions. Figure 9F is the knock-in efficiency of REDITv3 in hESCs. Figure 9G is an image of mKate knock-in using REDITv3 in hESCs. [Figure 9-2] Same as above. [Figure 9-3] Same as above. [Figure 9-4] Same as above. [Figure 9-5] Same as above.
[0072] [Figure 10-1] 10A and 10B are a schematic and graph of the relative mKate knock-in efficiency of select embodiments of REDITv2N (FIG. 10A) and REDITv2D (FIG. 10B) at the DYNLT1 and HSP90AA1 loci. [Figure 10-2] Same as above.
[0073] [Figure 11] Figures 11A-11D are agarose gel images showing junction PCR of mKate knock-in at the DYNLT1 and HSP90AA1 loci for selected REDITv2N lines, and Figure 11E is a chromatogram sequence of the junction PCR product at the DYNLT1 locus.
[0074] [Figure 12-1] Figures 12A and 12B are graphs of the genomic distribution of detected off-target cleavages for select embodiments of REDITv2 (Figure 12A) and REDITv2N (Figure 12B). Pileups include alignments with two or more reads that overlap each other. Adjacent pairs include alignments that appear on opposite strands within 200 bp upstream of each other. Matched targets include alignments that match the treated target in the upstream sequence (up to six mismatches are allowed in the target sequence, including one mismatch at the PAM). Figure 12C is a graph of HTS HDR and indel reads at the EMX1 locus for the REDITv2N system. [Figure 12-2] Same as above. [Figure 12-3] Same as above.
[0075] [Figure 13-1] Figure 13A is an agarose gel image showing junction PCR of mKate knock-in at the DYNLT1 locus for the REDITv2D system. Figure 13B is a chromatogram sequence of the junction PCR product at the DYNLT1 locus. [Figure 13-2] Same as above.
[0076] [Figure 14] 14A-14C are graphs of mKate knock-in efficiency at the HSP90AA1 locus in REDITv2 (FIG. 14A), REDITv2N (FIG. 14B), and REVITv2D (FIG. 14C) when treated with different FBS concentrations. 14A-14C are graphs of mKate knock-in efficiency at the HSP90AA1 locus in REDITv2 (FIG. 14D), REDITv2N (FIG. 14E), and REVITv2D (FIG. 14F) when treated with different serum FBS concentrations.
[0077] [Figure 15]Figure 15 shows images of nuclear localization of RecE_587 and RecT after EGFP fusion to the REDITv1 system. Nuclei were stained with NucBlue Live Ready Probes Reagent.
[0078] [Figure 16-1] Figures 16A and 16B show the relative mKate knock-in efficiencies at the HSP90AA1 and DYNLT1 loci after fusing different nuclear localization sequences to either the N- or C-terminus of RecT and RecE_587. Figures 16C and 16D show graphs of the absolute mKate knock-in efficiencies of the constructs from Figures 16A and 16B for the DYNLT1 locus (Figure 16C) and HSP90AA1 locus (Figure 16D). [Figure 16-2] Same as above. [Figure 16-3] Same as above. [Figure 16-4] Same as above.
[0079] [Figure 17-1] Figures 17A-17D are graphs of relative (Figures 17A and 17B) and absolute (Figures 17C and 17D) mKate knock-in efficiencies for the DYNLT1 locus (Figures 17A and 17C) and HSP90AA1 locus (Figures 17B and 17D) after fusion of novel NLS sequences and optimal linkers for REDITv2 and REDITv3 variants. REDITv2 versions using REDITv2N (D10A or H840A) and REDITv2D (dCas9) are shown on the horizontal axis along with the number of guides used. Different colors represent different controls and REDIT versions. [Figure 17-2] Same as above.
[0080] [Figure 18] Figure 18 is a graph of the relative editing efficiency of the REDITv3N system at the HSP90AA1 locus in hES-H9 cells.
[0081] [Figure 19-1] Figure 19A is a diagram of an exemplary saCas9 expression vector. Figures 19B-19D are graphs of the relative mKate knock-in efficiency and absolute efficiency (Figures 19D and 19E, respectively) at the AAVS1 locus (Figure 19B) and HSP90AA1 locus (Figure 19D) for different effectors in the saCas9 system. NC, no cleavage control. NR, no recombinator control. [Figure 19-2] Same as above.
[0082] [Figure 20-1] Figure 20A is a schematic diagram of RecT truncation. Figures 20B and 20C are graphs of relative mKate knock-in efficiency at the DYNLT1 locus for wild-type Streptococcus pyogenes Cas9 and Streptococcus pyogenes Cas9n(D10A) with single and double nicking. [Figure 20-2] Same as above.
[0083] [Figure 21-1] Figure 21A is a schematic diagram of RecE_587 truncation. Figures 21B and 21C are graphs of relative mKate knock-in efficiency at the DYNLT1 locus for wild-type Streptococcus pyogenes Cas9 and Streptococcus pyogenes Cas9n(D10A) with single and double nicking. [Figure 21-2] Same as above.
[0084] [Figure 22]Figures 22A and 22B are graphs comparing the efficiency of recombineering-based editing using various exonucleases (Figure 22A) and single-stranded DNA annealing proteins (SSAPs) (Figure 22B) from a naturally occurring recombineering system, including NR (no recombinator) as a negative control. Gene editing activity was measured using mKate knock-in assays at genomic loci (DYNLT1 and HSP90AA1). Data shown are the percentage of successful mKate knock-ins using human HEK293 cells, and each experiment was performed in triplicate (n=3).
[0085] [Figure 23-1] Figures 23A-23E show a compact mobilization system using boxB and N22. The REDIT recombinator protein was fused to the N22 peptide, and the sgRNA contained boxB, a short recognition sequence for the N22 peptide (Figure 23A). Figures 23B-23E show graphs of gene editing efficiency using mKate knock-in assays with wild-type SpCas9, compared side-by-side with the MS2-MCP mobilization system. Figures 23B and 23D show absolute mKate knock-in efficiencies at the DYNLT1 and HSP90AA1 loci, while Figures 23C and 23E show relative efficiencies. Data shown are the percentage of successful mKate knock-in using HEK293 human cells; each experiment was performed in triplicate (n=3). [Figure 23-2] Same as above. [Figure 23-3] Same as above.
[0086] [Figure 24-1]Figures 24A-C show the SunTag mobilization system. The REDIT recombinator protein was fused to an scFV antibody, and the GCN4 peptide was fused in tandem (10 copies of the GCN4 peptide separated by a linker) to the Cas9 protein (Figure 24A). Gene editing knock-in efficiency was measured using an mKate knock-in experiment with the DYNLT1 locus (Figure 24B) (Figure 24C). All data are measurements of gene editing efficiency using an mKate knock-in assay with wild-type SpCas9. The absolute mKate knock-in efficiency at DYNLT1 is shown in the lower right corner of each flow cytometry plot, where the control was no recombinator (NR), which included an scFV fused to GFP protein as a negative control. All experiments were performed in HEK293 human cells. [Figure 24-2] Same as above. [Figure 24-3] Same as above.
[0087] [Figure 25-1] Figures 25A and 25B illustrate REDIT using the Cas12A system. A Cpf1 / Cas12a-based REDIT system was created for two different Cpf1 / Cas12a proteins via the SunTag mobilization design (Figure 25A). An mKate knock-in assay was used to measure efficiency at two endogenous loci (DYNLT1 and AAS1) (Figure 25B). Absolute mKate knock-in efficiency, as measured by the percentage of mKate+ cells using HEK293 human cells, is shown; each experiment was performed in triplicate (n=3), with a negative control of no recombinator (NR). [Figure 25-2] Same as above.
[0088] [Figure 26-1]Figures 26A and 26B show the measurement of precise recombineering activity through mKate knock-in gene editing assays using RecE and RecT homologs at the DYNLT1 locus (Figure 26A) and HSP90AA1 locus (Figure 26B). Absolute mKate knock-in efficiency, as measured by the percentage of mKate+ cells, is shown using HEK293 human cells, with each experiment performed in triplicate (n=3). Negative controls are no recombinator (NR) and no cleavage (NC). Original RecE and RecT from E. coli were also included as positive controls. [Figure 26-2] Same as above.
[0089] [Figure 27] Figures 27A and 27B are a schematic showing SunTag-based recruitment of SSAP RecT to the Cas9-gRNA complex for gene editing (Figure 27A) and a graph quantifying the editing efficiency of SunTag compared to MS2-based strategies (Figure 27B).
[0090] [Figure 28]Figures 28A-28C show a comparison of REDIT with alternative HDR-enhancing gene editing approaches. Figure 28A is a schematic diagram showing alternative HDR-enhancing approaches by fusing the functional domains CtIP or Geminin (Gem) to the Cas9 protein (left) and in combination with REDIT (right). Figure 28B shows an alternative small molecule HDR-enhancing approach through cell cycle control. Nocodazole was used to synchronize cells at the G2 / M boundary (left) according to the timeline shown (right). Figure 28C shows a comparison of gene editing efficiency using REDIT and alternative HDR-enhancing tools, Cas9-HE (CtIP fusion), Cas9-Gem (Geminin fusion), and nocodazole (noc), as well as the combination of REDIT and these methods (Cas9-HE / Cas9-Gem / noc+REDIT). Donor DNA contains 200 + 400 bp (DYNLT1) or 200 + 200 bp (HSP90AA1) of HA. All assays were performed with no donor, NTC, and Cas9 (no enhancement) controls. #P<0.05 compared to REDIT; ##P<0.01 compared to REDIT.
[0091] [Figure 29-1]Figures 29A-29D show the template design guidelines, ligation accuracy, and performance of the REDIT gene editing method. Figure 29A is a graph of homology arm (HA) length testing comparing different template designs for HDR donors (longer HA) or NHEJ / MMEJ donors (zero / shorter HA) using REDIT and Cas9 references. Above and below are two genomic loci tested using the mKate knock-in assay. Figure 29B shows the design of an exemplary junction profiling assay using genomic PCR to isolate knock-in clones and subsequently use primers (fwd, rev) that bind outside the donor to avoid template amplification. Paired Sanger sequencing of the PCR products reveals homologous and non-homologous editing at the 5'- and 3'-junctions. Figure 29C is a graph of the percentage of colonies with the indicated ligation profiles from Sanger sequencing of knock-in clones like those in Figure 29B. The editing methods and donor DNA are listed below (HA lengths are shown in parentheses). Figure 29D is a graph of knock-in efficiency using a 2 kb cassette to insert a dual GFP / mKate tag to validate the Cas9-based REDIT method. The HA length of the donor DNA is indicated below. [Figure 29-2] Same as above.
[0092] [Figure 30-1]Figures 30A-30C (Figures 6C-6E) show GIS-seq results demonstrating that REDIT is an efficient method with the ability to insert kilobase-long sequences with fewer unwanted editing events. Figure 30A is a schematic diagram showing the design, procedure, and analysis steps for GIS-seq to measure genome-wide insertion sites of knock-in cassettes. High-molecular-weight (HMW) genomic DNA purification was required to remove potential contamination from donor DNA. The donor DNA had 200 bp of HA on each side. Figure 30B shows representative GIS-seq results showing plus / minus reads at the on-target locus DYNLT1. The predicted 2A-mKate knock-in site before the stop codon of the last exon is centered in the trimmed read (the read clipped to remove the 2A-mKate cassette). Template mutations avoid gRNA targeting and help distinguish between labeled and edited genomic reads. Figure 30C is a summary of the top GIS-seq insertion sites comparing the Cas9dn and REDITdn groups, showing the predicted on-target insertion sites (highlighted) and the reduced number of identified off-target insertion sites when using REDITdn. (Left) DYNLT1 and (right) ACTB loci, with MLE calculated from the distribution of filtered and trimmed GIS-seq reads. [Figure 30-2] Same as above.
[0093] [Figure 31-1]Figures 31A-31F show the dependence of REDIT gene editing on endogenous DNA repair and the application of the REDIT method for human stem cell engineering. Figure 31A is a model showing the editing process and major repair pathways involved when using REDIT or Cas9 for gene editing, with the HDR pathway highlighted relative to chemical perturbation (inhibition of RAD51). Donor DNA with 200 + 200 bp of HA is used for all inhibitor experiments. Figures 31B and 31C are graphs showing the relative knock-in efficiency of the REDIT tool compared to a Cas9 control treated with the RAD51 inhibitors B02 and RI-1 or vehicle for wtCas9-based REDIT and Cas9 (Figure 31B) and Cas9 nickase-based REDITdn and Cas9dn (Figure 31C). All conditions were measured in 1 kb knock-in assays at two genomic loci (DYNLT1 and HSP90AA1). Figure 31D is a graph of knock-in efficiency in hESCs (H9) using REDIT and REDITdn tested across three genomic loci compared to the corresponding Cas9 and Cas9dn references. Figures 31E and 31F are flow cytometry plots of the results of mKate knock-in in hESCs using REDIT, REDITdn, along with Cas9, Cas9dn, and an NTC control. The donor DNA in the hESC experiments has 200 + 200 bp of HA across all loci tested. [Figure 31-2] Same as above. [Figure 31-3] Same as above.
[0094] [Figure 32-1] Figures 32A-32B show chemical perturbations of dCas9 REDIT. Gene editing efficiency was determined upon treatment with mammalian DNA repair pathway inhibitors (Mirin, RI-1, and B02) with (Figure 32A) and without (Figure 32B) cell cycle inhibitor (Thy, double thymidine) blocking. Statistical analysis was performed using a t-test at 1% FDR via a two-step step-up method. [Figure 32-2] Same as above.
[0095] [Figure 33-1] Figures 33A and 33B are schematic diagrams of the DNA components (gene editing vector and template DNA) and tail vein injection in mice, respectively. [Figure 33-2] Same as above.
[0096] [Figure 34-1] Figures 34A-34C show the results of tail vein injection of mice with gene editing vectors. Figure 34A shows a schematic of PCR analysis and gel electrophoresis of hepatocytes from injected mice. Figure 34B shows the results of Sanger sequencing of PCR amplicons. Figure 34C shows a schematic of next-generation sequencing and a graph of quantification of knock-in junction errors. [Figure 34-2] Same as above. [Figure 34-3] Same as above.
[0097] [Figure 35-1] Figures 35A and 35B are schematic diagrams of DNA components (gene editing and control vectors) and adeno-associated virus (AAV) treatment, respectively. Figure 35C is a graph of fluorescent images of lungs from AAV-treated mice and corresponding quantification of tumor numbers. [Figure 35-2] Same as above. [Figure 35-3] Same as above.
[0098] [Figure 36] Figures 36A to 36C show the predicted structures of E. coli RecT (EcRecT) alone (Figure 36A) and with single-stranded DNA bound thereto (Figures 36B and 36C).
[0099] [Figure 37-1] Figures 37A-B show predicted interactions of EcRecT SSAP amino acids with ssDNA. [Figure 37-2] Same as above.
[0100] [Figure 38-1]Figures 38A-38F show the development of a dCas9 gene editor by mining microbial SSAPs. (Figure 38A) Schematic model of the dCas9 editor using single-strand annealing proteins (SSAPs). (Figure 38B) Design of a genome knock-in assay to measure gene editing efficiency (left); workflow of the SSAP screening experiment (right). (Figure 38C) Construct design for screening the gene editing efficiency of SSAPs using a 2A-mKate knock-in assay with an 800-bp transgene. (Figure 38D) Initial screening results for three SSAPs: the Bet protein from lambda phage (LBet), the RecT protein from Rac prophage (RacRecT), and gp2.5 from T7 phage (T7gp2.5). (Figure 38E) Screening of RecT-like SSAP candidates by metagenomic homolog mining and knock-in assay. The most active candidate is labeled dCas9-SSAP. NTC: non-targeting control. The lengths of the homology arms (HA) were: DYNLT1, 200 + 200 bp; HSP90AA1, 200 + 400 bp; ACTB, 200 + 400 bp. Donor templates were added to all groups except the no-donor control. (Figure 38F) Gene editing efficiency was measured using three donor designs with different HA lengths at the DYNLT1 (left) and HSP90AA1 (right) loci in HEK293T cells. All results in this and the following figures are from replicate experiments (n = 3), with error bars representing the standard error of the mean (SEM) unless otherwise noted. [Figure 38-2] Same as above. [Figure 38-3] Same as above. [Figure 38-4] Same as above. [Figure 38-5] Same as above.
[0101] [Figure 39-1]Figures 39A-H show on-target and off-target editing errors of dCas9-SSAP. (Figure 39A) Deep sequencing to measure the level of indel formation when using dCas9-SSAP and the Cas9 reference in endogenous targets. The donor template used is a 200-bp-HA HDR template. Assay details are described in Methods. (Figure 39B) Clonal Sanger sequencing to analyze the accuracy of knock-in editing using dCas9-SSAP and the Cas9 reference with different HDR and MMEJ donors. The donor templates used were a 200-bp-HA HDR template and a 25-bp-HA MMEJ template (Methods and Supplementary Material). (Figure 39C-E) Genome-wide detection of knock-in cassette insertion sites using unbiased sequencing. (Figure 39C) Representation of the workflow. (Figure 39D) Representative reads aligned to the knock-in genomic site. (e) Summary of detected on-target and off-target insertion sites. (Figure 39F-G) Workflow and results for measuring cytocompatibility effects, defined by the percentage of viable cells after editing (normalized to mock control). (Figure 39H) Summary analysis of knock-in accuracy of the dCas9-SSAP editor compared to the Cas9 HDR and Cas9 MMEJ methods. Accuracy is defined as the overall yield (%) of correct knock-ins within all edited results (correct knock-ins, knock-ins with indels, and NHEJ indels). [Figure 39-2] Same as above. [Figure 39-3] Same as above. [Figure 39-4] Same as above.
[0102] [Figure 40-1]Figures 40A-40G show validation of the dCas9-SSAP editor and comparison with the Cas9 reference and other HDR enhancement methods. (Figure 40A) Comparison of efficiency using dCas9-SSAP and other alternative Cas9, nCas9, and HDR enhancement tools: Cas9-HE (CtIP-fused Cas9), Cas9-Gem (geminin-fused Cas9), nCas9 (Cas9-D10A nickase reference), and nCas9-hRAD51 (improved Cas9 nickase editor). The donor template is the same as in Figure 1. (Figure 40B) Imaging validation of mKate knock-in at the endogenous genomic locus using the dCas9-SSAP editor. (Figure 40C) Design of knock-in donors with different transgene lengths. (Figure 40D) Knock-in efficiency for different transgene lengths using the dCas9-SSAP editor. The donor HA lengths are 200bp + 200bp for DYNLT1 and 200bp + 400bp for HSP90AA1. (Figure 40E) Performance of the dCas9-SSAP editor compared to the Cas9 reference across seven endogenous loci in HEK293T cells. ND, no donor control; NT, non-targeted control. (Figure 40F-G) Knock-in gene editing in human embryonic stem cells (hESCs, H9) using the dCas9-SSAP editor, with quantified HDR efficiency (Figure 40F) and flow cytometry analysis (Figure 40G). All statistical analyses are performed using multiple t-tests to compare across all genomic targets at a false discovery rate (FDR) of 1% using the Benjamini, Krieger, and Yekutieli two-step step-up method. [Figure 40-2] Same as above. [Figure 40-3] Same as above. [Figure 40-4] Same as above. [Figure 40-5] Same as above. [Figure 40-6] Same as above.
[0103] [Figure 41-1]Figures 41A-41D show chemical perturbations to probe the editing mechanism of the dCas9-SSAP editor. Gene editing efficiency of the dCas9-SSAP editor when treated with DNA repair pathway inhibitors (Mirin, RI1, and B02) without (Figures 41A, 41B) or with (Figures 41C, 41D) cell cycle synchronization (DTB, double thymidine blockade). All donor templates are the same as in Figure 38. Statistical analysis is from a t-test at 1% FDR using the Benjamini, Krieger, and Yekutieli two-step step-up method. [Figure 41-2] Same as above.
[0104] [Figure 42-1] Figures 42A-D show the minimization of the dCas9-SSAP editor as a compact CRISPR knock-in tool for convenient delivery. (Figure 42A) Schematic showing the EcRecT predicted secondary structure and priming sites for constructing truncated EcRecT proteins based on structural prediction. (Figure 42B) Relative knock-in efficiencies of various truncation designs. All groups were normalized to the Cas9 standard (separately for each target). (Figure 42C) Schematic of the dSaCas9-mSSAP system in an AAV construct using compact SaCas9 (left, element sizes not shown to scale) and (Figure 42D) knock-in efficiencies at the AAVS1 and HSP90AA1 endogenous targets via in vitro delivery of AAV2 vectors carrying the original and minimized dSaCas9-SSAP editors in HEK293T cells. [Figure 42-2] Same as above. [Figure 42-3] Same as above.
[0105] [Figure 43-1]Figures 43A-E show gel electrophoresis and sequencing verification of knock-in-specific PCR products using dCas9-SSAP. (Figure 43A) Agarose gel results of knock-in-specific junction PCR at the DYNLT1 locus. (Figures 43B-E) Sanger sequencing chromatograms of genomic junctions from knock-in experiments at the DYNLT1 locus. For all samples, Applicant amplified the 5' (Figure 43B, Figure 43C) and 3' (Figure 43D, Figure 43E) ends of genomic DNA using primers spanning the outside junction of the donor DNA to confirm knock-in. [Figure 43-2] Same as above.
[0106] [Figure 44] Figure 44 shows a phylogenetic tree and amino acid alignment of representative RecT homologs, along with annotated protein conserved domains.
[0107] [Figure 45] Figures 45A-B show deep sequencing of short sequence edits comparing the dCas9-SSAP editor and the Cas9 editor. (Figure 45A) Donor design for a 16-bp substitution in EMX1. (Figure 45B) Analysis of accurate HDR and indel editing results using deep sequencing at the EMX1 genomic locus. The first PCR used sequencing primers completely outside the donor to ensure that sequencing results were free of donor template contamination, which was verified by a non-targeting control (when donor DNA was delivered into cells).
[0108] [Figure 46]Figures 46A-B are schematic diagrams showing the workflow used in Sanger sequencing of knock-in products (Figure 46A) and the sequencing method used in the deep on-target indel assay (Figure 46B). The assays described here correspond to Figure 41. gPCR, genomic PCR. Sequence F / sequence R are primers for Sanger sequencing that bind upstream / downstream of the knock-in template.
[0109] [Figure 47-1] Figures 47A-B show Sanger sequencing chromatograms of genomic junctions from dCas9-SSAP experiments at the DYNLT1 locus. The sequences in the red boxes were not correctly repaired. For all samples, the 5' (Figure 47A) and 3' (Figure 47B) ends of the genomic DNA were amplified using primers spanning the junction to confirm knock-in accuracy. The genome-binding primers used are completely outside the donor DNA to avoid contamination. [Figure 47-2] Same as above.
[0110] [Figure 48-1] Figures 48A-B show Sanger sequencing chromatograms of genome junctions from dCas9-SSAP experiments at the HSP90AA1 locus. For all samples, the 5' (Figure 48A) and 3' (Figure 48B) ends of the genomic DNA were amplified using primers spanning the junction to confirm knock-in accuracy. The genome-binding primers used are completely outside the donor DNA to avoid contamination. [Figure 48-2] Same as above.
[0111] [Figure 49]Figures 49A-B show genome-wide insertion site mapping and quantification. (Figure 49A) Overall workflow for the unbiased genome-wide insertion site mapping process. On-target and off-target insertion sites are recovered from reads aligning to the reference genome (hg38). The complete protocol and data analysis pipeline are detailed in the Methods. (Figure 49B) Genome-wide insertion site quantification, counting all aligned reads (with valid UMIs), showed a decrease in insertion site abundance using Cas9-SSAP compared to Cas9 HDR across two genomic loci (DYNLT1 and HSP90AA1). Insertion site abundance is measured as RPKU or Reads Per Thousand UMIs.
[0112] [Figure 50] Figures 50A-50B show testing of the dCas9-SSAP editor tool using single-guide (Figure 50A) and dual-guide (Figure 50B) designs across three genomic targets (shown above). The donor DNA used is the same as shown in Figure 3a with an 800 bp knock-in design.
[0113] [Figure 51] Figures 51A-C show validation of dCas9-SSAP knock-in efficiency in three additional cell lines: HepG2 (Figure 51A), HeLa (Figure 51B), and U2OS (Figure 51C). The knock-in experiments used the same donor DNA carrying an approximately 800 bp cassette encoding the 2A-mKate transgene for all cell lines tested.
[0114] [Figure 52]Figures 52A-C show a complete set of flow cytometry analysis data using the dCas9-SSAP editor for human stem cell engineering. Flow cytometry analysis of knock-in gene editing at the HSP90AA1 (Figure 52A), ACTB (Figure 52B), and OCT4 (Figure 52C) endogenous loci in human embryonic stem cells (hESCs, H9) using dCas9-SSAP compared to a non-targeting control and a Cas9 (Cas9 HDR) reference.
[0115] [Figure 53] Figure 53 is a schematic diagram showing the RecT protein secondary structure predicted using an online tool (CFSSP, see Methods). The prediction results (secondary structure visualized above, alignment below) formed the basis for developing truncated functional RecT variants.
[0116] [Figure 54-1]Figures 54A-C show optimization of dCas9-SSAP for efficient and durable gene editing. (Figure 54A) Knock-in efficiency for SSAP dosage optimization. The donor HA length is approximately 200 bp for DYNLT1 and approximately 300 bp for HSP90AA1. n=3 biologically independent experiments. (Figure 54B) Performance of the dCas9-SSAP editor compared to the Cas9 reference across seven endogenous loci in HEK293T cells after SSAP dosage optimization and donor HA extension. The donor HA lengths are approximately 200 bp for DYNLT1, 673 bp + 750 bp for HSP90AA1, 500 bp + 800 bp for ACTB, 608 bp + 740 bp for BCAP31, 212 bp + 413 bp for HIST1H2BK, 705 bp + 602 bp for CLTA, and 464 bp + 440 bp for RAB11A. All knock-in donors target the C-terminus of the endogenous protein, except for the CLTA / RAB11A donor, which targets the N-terminus. n = 3 biologically independent experiments. (Figure 54C) Stability of transgene expression in HSP90AA1 (left) and ACTB (right) after sorting 3 days after dCas9-SSAP knock-in. Variable sorting efficiency resulted in different starting mKate+ rates (Figure 57). Error bars in a and b represent the standard error of the mean (SEM). [Figure 54-2] Same as above.
[0117] [Figure 55-1]Figures 55A-C show optimization of donor dosage and homology arms of donor DNA. (Figure 55A) Quantification of genomic mKate knock-in efficiency at the DYNLT1, HSP90AA1, and ACTB loci for donor dosage optimization using the dCas9-SSAP editor. Non-targeted, non-targeted control. The donor HA lengths are 200bp + 200bp for DYNLT1, 200bp + 400bp for HSP90AA1, and 200bp + 400bp for ACTB. Quantification of mKate knock-in efficiency at the HSP90AA1 (Figure 55B) and ACTB (Figure 55C) loci for donor homology arm (HA) optimization using the dCas9-SSAP editor. Non-targeted, non-targeted control. The donor HA lengths were 200bp + 200bp or 673bp + 750bp for HSP90AA1 and 200bp + 400bp or 500bp + 800bp for ACTB. n = 3 biologically independent experiments. All results in this figure are from replicate experiments, with error bars representing the standard error of the mean (SEM). [Figure 55-2] Same as above.
[0118] [Figure 56-1]Figures 56A-D show validation of the dCas9-SSAP editor by protein function assays. (Figure 56A) Design of a genomic puromycin / blasticidin-resistant cassette knock-in assay to validate functional on-target editing by dCas9-SSAP. (Figure 56B) Immunoblotting confirms the presence and size of on-target dCas9-SSAP knock-in products at the HSP90AA1 and ACTB loci, performed with an anti-V5 antibody that recognizes in-frame fusions with endogenous proteins. Data shown represent three biologically independent experiments. (Figures 56C-D) Validation and quantification of on-target knock-in using dCas9-SSAP by colony formation assay. Cells were selected by the knock-in resistance cassette and then quantified using CrystalViolet staining. Scale bar = 500 μm. n = 4 biologically independent experiments. Error bars represent the standard error of the mean (SEM). [Figure 56-2] Same as above. [Figure 56-3] Same as above.
[0119] [Figure 57-1] Figures 57A-57E show validation of on-target editing stability. (Figure 57A) Workflow of a long-term time-course experiment to assess the stability of editing results using the dCas9-SSAP editor. (Figure 57B) Flow cytometry analysis of knock-in gene editing at the HSP90AA1 and ACTB endogenous loci at different time points after delivery of dCas9-SSAP and donor DNA. (Figures 57C-57D) Representative crystal violet-stained images of on-target puromycin knock-in at the HSP90AA1 and ACTB loci. Scale bar = 500 μm. The assay was performed three times with similar results. (Figure 57E) Quantification of HSP90AA1 and ACTB gene expression levels in HEK293T cells by bulk RNA-seq analysis, demonstrating significantly higher levels of HSP90AA1 expression. This resulted in better cell survival in the HSP90AA1 group compared to the ACTB group. Data from two biologically independent experiments are presented. [Figure 57-2] Same as above. [Figure 57-3] Same as above.
[0120] [Figure 58] Figure 58 shows SSAP+Cas9-mediated knock-in editing with inactivating guide RNA (dgRNA). SSAP+Cas9 contains RecT and wtCas9. mKate knock-ins are shown for DYNLT1, HSP90AA1, and ACTB.
[0121] [Figure 59] Figure 59 shows dCas9-SSAP-mediated knock-in of a 600-bp transgene expressing luciferase or mKate at the human albumin (ALB) locus (top) or AAVS1 locus (bottom) in human HEK293T cells or human hepatocytes. The transgene knock-in at the albumin locus was highly expressed in hepatocytes but not in HEK293T (top). The transgene knock-in at the AAVS1 locus was highly expressed in HEK293T but not in hepatocytes (bottom).
[0122] [Figure 60] Figure 60 shows dCas9-SSAP-mediated knock-in of an 800-bp transgene expressing luciferase or mKate at the human albumin (ALB) locus (top) or AAVS1 locus (bottom) in human HEK293T cells or human hepatocytes. The transgene knock-in at the albumin locus was highly expressed in hepatocytes but not in HEK293T (top). The transgene knock-in at the AAVS1 locus was highly expressed in HEK293T but not in hepatocytes (bottom).
[0123] [Figure 61]Figure 61 shows electroporation of RNPs containing an 800-bp mKate-encoding transgene in K562 cells. Cells were electroporated with RNPs containing purified Cas9 or dCas9 protein complexed with guide RNA, a double-stranded 800-bp mKate transgene, with and without RecT. Knock-in was at the HSP90AA1 (left) locus or the HIST1H2BK (right) locus.
[0124] [Figure 62] Figure 62 shows delivery of RNPs containing Cas9 or dCAS9, with or without SSAP, into mouse primary hematopoietic stem cells (HSCs) and AAV6 to knock-in a GFP-expressing transgene.
[0125] [Figure 63] Figure 63 shows transgene expression.
[0126] [Figure 64-1] Figures 64A-B show SSAP-mediated knock-in of a transgene using an R-loop-forming guide without CRISPR components. (Figure 64A) Model of guide RNA-SSAP-mediated gene editing showing MCP-MS2 aptamer pairing of SSAP and R-loop-gRNA. (Figure 64B) Vector / plasmid design for expressing guide RNA, dsDNA donor, and SSAP. (Figure 64C) Raw FACS plot showing identification of mCherry-expressing subsets. (Figure 64D) Transgene knock-in at the ACTB locus using various guide lengths (18 nt, 20 nt, 25 nt) with and without RecT. [Figure 64-2] Same as above. [Figure 64-3] Same as above.
[0127] [Figure 65]Figures 65A-65B show R-loop guide RNA designs. R-loop-guide RNAs contain two components, a guide and a scaffold, shown in guide-scaffold and scaffold-guide configurations (e.g., a guide at the 5' or 3' end of the scaffold). The guide sequence is designed to match the target DNA. One or more aptamers can be incorporated, including, but not limited to, MS2, PP7, and BoxB. MS2 is shown.
[0128] [Figure 66] Figure 66 shows a chimeric guide RNA containing the MS2 / PP7-aptamer.
[0129] [Figure 67] Figure 67 shows the effect of varying guide lengths on knock-in efficiency at the ACTB locus comparing R-loop-SSAP (without CRISPR), Cas9 HDR, and dCas9-SSAP. The donor only contained the mKate knock-in donor.
[0130] [Figure 68] Figure 68 shows the effect of varying guide lengths on knock-in efficiency at the HIST locus comparing R-loop-SSAP (without CRISPR), Cas9 HDR, and dCas9-SSAP. The donor only contained the mKate knock-in donor.
[0131] [Figure 69] Figure 69 shows the effect of varying guide lengths on knock-in efficiency at the HSP90AA1 locus comparing R-loop-SSAP (without CRISPR), Cas9 HDR, and dCas9-SSAP. The donor only contained the mKate knock-in donor.
[0132] [Figure 70]Figure 70 shows R-loop-SSAP-mediated knock-in of a 600-bp transgene expressing luciferase or mKate at the human albumin (ALB) locus (top) or AAVS1 locus (bottom) in human HEK293T cells or human hepatocytes. The transgene knock-in at the albumin locus was highly expressed in hepatocytes but not in HEK293T (top). The transgene knock-in at the AAVS1 locus was highly expressed in HEK293T but not in hepatocytes (bottom).
[0133] [Figure 71] Figure 71 shows R-loop-SSAP-mediated knock-in of an 800-bp transgene expressing luciferase or mKate at the human albumin (ALB) locus (top) or AAVS1 locus (bottom) in human HEK293T cells or human hepatocytes. The transgene knock-in at the albumin locus was highly expressed in hepatocytes but not in HEK293T (top). The transgene knock-in at the AAVS1 locus was highly expressed in HEK293T but not in hepatocytes (bottom).
[0134] [Figure 72] Figure 72 shows a schematic comparing RNA-mediated SSAP editing without reverse transcriptase (top) and with reverse transcriptase (bottom). The RNA template / donor shown contains a homology arm (HA) region with one HA. In some embodiments, the RNA template / donor contains two HA regions, one at each end. Because the HA matches the genomic region next to the editing site, SSAP can facilitate editing.
[0135] [Figure 73]Figure 73 shows the insertion rate of 4-bp sequences inserted into the human EMX1 locus using in vitro transcribed (IVT) RNA templates. The components of the system are: 1. gE3, sg2, or sgHE SpCas9 guide RNAs targeting human EMX1; 2. dg2, a dg3 inactivation / inactivation guide RNA that binds to a region near gE3 / sg2 / sgHE; and 3. 100 / 200 HA: a 100-bp or 200-bp HA region flanking the 4-bp edit on both ends.
[0136] [Figure 74] Figure 74 shows the sense- and antisense-oriented U6-expressing RNA templates used to replace a 16-bp sequence (introduce a 16-bp edit) in human EMX1. The components of the system include: 1. an SpCas9 guide RNA targeting human EMX1; 2. an inactive / inactivating guide RNA that binds to the region. The numbers at the top of each gel lane indicate the length of the homology arm (HA) region next to the edit on both ends.
[0137] [Figure 75] Figure 75 shows the dosage relationship of sense- or antisense-oriented U6-expressing RNA templates used to replace a 16-bp sequence (introduce a 16-bp edit) in human EMX1. The components of the system include: 1. an SpCas9 guide RNA targeting human EMX1; 2. an inactive / inactivating guide RNA that binds to the region.
[0138] [Figure 76] Figure 76 shows a system of the invention inserted into the human AAVS1 locus (Figure 76A) and repair of a defective Venus (green fluorescent protein) locus (Figure 76B).
[0139] [Figure 77]Figure 77 shows a schematic of the sgRNA+dgRNA system (top) and an example based on a TLR locus (bottom). SpCas9 guide_20bp sgRNA: A 20bp guide for the TLR-targeting sgRNA used in the first guide RNA. SaCas9 guide designs are also shown. dg1-dg6: Different 15bp / 16bp inactive / inactivating guides in the TLR-targeting dgRNA used in the second guide RNA containing an aptamer for recruitment.
[0140] [Figure 78-1] Figure 78 shows a demonstration of the sgRNA + dgRNA system, with a signal achieved in the GFP (green) channel indicating repair of the Venus protein. pA19 encodes Cas9 and a guide RNA with two guides, sg318 / sg530, that target the TLR gene editing reporter genomic region. BB is the main strand and serves as a negative control. Dg532 / 534 / 536 / 538 are different inactive guide RNAs containing the MS2 / PP7 aptamer. The red box indicates that the design with the RNA repair template / donor for Venus at the 3' end of the RNA performed best. [Figure 78-2] Same as above. [Figure 78-3] Same as above. [Figure 78-4] Same as above. [Figure 78-5] Same as above.
[0141] [Figure 79] Figure 79 shows a schematic diagram of an sgRNA + dgRNA system involving direct fusion of an RNA template / donor with a dgRNA. In certain embodiments, one or both of the sgRNA and dgRNA can be circular RNA (e.g., a configuration in which the 5' and 3' ends of the RNA are covalently linked). Circular RNA can increase stability and efficiency.
[0142] [Figure 80-1] Figure 80 shows a demonstration of the system in which dgRNA was fused to an RNA template donor and SSAP. The box highlights significantly higher Venus repair than the control. [Figure 80-2] Same as above. [Figure 80-3] Same as above.
[0143] [Figure 81-1] Figure 81 shows a study of pol2 (CMV) versus pol3 (U6) promoter TLR genome editing. [Figure 81-2] Same as above.
[0144] [Figure 82] Figure 82 shows an example of an sgRNA+dgRNA system based on the EMX1 locus. gE3: 20 bp guide for sgRNA targeting EMX1 used in the first guide RNA; dg1-dg8 (15 bp / 16 bp inactive / inactivating guide in dgRNA targeting EMX1 used in the second guide RNA with a fusion RNA template / donor.
[0145] [Figure 83] Figure 83 shows a dgRNA with a fusion RNA template donor and SSAP targeted to the human EMX1 site. pA19 contains Cas9 and a guide RNA with two guides, sg334 / sg516, that target the EMX1 gene editing reporter genomic region. BB is the main strand and serves as a negative control. dg518 / 520 / 522 / 526 is a different inactive guide RNA with an MS2 / PP7 aptamer that binds to a position near sg334 / sg516 and has a 36-bp linker (36L) between the fusion RNA template / donor. All designs are antisense and contain a 300-bp homology arm region (a300+300). Red box: Designs with optimal positions for the dg guide RNA supported higher editing efficiency.
[0146] [Figure 84]Figure 84 shows a schematic diagram of a system incorporating SSAP and prime editing. The MS2 aptamer recruits SSAP-MCP to the Cas9-RT complex. This system 1) provides a locally reverse-transcribed template donor for SSAP, 2) bypasses the endogenous HDR machinery restricted to dividing cells, and 3) benefits from the use of the homology arm regions of the template / donor, enabling SSAP editing by Cas / nCas / dCas.
[0147] [Figure 85-1] Figure 85 shows SSAP+ prime-mediated editing at the HEK3 locus (top) and RFN2 locus (bottom). 293T cells were transfected with a Cas9n-RT construct, a pegRNA construct, a nicking / mobilization sgRNA-MS2 construct, and an SSAP-MCP construct. [Figure 85-2] Same as above.
[0148] [Figure 86-1] Figure 86 shows different length edits mediated by SSAP+ prime editing at the HEK3 locus (top) and RFN2 locus (bottom). 293T cells were transfected with a Cas9n-RT construct, a pegRNA construct, a nicking / mobilization sgRNA-MS2 construct, and an SSAP-MCP construct. [Figure 86-2] Same as above.
[0149] [Figure 87]Figure 87 shows a schematic diagram of an editing system incorporating SSAP and retron. The retron-SSAP editor has three components: (1) the retron-sgRNA, which can be subdivided into three regions: the region of RNA that is reverse transcribed (termed "msd") and the region that remains as RNA in the final molecule (termed "msr"), and finally the guide RNA region (guide RNA with MS2 or other aptamers to recruit SSAP). The gRNA region can be derived from the Cas9 scaffold. This msr / msd RNA helps initiate the RT process, generating reverse-transcribed ssDNA directly linked to the sgRNA. (2) The retron's RT recognizes the retron RNA and completes reverse transcription of the donor template (the ligated RNA-DNA hybrid molecule). This RT can be fused to Cas9 or MCP-SSAP. (3) The SSAP protein fused to the MCP is optionally also fused to the RT if the RT is not fused to Cas9.
[0150] [Figure 88] Figures 88A-D show SSAP array screening showing cell viability versus editing efficiency (fold over negative control (Figure 88A, Figure 88C) or percent mKate knock-in (Figure 88B, Figure 88D)) for ACTB (Figure 88A, Figure 88B) and HSP90AA1 (Figure 88C, Figure 88D) targets. The positive control is EcRecT.
[0151] [Figure 89-1] Figures 89A-89C show normalized (Figure 89A) and absolute (Figure 89B) editing efficiencies comparing activity at two targets, HSP90AA and ACTB. Figure 89C shows cell viability using SSAP for HSP90AA1 knock-in compared to ACTB knock-in. The positive control is EcRecT. [Figure 89-2] Same as above.
[0152] [Figure 90-1]Figure 90 shows a scatter plot comparing cell viability versus normalized (A) or absolute (B) editing efficiency for all targets combined. Bar graphs compare the normalized (C) or absolute (D) editing efficiency for two targets, HSP90 and QCTB, for each candidate. The positive control is EcRecT. [Figure 90-2] Same as above.
[0153] [Figure 91] Figure 91 shows the tree and sequence alignment of SSAP_16 (1, SEQ ID NO: 185), SSAP_10 (2, SEQ ID NO: 179), SSAP_36 (3, SEQ ID NO: 205), SSAP_152 (4, SEQ ID NO: 321), and SSAP_184 (5, SEQ ID NO: 353) compared to EcRecT (SEQ ID NO: 171). See Table 12.
[0154] [Figure 92-1] Figure 92 shows the tree and sequence alignment of SSAP_16 (1, SEQ ID NO: 185), SSAP_10 (2, SEQ ID NO: 179), SSAP_36 (3, SEQ ID NO: 205), SSAP_152 (4, SEQ ID NO: 321), SSAP_184 (5, SEQ ID NO: 353), SSAP_197 (6, SEQ ID NO: 366), SSAP_305 (7, SEQ ID NO: 424), SSAP_210 (8; SEQ ID NO: 379), and SSAP_190 (9, SEQ ID NO: 359) compared to EcRecT (SEQ ID NO: 171). See Table 12. [Figure 92-2] Same as above.
[0155] [Figure 93] Figure 93 shows the tree and sequence alignment of SSAP_16 (1, SEQ ID NO: 185), SSAP_10 (2, SEQ ID NO: 179), SSAP_36 (3, SEQ ID NO: 205), SSAP_197 (6, SEQ ID NO: 366), and SSAP_210 (8; SEQ ID NO: 379) compared to EcRecT (SEQ ID NO: 171). See Table 12.
[0156] [Figure 94-1] Figure 94 shows the evolutionary tree of candidate SSAPs. A set of filters was applied to select 296 candidates that maximize evolutionary distance. The SSAPs cover diverse phylogenetic families (branches) within the SSAP family. [Figure 94-2] Same as above. [Figure 94-3] Same as above. [Figure 94-4] Same as above.
[0157] [Figure 95] Figures 95A-C show the editing efficiencies of the top 10 SSAPs compared to EcRecT and negative controls using the dCas9 editing system. (Figure 95A) mKate knock-in at the ACTB locus. (Figure 95B) mKate knock-in at the HSP90 locus. (Figure 95C) Scatter plot showing the editing efficiencies of candidate SSAPs at the ACTB and HSP90AA1 targets.
[0158] [Figure 96] Figure 96 compares the editing efficiencies of the top 10 SSAPs in a Cas-free system with a negative control (pA25 expressing MCP-EBFP) and pCK914 expressing MCP-EcRecT. Gene editing efficiencies are measured for knock-in of an 800-bp mKate donor with homology arms and a 20-nt guide RNA matching the target genomic insertion site, HSP90AA1 (top) or ACTB (bottom), in HEK293 human cells. The constructs used are: i) a guide RNA with an MS aptamer to recruit SSAP; ii) an MCP-SSAP fusion protein; and iii) donor DNA inserting the mKate cargo (without a promoter).
[0159] [Figure 97]Figures 97A-97C show the editing efficiency of top SSAPs compared to EcRecT using the dCas9 editing system with a transcribed AAV donor in primary hepatocytes (mouse). dCas9 is delivered virally separately using adenovirus-Cas9 (adeno-CMV-Cas9) under the control of a strong CMV promoter. (Figure 97A) AAV donor design: (top) typical AAV donor DNA; (bottom) AAV vector contains a promoter that transcribes the donor cargo into RNA. The donor RNA is transcribed in the antisense orientation to avoid cargo expression. A 600 bp luciferase cargo was knocked in at the mouse albumin locus (Figure 97B) or ACTB locus (Figure 97C).
[0160] [Figure 98] Figures 98A-98C show AAV donor designs and editing efficiencies. dCas9 is delivered virally separately using adenovirus-Cas9 (adeno-CMV-Cas9) under the control of a strong CMV promoter. (Figure 98A) Top: 5'-releasing AAV design. A second guide RNA is provided to bind / cleave the 5' end of the cargo (hsgRNA cleavage site adjacent to the left homology arm (left HA)). Middle: 3'-releasing AAV design. A second guide RNA is provided to bind / cleave the 3' end of the cargo (hsgRNA cleavage site adjacent to the right homology arm (right HA)). Bottom: Intact AAV design. A 600 bp luciferase cargo was knocked in at the mouse albumin locus (Figure 98B) or ACTB locus (Figure 98C).
[0161] [Figure 99] Figure 99 shows genome engineering across multiple human targets using SSAPs. The editing system included dCas9 (dSpCas9), guide RNA with the MS2 aptamer, MCP protein fused to a candidate SSAP, and donor DNA that inserted the mKate fluorescent protein cargo in-frame into the indicated endogenous genomic loci.
[0162] [Figure 100]Figures 100A-100C show a comparison of SSAP clones with RecT (LiRecT) from Listeria innoccua phage A118, and a model of the RecT complex during DNA annealing. (Figure 100A) Evolutionary tree showing clones SSAP-10, SSAP-16, SSAP-152, SSAP-198, EcRecT, and LiRecT. (Figure 100B) Model of the RecT-dsDNA complex. (Figure 100C) Model of RecT in complex with the duplex intermediate of DNA annealing. The model is based on the stacking structures 7UBB and 7UB2.
[0163] [Figure 101] Figures 101A-101B show the operation of RecT. (A) Model of RecT in complex with the duplex intermediate of DNA annealing. The inset shows that the RecT and DNA strands have extensive interactions (yellow dashed lines) with selected protein residues, including the core tyrosine amino acid highlighted in the upper right. (B) Model of RecT showing interactions with dsDNA (highlighted) consistent with the knock-in efficiency of N-terminally truncated mini-SSAP (mSSAP) (see, e.g., Figure 42). DETAILED DESCRIPTION OF THE INVENTION
[0164] Detailed Description of the Invention The present invention relates to a system and components for DNA editing.In particular, the disclosed system is based on CRISPR targeting and homology-directed repair by phage recombinase.This system provides excellent recombination efficiency and precision at kilobase scale.
[0165] To facilitate understanding of the present technology, several terms and phrases are defined below. Additional terms are described throughout the detailed description.
[0166] As used herein, the terms "comprise(s)," "include(s)," "having," "has," "can," "contain(s)," and variations thereof are intended to be open-ended transitional phrases, terms, or words that do not exclude additional acts or structures. The singular forms "a," "and," and "the" include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments that comprise, "consist of," and "consist essentially of" the embodiments or elements presented herein, whether explicitly stated or not.
[0167] For the recitation of numerical ranges herein, each intervening number is expressly contemplated to the same degree of precision. For example, in the range of 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and in the range of 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.
[0168] Unless otherwise defined herein, scientific and technical terms used in connection with the present invention shall have the meanings commonly understood by those skilled in the art. For example, any nomenclature used in connection with cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein, as well as the techniques thereof, are well known and commonly used in the art. The meaning and scope of terms should be clear; however, in the event of any potential ambiguity, the definitions provided herein shall prevail over dictionary or external definitions. Furthermore, unless otherwise required by context, singular terms shall include plurals, and plural terms shall include the singular.
[0169] The terms "complementary" and "complementarity" refer to the ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence, either through conventional Watson-Crick base pairing or other non-conventional types of pairing. The degree of complementarity between two nucleic acid sequences can be indicated by the percentage of nucleotides in a nucleic acid sequence that can form hydrogen bonds (e.g., Watson-Crick base pairing) with the second nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementary). Two nucleic acid sequences are "fully complementary" if all contiguous nucleotides of a nucleic acid sequence hydrogen bond with the same number of contiguous nucleotides in the second nucleic acid sequence. Two nucleic acid sequences are "substantially complementary" if the degree of complementarity between them is at least 60% (e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100%) over a region of at least 8 nucleotides (e.g., 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides), or if the two nucleic acid sequences form hybrids under at least moderate, and preferably high, stringency conditions. Exemplary moderate stringency conditions include overnight incubation at 37°C in a solution containing 20% formamide, 5x SSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5x Denhardt's solution, 10% dextran sulfate, and 20 mg / ml denatured, sheared salmon sperm DNA, followed by washing the filters at about 37-50°C in 1x SSC, or substantially similar conditions, e.g., the moderately stringent conditions described in Sambrook et al., infra.High stringency conditions include, for example, (1) using low ionic strength and high temperature, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate (SDS) at 50°C for washing, (2) using a denaturing agent during hybridization, such as formamide, e.g., 50% (v / v) formamide containing 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1% polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer at pH 6.5 with 750 mM sodium chloride and 75 mM sodium citrate at 42°C, or (3) 50% formamide, 5x SSC (0.75 M The hybridization conditions are: 50 mM NaCl, 0.075 M sodium citrate, 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5× Denhardt's solution, sonicated salmon sperm DNA (50 μg / ml), 0.1% SDS, and 10% dextran sulfate at 42° C., followed by washing at (i) 0.2× SSC at 42° C., (ii) 50% formamide at 55° C., and (iii) 0.1× SSC at 55° C. (preferably in combination with EDTA). Further details and explanations of the stringency of hybridization reactions are provided, for example, in Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Press, Cold Spring Harbor, NY (2001) and Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates and John Wiley & Sons, New York (1994).
[0170] A cell has been "genetically modified," "transformed," or "transfected" with exogenous DNA, such as a recombinant expression vector, when the exogenous DNA has been introduced inside the cell. The presence of the exogenous DNA results in a permanent or transient genetic change. The transforming DNA may or may not be integrated (covalently linked) into the genome of the cell. For example, in prokaryotes, yeast, and mammalian cells, the transforming DNA may be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming DNA has integrated into a chromosome so that it is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones comprising a population of daughter cells containing the transforming DNA. A "clone" is a population of cells arising from a single cell or common ancestor by mitosis. A "cell line" is a clone of a primary cell capable of stable growth in vitro for many generations.
[0171] As used herein, "nucleic acid" or "nucleic acid sequence" refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively. The present technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, as well as any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. The polymer or oligomer may be heterogeneous or homogeneous in composition and may be isolated from naturally occurring sources or artificially or synthetically produced. Furthermore, the nucleic acid may be DNA or RNA, or a mixture thereof, and may exist permanently or transitionally in single- or double-stranded form, including homoduplexes, heteroduplexes, and hybrid states. In some embodiments, the nucleic acid or nucleic acid sequence comprises other types of nucleic acid structures, such as, for example, a DNA / RNA helix, a peptide nucleic acid (PNA), a morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41(14):4503-4510 (2002) and U.S. Pat. No. 5,034,506, which are incorporated herein by reference), a locked nucleic acid (LNA; see Wahlestedt et al., Proc. Natl. Acad. Sci. USA, 97:5633-5638 (2000), which are incorporated herein by reference), a cyclohexenyl nucleic acid (see Wang, J. Am. Chem. Soc., 122:8595-8602 (2000), which are incorporated herein by reference), and / or a ribozyme. Thus, the term "nucleic acid" or "nucleic acid sequence" can also encompass strands that contain non-naturally occurring nucleotides, modified nucleotides and / or non-nucleotide building blocks (e.g., "nucleotide analogs"), which can exhibit the same function as natural nucleotides; furthermore, as used herein, the term "nucleic acid sequence" refers to oligonucleotides, nucleotides or polynucleotides and fragments or portions thereof, as well as DNA or RNA of genomic or synthetic origin, which can be single-stranded or double-stranded and represent the sense or antisense strand.The terms "nucleic acid," "polynucleotide," "nucleotide sequence," and "oligonucleotide" are used interchangeably. The terms "nucleic acid," "polynucleotide," "nucleotide sequence," and "oligonucleotide" refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof.
[0172] A "peptide" or "polypeptide" is a linked sequence of two or more amino acids linked by peptide bonds. A peptide or polypeptide can be natural, synthetic, or a modified or combination of natural and synthetic. Polypeptides include proteins such as binding proteins, receptors, and antibodies. Proteins can be modified by the addition of sugars, lipids, or other moieties not contained in the amino acid chain. The terms "polypeptide" and "protein" are used interchangeably herein.
[0173] As used herein, the term " percent sequence identity " refers to the percentage of nucleotides or nucleotide analogs in a nucleic acid sequence, or the percentage of amino acids in an amino acid sequence, that are identical with the corresponding nucleotides or amino acids in a reference sequence, after aligning the two sequences and, if necessary, introducing gaps to achieve maximum percent identity.Therefore, if the nucleic acid obtained by this technology is longer than the reference sequence, the additional nucleotides in the nucleic acid that are not aligned with the reference sequence are not considered for determining sequence identity.Methods and computer programs for alignment include BLAST, Align2 and FASTA, and are well known in the art.
[0174] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, to which another DNA segment, e.g., an "insert," can be attached or incorporated so as to bring about replication of the attached segment in a cell.
[0175] The term "wild-type" refers to a gene or gene product that has the characteristics of that gene or gene product when isolated from a naturally occurring source. A wild-type gene is that gene that is most frequently observed in a population and is therefore arbitrarily designated the "normal" or "wild-type" form of the gene. In contrast, the terms "modified," "mutant," or "polymorphic" refer to a gene or gene product that exhibits modified sequence and / or functional properties (e.g., altered properties) when compared to the wild-type gene or gene product. It should be noted that naturally occurring mutants can be isolated and are identified by the fact that they have altered properties when compared to the wild-type gene or gene product.
[0176] RNA-guided CRISPR recombineering system In bacteria and archaea, CRISPR / Cas systems provide immunity by integrating fragments of invading phage, viral, and plasmid DNA into CRISPR loci and directing the degradation of homologous sequences using the corresponding CRISPR RNA ("crRNA"). Each CRISPR locus encodes an acquired "spacer" that is separated by repeat sequences. Transcription of the CRISPR locus generates a "pre-crRNA," which is processed to generate a crRNA containing a spacer-repeat fragment that guides an effector nuclease complex to cleave dsDNA sequences complementary to the spacer. Three different types of CRISPR systems, type I, type II, or type III, are known and are classified based on the Cas protein type and the use of a protospacer adjacent motif (PAM) for selection of the protospacer in the invading DNA. The endogenous type II system contains four genes, including the Cas9 protein and the gene encoding it, and two non-coding crRNAs: a trans-activating crRNA (tracrRNA) and a precursor crRNA (pre-crRNA) array containing nuclease guide sequences (also called "spacers") spaced by identical direct repeats (DRs). The tracrRNA is important for pre-crRNA processing and Cas9 complex formation. First, the tracrRNA hybridizes to the repeat region of the pre-crRNA. Second, endogenous RNase III cleaves the hybridized crRNA-tracrRNA, and a second event removes the 5' end of each spacer to generate mature crRNAs that remain associated with both the tracrRNA and Cas9. Third, each mature complex locates the target double-stranded DNA (dsDNA) sequence and cleaves both strands using the nuclease activity of Cas9.
[0177] CRISPR / Cas gene editing systems have been developed to enable targeted modification of specific genes of interest in eukaryotic cells. CRISPR / Cas gene editing systems are generally based on the RNA-guided Cas9 nuclease derived from the type II prokaryotic clustered regularly interspaced short palindromic repeats (CRISPR) adaptive immune system. Engineering CRISPR / Cas systems for use in eukaryotic cells typically involves reconstitution of the crRNA-tracrRNA-Cas9 complex. In human cells, the Cas9 amino acid sequence can be codon-optimized and modified to include appropriate nuclear localization signals, for example, and the crRNA and tracrRNA sequences can be expressed individually or as a single chimeric molecule via an RNA polymerase II promoter. Typically, the crRNA and tracrRNA sequences are expressed as chimeras and are collectively referred to as "guide RNA" (gRNA) or single guide RNA (sgRNA). Thus, the terms "guide RNA," "single guide RNA," and "synthetic guide RNA" are used interchangeably herein to refer to a nucleic acid sequence comprising an array of tracrRNA and pre-crRNA containing a guide sequence. The terms "guide sequence," "guide," and "spacer" are used interchangeably herein to refer to an approximately 20-nucleotide sequence within a guide RNA that specifies a target site. In the CRISPR / Cas9 system, the guide RNA contains an approximately 20-nucleotide guide sequence followed by a protospacer adjacent motif (PAM), which directs Cas9 to the target sequence via Watson-Crick base pairing.
[0178] In some embodiments or the present invention, a system or composition for RNA-guided recombineering is provided that utilizes tools from the CRISPR gene editing system. The system includes a Cas protein, a nucleic acid molecule comprising a guide RNA sequence that is complementary to a target DNA sequence, and a recombination protein. In certain embodiments, the recombination protein includes a microbial recombination protein. In certain embodiments, the recombination protein includes a viral recombination protein. In certain embodiments, the recombination protein includes a eukaryotic recombination protein. In certain embodiments, the recombination protein includes a mitochondrial recombination protein.
[0179] The Cas protein family is described in further detail in, for example, Haft et al., PLoS Comput. Biol., 1(6):e60 (2005), incorporated herein by reference. The Cas protein can be any Cas endonuclease. In some embodiments, the Cas protein is Cas9 or Cas12a, otherwise referred to as Cpf1. In one embodiment, the Cas9 protein is a wild-type Cas9 protein. The Cas9 protein can be obtained from any suitable microorganism, and many bacteria express Cas9 protein orthologs or variants. In some embodiments, Cas9 is derived from Streptococcus pyogenes or Staphylococcus aureus. Cas9 proteins from other species are known in the art (see, e.g., U.S. Patent Application Publication No. 2017 / 0051312, incorporated herein by reference) and can be used in connection with the present invention. The amino acid sequences of Cas proteins from various species are publicly available through the GenBank and UniProt databases.
[0180] In some embodiments, the Cas9 protein is a Cas9 nickase (Cas9n). Wild-type Cas9 has two catalytic nuclease domains that facilitate double-stranded DNA cleavage. Cas9 nickase proteins are typically engineered by using the remaining active nuclease domain to nick Cas9 or by inactivating a point mutation(s) in one of the catalytic nuclease domains that enzymatically cleaves only one of the two DNA strands. Cas9 nickases are known in the art (see, for example, U.S. Patent Application Publication No. 2017 / 0051312, incorporated herein by reference), including, for example, Streptococcus pyogenes with point mutations at D10 or H840. In selected embodiments, the Cas9 nickase is Streptococcus pyogenes Cas9n (D10A).
[0181] In some embodiments, the Cas protein is a catalytically inactive Cas. For example, catalytically inactive Cas9 is essentially a DNA-binding protein, typically with two or more mutations in its catalytic nuclease domain, resulting in the protein having little or no catalytic nuclease activity. Streptococcus pyogenes Cas9 can be rendered catalytically inactive by mutation of D10 and at least one of E762, H840, N854, N863, or D986, typically H840 and / or N863 (see, e.g., U.S. Patent Application Publication No. 2017 / 0051312, incorporated herein by reference). Mutations in corresponding orthologs, such as N580 in Staphylococcus aureus Cas9, are known. In many cases, such mutations result in catalytically inactive Cas proteins having 3% or less of their normal nuclease activity.
[0182] In certain embodiments, the system includes a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence, which specifies a target site with an approximately 20-nucleotide guide sequence, as described above, followed by a protospacer adjacent motif (PAM) that directs Cas9 to the target sequence via Watson-Crick base pairing.
[0183] In certain embodiments, the system includes a nucleic acid molecule comprising an inactivated guide RNA (dgRNA) sequence complementary to the target DNA sequence. The inactivated guide is shortened or modified so that a CRISPR complex containing the dgRNA binds to the target DNA but does not cleave or nick it. Non-limiting examples include guides such as those described in International Publication No. 2016 / 094872, where the guide is modified to allow the formation of a CRISPR complex and successful binding to the target while not allowing successful nuclease activity (e.g., no nuclease activity / no indel activity). The guide nucleic acid can be considered catalytically inactive or conformationally inactive with respect to nuclease activity. A dgRNA with a short target recognition sequence can dramatically improve Cas9-mediated editing specificity by binding to the active Cas9 sgRNA complex and shielding it from off-target sites. (Rose et al., Suppression of unwanted CRISPR-Cas9 editing by co-administration of catalytically inactivating truncated guide RNAs. Nature Communications (2020) vol. 11, article 2697). Shortened / modified dgRNAs are used in accordance with the present invention to recruit Cas9-SSAP for cleavage-free knock-in of long sequences.
[0184] The terms "target DNA sequence," "target nucleic acid," "target sequence," and "target site" are used interchangeably herein to refer to a polynucleotide (such as a nucleic acid, gene, chromosome, or genome) to which a guide sequence (e.g., a guide RNA) is designed to have complementarity; hybridization between the target sequence and the guide sequence promotes the formation of a Cas9 / CRISPR complex, provided sufficient conditions for binding exist. In some embodiments, the target sequence is a genomic DNA sequence. The term "genomic," as used herein, refers to a nucleic acid sequence (e.g., a gene or locus) located on a chromosome within a cell. The target sequence and the guide sequence need not exhibit perfect complementarity, provided sufficient complementarity exists to cause hybridization and promote the formation of a CRISPR complex. The target sequence may comprise any polynucleotide, such as DNA or RNA. Suitable DNA / RNA binding conditions include physiological conditions normally present in a cell. Other suitable DNA / RNA binding conditions (e.g., conditions in a cell-free system) are known in the art; see, e.g., Sambrook, referenced herein and incorporated by reference. The strand of target DNA that is complementary to and hybridizes with a DNA-targeting RNA is called the "complementary strand," and the strand of target DNA that is complementary to the "complementary strand" (and therefore not complementary to the DNA-targeting RNA) is called the "noncomplementary strand" or "non-complementary strand."
[0185] The target genomic DNA sequence may encode a gene product. As used herein, the term "gene product" refers to any biochemical product resulting from the expression of a gene. A gene product may be RNA or protein. RNA gene products include non-coding RNAs such as tRNA, rRNA, microRNA (miRNA), and small interfering RNA (siRNA), and coding RNAs such as messenger RNA (mRNA). In some embodiments, the target genomic DNA sequence encodes a protein or polypeptide.
[0186] In some embodiments, for example, when the system includes a Cas9 nickase or a catalytically inactive Cas9, two nucleic acid molecules containing guide RNA sequences can be utilized. The two nucleic acid molecules can have the same or different guide RNA sequences and thus can be complementary to the same or different target DNA sequences. In some embodiments, the guide RNA sequences of the two nucleic acid molecules are complementary to the target DNA sequence at opposite ends (e.g., 3' or 5') of the insertion position and / or on opposite strands.
[0187] In some embodiments, the system further comprises a recruitment system comprising at least one aptamer sequence and an aptamer-binding protein operably linked to the microbial recombinant protein as part of a fusion protein.
[0188] In some embodiments, the aptamer sequence is an RNA aptamer sequence. In some embodiments, the nucleic acid molecule containing the guide RNA also contains one or more RNA aptamers, or distinct RNA secondary structures or sequences, that can recruit and bind to another molecular species, an adapter molecule, such as a nucleic acid or a protein. Some CRISPR systems are compatible with guide RNA insertion and extension, including, but not limited to, SpCas9, SaCas9, and LbCas12a (also known as Cpf1). The RNA aptamer can be a natural or synthetic oligonucleotide engineered by repeated rounds of in vitro selection or SELEX (systematic evolution of ligands by exponential enrichment) to bind to a specific target molecular species. In some embodiments, the nucleic acid contains two or more aptamer sequences. The aptamer sequences can be the same or different and can target the same or different adapter proteins. In selected embodiments, the nucleic acid contains two aptamer sequences.
[0189] Any known RNA aptamer / aptamer-binding protein pair can be selected and used in connection with the present invention (see, e.g., Jayasena, S.D., Clinical Chemistry, 1999. 45(9): p. 1628-1650; Gelinas et al., Current Opinion in Structural Biology, 2016. 36: p. 122-132; and Hasegawa, H., Molecules, 2016; 21(4): p. 421, which are incorporated herein by reference).
[0190] There are many RNA aptamer binding or adaptor proteins, including a diverse series of bacteriophage coat proteins. Examples of such coat proteins include, but are not limited to, MS2, Qβ, F2, GA, fr, JP501, M12, R17, BZ13, JP34, JP500, KU1, M11, MX1, TW18, VK, SP, FI, ID2, NL95, TW19, AP205, φCb5, φCb8r, φCb12r, φCb23r, 7s and PRR1. In some embodiments, the RNA aptamer binds to MS2 bacteriophage coat protein or its functional derivative, fragment or variant. MS2-binding RNA aptamers generally have a simple stem-loop structure and are classically defined as 19-nucleotide RNA molecules with a single bulged adenine at the 5' end of the stem (Witherall GW et al., (1991) Prog. Nucleic Acid Res. Mol. Biol., 40, 185-220, incorporated herein by reference). However, many widely different primary sequences have been found to be capable of binding to MS2 coat proteins (Parrott AM et al., Nucleic Acids Res. 2000; 28(2): 489-497; Buenrostro JD et al., Natura Biotechnology 2014; 32, 562-568, incorporated herein by reference). Any of the RNA aptamer sequences known to bind to MS2 bacteriophage coat proteins can be utilized in the context of the present invention to bind to fusion proteins containing MS2. In selected embodiments, the MS2 RNA aptamer sequence comprises AACAUGAGGAUCACCCAUGUCUGCAG (SEQ ID NO: 145), AGCAUGAGGAUCACCCAUGUCUGCAG (SEQ ID NO: 146) or AGCGUGAGGAUCACCCAUGCCUGCAG (SEQ ID NO: 147).
[0191] Bacteriophage N-proteins (Nut utilization site proteins) contain a conserved, approximately 20-amino acid, arginine-rich RNA recognition motif called the N-peptide. RNA aptamers can bind to phage N-peptides or their functional derivatives, fragments, or variants. In some embodiments, the phage N-peptide is a lambda or P22 phage N-peptide or a functional derivative, fragment, or variant thereof.
[0192] In selected embodiments, the N peptide is the lambda phage N22 peptide or a functional derivative, fragment, or variant thereof. In some embodiments, the N22 peptide comprises an amino acid sequence having at least 70% similarity to the amino acid sequence GNARTRRRERRAEKQAQWKAAN (SEQ ID NO: 149). The N22 peptide, the 22-amino acid RNA-binding domain of the λ bacteriophage antiterminator protein N (λN-(1-22) or λN peptide), can specifically bind to specific stem-loop structures, including but not limited to, BoxB stem-loops. See, e.g., Cilley and Williamson, RNA 1997;3(1):57-67, incorporated herein by reference. Several different BoxB stem-loop primary sequences are known to bind to the N22 peptide, any of which may be utilized in connection with the present invention. In some embodiments, the N22 peptide RNA aptamer sequence comprises a nucleotide sequence having at least 70% similarity to an RNA sequence selected from the group consisting of GCCCUGAAAAAGGGC (SEQ ID NO: 150), GCCCUGAAGAAGGGC (SEQ ID NO: 151), GCGCUGAAAAAGCGC (SEQ ID NO: 152), GCCCUGACAAAGGGC (SEQ ID NO: 153), and GCGCUGACAAAGCGC (SEQ ID NO: 154). In some embodiments, the N22 peptide RNA aptamer sequence is selected from the group consisting of SEQ ID NOs: 150-154.
[0193] In selected embodiments, the N-peptide is a P22 phage N-peptide or a functional derivative, fragment, or variant thereof. Several different BoxB stem-loop primary sequences are known to bind to P22 phage N-peptides and their variants, any of which may be utilized in connection with the present invention. See, e.g., Cocozaki, Ghattas, and Smith, Journal of Bacteriology 2008;190(23):7699-7708, incorporated herein by reference. In some embodiments, the P22 phage N-peptide comprises an amino acid sequence having at least 70% similarity to the amino acid sequence GNAKTRRHERRRKLAIERDTI (SEQ ID NO: 155). In some embodiments, the P22 phage N-peptide RNA aptamer sequence comprises a sequence having at least 70% similarity to an RNA sequence selected from the group consisting of GCGCUGACAAAGCGC (SEQ ID NO: 156) and CCGCCGACAACGCGG (SEQ ID NO: 157). In some embodiments, the P22 phage N-peptide RNA aptamer sequence is selected from the group consisting of SEQ ID NOs: 156-157, UGCGCUGACAAAGCGCG (SEQ ID NO: 158), or ACCGCCGACAACGCGGU (SEQ ID NO: 159).
[0194] In certain embodiments, different aptamer / aptamer-binding protein pairs can be selected to result in a combination of recombinant proteins and functions.
[0195] In some embodiments, the aptamer sequence is a peptide aptamer sequence. Peptide aptamers can be natural or synthetic peptides that are specifically recognized by affinity agents. Such aptamers include, but are not limited to, c-Myc affinity tag, HA affinity tag, His affinity tag, S affinity tag, methionine-His affinity tag, RGD-His affinity tag, 7xHis tag, FLAG octapeptide, strep tag or strep tag II, V5 tag or VSV-G epitope. Corresponding aptamer-binding proteins are well known in the art, including, for example, primary antibodies, biotin, affimers, single-domain antibodies and antibody mimics.
[0196] Exemplary peptide aptamers include GCN4 peptides (Tanenbaum et al., Cell, 2014; 159(3):635-646, incorporated herein by reference). Antibodies or GCN4 binding proteins can be used as aptamer binding proteins.
[0197] In some embodiments, the peptide aptamer sequence is conjugated to a Cas protein. The peptide aptamer sequence can be fused to Cas in any orientation (e.g., N-terminal to C-terminal, C-terminal to N-terminal, N-terminal to N-terminal). In selected embodiments, the peptide aptamer is fused to the C-terminus of the Cas protein.
[0198] In some embodiments, 1 to 24 peptide aptamer sequences can be conjugated to the Cas protein. The aptamer sequences can be the same or different and can target the same or different aptamer-binding proteins. In select embodiments, 1 to 24 tandem repeats of the same peptide aptamer sequence are conjugated to the Cas protein. In preferred embodiments, 4 to 18 tandem repeats are conjugated to the Cas protein. Individual aptamers can be separated by a linker region. Suitable linker regions are known in the art. The linker can be flexible or configured to allow binding of an affinity agent to adjacent aptamers with no or low steric hindrance. The linker sequence can provide an unstructured or linear region of the polypeptide, for example, containing one or more glycine and / or serine residues. The linker sequence can be at least about 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids in length.
[0199] In some embodiments, the fusion protein comprises a recombinant protein operably linked to an aptamer-binding protein. In some embodiments, the recombinant protein comprises a microbial recombinant protein. In some embodiments, the recombinant protein comprises a recombinase. In certain embodiments, the recombinant protein comprises 5'-3' exonuclease activity. In certain embodiments, the recombinant protein comprises 3'-5' exonuclease activity. In certain embodiments, the recombinant protein comprises ssDNA binding activity. In certain embodiments, the recombinant protein comprises ssDNA annealing activity.
[0200] The bacteriophage λ-encoded genetic recombination machinery, referred to as the λ red system, contains the exo and bet genes, assisted by the gam gene, which together are referred to as the λ red genes. Exo is a 5'-3' exonuclease that targets dsDNA, and Bet is an ssDNA-binding protein. Bet functions include protecting ssDNA from degradation and promoting the annealing of complementary ssDNA strands. Another bacteriophage system found in E. coli is the Rac prophage system, which contains the recE and recT genes, which are functionally similar to exo and bet. In some embodiments, the microbial recombination proteins can be RecE, RecT, lambda exonuclease (Exo), Bet proteins (betA, redB), exonuclease gp6, single-stranded DNA-binding protein gp2.5, or derivatives or variants thereof.
[0201] Recombination proteins and functional fragments thereof useful in the present invention include nucleases, ssDNA binding proteins (SSBs), and ssDNA annealing proteins (SSABs). Among microbial proteins, these include, but are not limited to, E. coli proteins such as ExoI (xonA; sbcB), ExoIII (xthA), ExoIV (orn), ExoVII (xseA, xseB), ExoIX (ygdG), ExoX (exoX), DNA polI 5'Exo (ExoVI) (polA), DNA Pol I 3'Exo (ExoII) (polA), DNA Pol II 3'Exo (polB), DNA Pol III 3'Exo (dnaQ, mutD), RecBCD (recB, recC, recD), and RecJ (recJ), and functional fragments thereof.
[0202] Double-stranded DNA contains genetic information, but the use of this information involves a single-stranded intermediate. While single-stranded intermediates form secondary structures and are susceptible to chemical and nucleolytic degradation, cells encode ssDNA-binding proteins (SSBs) that bind to and stabilize ssDNA. Useful SSBs include, but are not limited to, prokaryotic, bacteriophage, eukaryotic, mammalian, mitochondrial, and viral SSBs. While SSBs are found in all organisms, the proteins themselves share surprisingly little sequence similarity and can differ in subunit composition and oligomerization state. SSB proteins can contain specific structural features. One is the use of an oligonucleotide / oligosaccharide-binding (OB) domain to bind ssDNA through a combination of electrostatic and base-stacking interactions with the phosphodiester backbone and nucleotide bases. Another feature is oligomerization, which brings together the DNA-binding OB fold. Eukaryotic SSBs are regulated by phosphorylation on serine and threonine residues. Tyrosine phosphorylation of microbial SSB has been observed in taxonomically distant bacteria and substantially increases its affinity for ssDNA. Human mitochondrial ssDNA-binding protein is structurally similar to Escherichia coli-derived SSB (EcoSSB) but lacks the C-terminal disordered domain. Eukaryotic replication protein A (RPA) shares function but not sequence homology with bacterial SSB. Herpes simplex virus (HSV-1) SSB, ICP8, is a nuclear protein required for viral DNA replication along with other replication proteins.
[0203] Without being bound by theory, it is believed that the exonuclease activity and ssDNA binding activity of the recombination protein of the present invention reveal and protect the single-stranded regions of template and target DNA, thereby promoting recombination.In addition, targeting can be cooperative, involving targeted CRISPR-mediated nicking of chromosomal DNA in concert with recombination by homologous arms designed in template DNA.In certain embodiments of the present invention, off-target effects are minimized.For example, targeted recombination involves cooperative CRISPR and recombination functions, but at off-target sites, there is no homology with HR template DNA, and nick repair can be promoted.
[0204] Single-stranded DNA annealing proteins (SSAPs) are also ubiquitous among organisms with diverse sequences and have been classified into families and superfamilies through bioinformatics and experimental analysis. Furthermore, phage-encoded SSAPs are recognized to encode their own SSAP recombinases, which replace the classical RecA protein while functioning in conjunction with host proteins to control DNA metabolism. Steczkiewicz classified SSAPs into seven families (RecA, Gp2.5, RecT / Redβ, Erf, Rad52 / 22, Sak3, and Sak4) organized into three superfamilies, including prokaryotes, eukaryotes, and phages (Steczkiewicz et al., 2021, Front. Microbiol. 12:644-622). Non-limiting examples of SSAPs that can be used in accordance with the present invention are listed in Table 7. Any one or more of the SSAPs can be used in the present invention.
[0205] In certain embodiments, the microbial recombinant protein is RecE or RecT, or a derivative or variant thereof. Derivatives or variants of RecE and RecT are functionally equivalent proteins or polypeptides that have substantially similar functions to wild-type RecE and RecT. Derivatives or variants of RecE and RecT include biologically active amino acid sequences that are similar to the wild-type sequences but differ due to amino acid substitutions, additions, deletions, truncations, post-translational modifications, or other modifications. In some embodiments, derivatives may improve translation, purification, biological half-life, activity, or eliminate or reduce any undesirable side effects or reactions. Derivatives or variants may be natural polypeptides, synthetic or chemically synthesized polypeptides, or genetically engineered peptide polypeptides. RecE and RecT biological activities are known and readily assayed by those skilled in the art and include, for example, exonuclease activity and single-stranded nucleic acid binding, respectively.
[0206] RecE or RecT can be derived from several organisms, including Escherichia coli, Pantoea breeneri, the F-type symbiont of Plautia stali, Providencia sp., MGF014, Shigella sonnei, and Pseudobacteriovorax antillogorgiicola, among others. Other non-limiting sources include Desulfotalea psychrophila, Lactococcus lactis, Flavobacterium psychrophilum, Mycobacterium smegmatis, Lactobacillus rhamnosus, Psychrobacter arcticus, Psychrobacter cryohalolentis, Psychromonas ingrahamii, Photobacterium profundum, Psychroflexus torquis, and Caulobacter crescentus. In certain embodiments, the RecE and RecT proteins are derived from Escherichia coli.
[0207] In some embodiments, the fusion protein comprises RecE or a derivative or variant thereof. RecE or a derivative or variant thereof may comprise an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-8. RecE or a derivative or variant thereof may comprise an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) similarity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-8. In select embodiments, RecE or a derivative or variant thereof comprises an amino acid sequence having at least 90% similarity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-8. In exemplary embodiments, RecE or a derivative or variant thereof comprises an amino acid sequence having at least 90% similarity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-3.
[0208] In some embodiments, the fusion protein comprises RecT or a derivative or variant thereof. RecT or a derivative or variant thereof may comprise an amino acid sequence selected from the group consisting of SEQ ID NOs:9-14. RecT or a derivative or variant thereof may comprise an amino acid sequence having at least 70% (e.g., 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%) similarity to an amino acid sequence selected from the group consisting of SEQ ID NOs:9-14. In select embodiments, RecT or a derivative or variant thereof comprises an amino acid sequence having at least 90% similarity to an amino acid sequence selected from the group consisting of SEQ ID NOs:9-14. In exemplary embodiments, RecT or a derivative or variant thereof comprises an amino acid sequence having at least 90% similarity to an amino acid sequence selected from the group consisting of SEQ ID NO:9.
[0209] In certain embodiments, the fusion protein comprises a recombinant protein comprising an amino acid sequence at least 75% similar or at least 75% identical to a recombinant protein of SEQ ID NO:166-491, a recombinant protein of Table 9, a recombinant protein of SEQ ID NO:179, SEQ ID NO:185, SEQ ID NO:205, SEQ ID NO:321, SEQ ID NO:353, SEQ ID NO:359, SEQ ID NO:366, SEQ ID NO:424, or SEQ ID NO:479, or a recombinant protein of SEQ ID NO:166, SEQ ID NO:168, SEQ ID NO:169, SEQ ID NO:170, SEQ ID NO:171, SEQ ID NO:241, SEQ ID NO:253, SEQ ID NO:290, SEQ ID NO:408, SEQ ID NO:411, or SEQ ID NO:442. In certain embodiments, the fusion protein comprises a recombinant protein comprising a sequence having at least 80%, at least 85%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, or 100% similarity or identity to the recombinant protein referenced above. Truncations can be from either the C-terminus or the N-terminus, or both. For example, as shown in Example 6 below, a set of diverse truncations from either terminus or both provided functional products. In some embodiments, one or more (2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120 or more) amino acids may be truncated from the C-terminus, N-terminus when compared to the wild-type sequence.
[0210] In some embodiments, the recombination protein comprises a tyrosine recombinase or a functional fragment thereof. In some embodiments, the recombination protein comprises a serine recombinase or a functional fragment thereof. In some embodiments, the recombination protein comprises an integrase, a resolvase, or an invertase, or a functional fragment thereof. In some embodiments, the recombinase protein comprises a site-specific recombinase protein or a functional fragment thereof. In some embodiments, the recombination protein comprises an exonuclease or a functional fragment thereof. In some embodiments, the recombination protein comprises a ssDNA binding protein or a functional fragment thereof. In certain embodiments, the fusion protein includes, but is not limited to, Hin, Gin, Tn3, β / six, CinH, Min, ParA, γδ, Bxb1, φC31, TP901-1, TGI, Wβ, φ370.1, φK38, φBT1, R4, φRV1, φFC1, MR11, A118, U153, Bxz2, gp29, Cre, Dre, Vika, Flp, Kw, SprA, HK022, P22, L1, or L5, or a homolog or functional fragment thereof of any such protein. Such recombinases may be classified in the art as integrases, resolvases, or invertases, may share partial structure and activity with exonucleases and SSBs, and may be used in accordance with the present invention.
[0211] The present invention provides a system comprising a reverse transcriptase, a guide nucleic acid, and a recombination protein, and optionally a Cas protein.
[0212] The term "reverse transcriptase" refers to a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require a primer to synthesize DNA transcripts from an RNA template. Historically, reverse transcriptases have primarily been used to transcribe mRNA into cDNA, which can then be cloned into vectors for further manipulation. Avian myeloblastosis virus (AMV) reverse transcriptase was the first widely used RNA-dependent DNA polymerase (Verma, Biochim. Biophys. Acta 473:1 (1977)). The enzyme possesses 5'-3' RNA-directed DNA polymerase activity, 5'-3' DNA-directed DNA polymerase activity, and RNase H activity. RNase H is a processive 5' and 3' ribonuclease specific for the RNA strand of an RNA-DNA hybrid (Perbal, A Practical Guide to Molecular Cloning, New York: Wiley & Sons (1984)). Known viral reverse transcriptases lack the 3'-5' exonuclease activity required for proofreading, so transcription errors cannot be corrected by the reverse transcriptase (Saunders and Saunders, Microbial Genetics Applied to Biotechnology, London: Croom Helm (1987)). A detailed study of the activity of AMV reverse transcriptase and its associated RNase H activity was presented by Berger et al., Biochemistry 22:2365-2372 (1983). Another reverse transcriptase widely used in molecular biology is the reverse transcriptase derived from Moloney murine leukemia virus (M-MLV). See, for example, Gerard, GR, DNA 5:271-279 (1986) and Kotewicz, ML et al., Gene 35:249-258 (1985). An M-MLV reverse transcriptase substantially lacking RNase H activity has also been described. See, for example, US Patent No. 5,244, 797. The present invention contemplates the use of any such reverse transcriptase, or variant or mutant (RT) thereof.With respect to RT and linkers or methods for operably linking components of embodiments of the invention, such as RT systems or compositions of the invention (as well as linkers or methods for operably linking components of systems or compositions discussed herein that do not include RT), including what is known as prime editing and twin prime editing, see WO 2020 / 191241, WO 2020 / 191153, WO 2020 / 191248, WO 2020 / 191153, WO 2020 / 191249 ... Reference is made to International Publication Nos. 20 / 191245, 2020 / 191243, 2020 / 191233, 2020 / 191246, 2020 / 191249, 2020 / 191239, 2020 / 191234, 2020 / 191242, 2020 / 191248, 2020191171, and 2021 / 226558. Each of International Publication No. 2020 / 191241, International Publication No. 2020 / 191153, International Publication No. 2020 / 191245, International Publication No. 2020 / 191243, International Publication No. 2020 / 191233, International Publication No. 2020 / 191246, International Publication No. 2020 / 191249, International Publication No. 2020 / 191239, International Publication No. 2020 / 191234, International Publication No. 2020 / 191242, International Publication No. 2020 / 191248, International Publication No. 2020191171, and International Publication No. 2021 / 226558 is incorporated herein by reference. The RTs of International Publication No. 2020 / 191241, International Publication No. 2020 / 191153, International Publication No. 2020 / 191245, International Publication No. 2020 / 191243, International Publication No. 2020 / 191233, International Publication No. 2020 / 191246, International Publication No. 2020 / 191249, International Publication No. 2020 / 191239, International Publication No. 2020 / 191234, International Publication No. 2020 / 191242, International Publication No. 2020 / 191248, International Publication No. 2020191171, and International Publication No. 2021 / 226558 can be used in the practice of the present invention.The linkers or operably linking methods of WO 2020 / 191241, WO 2020 / 191153, WO 2020 / 191245, WO 2020 / 191243, WO 2020 / 191233, WO 2020 / 191246, WO 2020 / 191249, WO 2020 / 191239, WO 2020 / 191234, WO 2020 / 191242, WO 2020 / 191248, WO 2020191171, and WO 2021 / 226558 can be used in the practice of the invention. Reference is also made to U.S. Patent Application Publication Nos. 2014 / 0349400 and 2018 / 0298391 (both of which are incorporated herein by reference) which contain systems containing Cas9, reverse transcriptase, guide RNA and RNA for reverse transcriptase activity, and the reverse transcriptase and other aspects of these earlier systems can be used in the practice of the present invention.
[0213] WO 2020 / 191153 describes a system comprising a CRISPR protein (e.g., Cas9 nickase) and reverse transcriptase for use with a guide RNA that specifies a target site, and template synthesis of a desired edit in the form of a replacement DNA strand by an extension (either DNA or RNA) engineered onto the guide nucleic acid (e.g., at the 5' or 3' end, or an internal portion of the guide RNA). Through DNA repair and / or replication mechanisms, the endogenous strand at the target site is replaced by the newly synthesized replacement strand containing the desired edit. The invention provides single-stranded binding proteins (e.g., SSAP or SSB) used with reverse transcriptase to edit without CRISPR-mediated nicking or cleavage of the target DNA.
[0214] Background on RT Systems and Compositions: Current genome editing technologies are limited by their low efficiency and precision for precise editing, and the ability to introduce precise substitutions, deletions, or insertions into mammalian cells using current tools such as the CRISPR system is highly unreliable. The typical process involves the delivery of a gene editing tool (such as CRISPR) and a DNA repair template to introduce the desired changes into the genomic sequence. However, DNA delivered to cells can be nonspecifically inserted into off-target genomic loci or unintended targets, leading to major challenges in ensuring safe and precise gene editing for therapeutic purposes.
[0215] Description of RT Systems and Compositions: Here, Applicants describe the invention using RNA as a molecular entity to mediate gene editing. Applicants have designed and validated components of systems and methods for applying RNA as a template (donor) to insert, delete, replace, or control genomic DNA sequences mediated through the activity of SSAP (single-stranded annealing protein exemplified by RecT, Lambda Red, T7gp2.5).
[0216] In a first embodiment, Applicants demonstrate the efficiency of gene editing by a process that delivers three components to cells: (1) Applicants use a CRISPR system consisting of CRISPR enzymes (corresponding to Cas9 / Cas9n / dCas9 or Cas12a / nCas12a / dCas12a for cleavage / nicking / R-loop formation, respectively) and guide RNAs to introduce localized DNA cleavage, nicking, or R-loop formation, where the guide RNAs contain aptamers (e.g., MS2, PP7, or BoxB) for recruiting SSAP proteins, and (2) an RNA sequence carrying the desired DNA change has one or more homology arm (HA) regions(s) fused / linked to the guide RNA of (1) or fused / linked to a second guide RNA. The HA regions are at least 20 bp long and provide homologous regions adjacent to the editing site for SSAP-mediated editing. When a second guide RNA is used, this second guide RNA binds to a nearby genomic site located 0-150 bp away from the first guide RNA. This second guide RNA then forms a complex with a CRISPR enzyme (such as Cas9 / nCas9 / dCas9 and Cas12a / nCas12a / dCas12a) and is recruited to the target genomic locus, serving to provide an RNA template / donor for editing. The enzyme can be either a regular CRISPR enzyme or a Cas protein, but it can also be a nicking or inactivating CRISPR enzyme (such as dCas9 or dCas12a) that binds only to the target locus. The guide can be a regular guide RNA or a shorter guide RNA (typically 2-6 bp shorter than the regular guide RNA, e.g., 14-18 bp) to enable efficient binding rather than target cleavage. (3) An SSAP protein fused to an RNA-aptamer-binding protein (RBP) via a linker. The RBP is the MS2 coat protein (MCP), the PP7 coat protein (PCP), or the BoxB-binding peptide from lambda phage (lambda N22 peptide).For this component, Applicants also identified an additional factor that enhances this RNA-templated SSAP gene editing; namely, when Applicants fused reverse transcriptase (RT) to the SSAP protein via a long peptide linker, creating this third component, RBP-SSAP-RT or RBP-RT-SSAP (- represents the linker), this further enhanced editing efficiency.
[0217] In a second embodiment, the Cas9 / nCas9 / dCas9 or Cas12a / nCas12a / dCas12a protein is fused to a reverse transcriptase (RT) via a linker. The guide RNA of this design also has a primer binding site (PBS) of at least 14 bp or more that is complementary to the region of the editing site. This PBS helps initiate RT activity. Alternatively, another design uses the same guide RNA as in the first embodiment to initiate RT activity and supply cells with a short oligo DNA (14 bp or more in length) that is complementary to the region of the editing site. This oligo DNA initiates RT activity and enables SSAP-mediated gene editing.
[0218] In a third embodiment, Cas9 / nCas9 / dCas9 or Cas12a / nCas12a / dCas12a proteins are fused to a reverse transcriptase (RT) from a retron system via a linker. This design of guide RNA also contains the msr / msd sequence from the retron and one or more homology arm (HA) regions complementary to the region of the editing site. The msr / msd sequence serves to initiate RT activity. The HA region serves to mediate SSAP gene editing.
[0219] Overall, this set of tools and methods provides a novel and innovative RNA-mediated / RNA-templated gene editing in eukaryotic / mammalian cells. Applicants further demonstrate that by designing cleavable RNA templates using endogenous tRNAs, ribozymes, or direct repeats from the Cas12a system, they also achieve multiplexed target gene editing using RNA as a template.
[0220] Features Related to the RT System and Composition: Applicant's RNA-templated SSAP gene editing system has five advantages: (1) reduced off-target or toxicity due to RNA, less immunogenic compared to DNA used in existing gene editing processes, and inability to directly integrate RNA into unintended genomic or off-target DNA sites; (2) Applicants easily multiplex precise gene editing methods by using a cleavable RNA template in their method; (3) RNA is easier to deliver into cells, easier to manufacture, and cheaper to scale up for clinical use; (4) RNA has many possibilities for manipulation by combining other regulatory or combinatorial payloads / components via chemical or biochemical coupling, allowing for more efficient delivery, editing, or synergism between RNA-templated gene editing and other types of gene editing or therapeutic modalities; and (5) the efficiency of RNA-templated gene editing can be enhanced via RNA and protein factors, making it orthogonal to normal DNA repair pathways that may be important for the health of target cells.
[0221] RNA-guided recombination protein systems. Certain embodiments or the present invention provide systems or compositions for RNA-guided recombineering that do not primarily rely on CRISPR proteins. In such embodiments, the system or composition comprises a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence and a recombination protein. In certain embodiments, the system or composition can promote R-loop formation. In certain embodiments, the system or composition is capable of recombination. In certain embodiments, the system or composition does not comprise a CRISPR protein. In certain embodiments, the recombination protein comprises a microbial recombination protein. In certain embodiments, the recombination protein comprises a viral recombination protein. In certain embodiments, the recombination protein comprises a eukaryotic recombination protein. In certain embodiments, the recombination protein comprises a mitochondrial recombination protein. In various embodiments, the recombination protein comprises a single-stranded DNA annealing protein (SSAP), a single-stranded DNA binding protein (SSB), an exonuclease, or a combination of two or more thereof. In certain embodiments, the system or composition does not comprise Cas9. In certain embodiments, the system or composition does not comprise Cas12a. In certain embodiments, the system or composition does not comprise Cas. In certain embodiments, the system or composition does not comprise CRISPR.
[0222] Without being bound by theory, the system can be thought of as comprising a guide nucleic acid that binds to the target DNA, thereby promoting R-loop formation, and a recombination protein that promotes recombination between the target nucleic acid and the donor nucleic acid.
[0223] In certain embodiments, the guide RNA and the recombinant protein are operatively linked. In some embodiments, the linkage is covalent. In some embodiments, the linkage is non-covalent. In certain embodiments, the guide nucleic acid comprises an aptamer sequence, and the recombinant protein comprises or is bound to an aptamer-binding domain. The following table provides non-limiting examples of R-loop guide nucleic acids for use in the present invention. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4]
[0224] R-loop-guide RNAs comprise guide and scaffold components in various configurations, such as guide-scaffold and scaffold-guide configurations. In certain embodiments, R-loop-guide RNAs comprise a guide at the 5' end of the scaffold. In certain embodiments, R-loop-guide RNAs comprise a guide at the 3' end of the scaffold.
[0225] The guide sequence is engineered to bind to a target DNA (genomic target). In certain embodiments, the guide is 17-160 bases long. The scaffold comprises one or more aptamer sequences. Aptamers for use in the present invention include, but are not limited to, MS2, PP7, BoxB, and the like. In certain embodiments, the fusion protein comprises an RNA-binding component that binds to such an aptamer and an SSAP protein, such as, but not limited to, RecT, Lambda Red, or T7gp2.5.
[0226] The donor nucleic acid may be single-stranded or double-stranded DNA and may contain (1) homology arms (HA) of various lengths that match the genomic target region, and (2) a transgene, such as a knock-in sequence or a replacement sequence. There is no limit to the size of the transgene. Inserts of 600 bp (Figure 70) and 800 bp (Figure 71) are exemplified herein.
[0227] In certain embodiments, the R-loop-guide RNA binds to an RNA-binding protein or domain fused to a recombinant protein, such as, but not limited to, an SSAP.
[0228] In various embodiments, the present invention provides fusion proteins. In the fusion protein, the recombinant protein can be linked to either end of the aptamer-binding protein in either orientation (e.g., N- to C-terminus, C- to N-terminus, N- to N-terminus). In selected embodiments, the recombinant protein N-terminus is linked to the aptamer-binding protein C-terminus. Thus, the overall N- to C-terminal fusion protein comprises the aptamer-binding protein (N- to C-terminus) linked to the recombinant protein (N- to C-terminus).
[0229] In some embodiments, a recombinant protein may be operably linked to a nuclease as a fusion protein or chimeric molecule. In some embodiments, a recombinant protein may be operably linked to an endonuclease as a fusion protein or chimeric molecule. In some embodiments, a recombinant protein may be operably linked to an exonuclease as a fusion protein or chimeric molecule. In some embodiments, a recombinant protein may be operably linked to a nuclease and / or a Cas or dCas as a fusion protein or chimeric molecule. In some embodiments, a recombinant protein may be operably linked to an endonuclease and / or a Cas or dCas as a fusion protein or chimeric molecule. In some embodiments, a recombinant protein may be operably linked to an exonuclease and / or a Cas or dCas as a fusion protein or chimeric molecule. In some embodiments, a recombinant protein may be operably linked to an exonuclease and / or a Cas or dCas as a fusion protein or chimeric molecule. In some embodiments, a recombinant protein may be expressed independently rather than as a fusion protein with a nuclease. In some embodiments, a recombinant protein may be expressed independently rather than as a fusion protein with an endonuclease. In some embodiments, the recombinant protein may be expressed independently and not as a fusion protein with an exonuclease. In some embodiments, the recombinant protein may be expressed independently and not as a fusion protein with a nuclease and / or Cas or dCas. In some embodiments, the recombinant protein may be operably linked to an aptamer and / or aptamer-binding protein as a fusion protein or chimeric or chimeric molecule. In some embodiments, the recombinant protein may be expressed independently and not as a fusion protein with an aptamer and / or aptamer-binding protein. In some embodiments, the recombinant protein may be operably linked to a nuclease and / or Cas or dCas and / or aptamer and / or aptamer-binding protein as a fusion protein or chimeric or chimeric molecule.In some embodiments, the recombinant protein can be expressed independently, rather than as a fusion protein with a nuclease and / or Cas or dCas and / or an aptamer and / or an aptamer-binding protein. In some embodiments, the aptamer and / or aptamer-binding protein is an MCP protein. In some embodiments, the recombinant protein can be an SSAP.
[0230] As used herein, the term "nuclease" refers to an agent, such as a protein or small molecule, that can cleave the phosphodiester bonds that link nucleotide residues within a nucleic acid molecule. In some embodiments, a nuclease is, for example, an enzyme that can bind to a nucleic acid molecule and cleave the phosphodiester bonds that link nucleotide residues within a nucleic acid molecule, but is woven. A nuclease may be an endonuclease that cleaves phosphodiester bonds in a polynucleotide chain, or an exonuclease that cleaves phosphodiester bonds at the end of a polynucleotide chain. In some embodiments, a nuclease is a site-specific nuclease that binds to and / or cleaves specific phosphodiester bonds within a specific nucleotide sequence (also referred to herein as a "recognition sequence," "nuclease target site," or "target site"). In some embodiments, a nuclease is an RNA-guided (e.g., RNA-programmable) nuclease that complexes (e.g., binds) to RNA having a sequence complementary to the target site, thereby providing the sequence specificity of the nuclease. In some embodiments, the nuclease recognizes a single-stranded target site, while in other embodiments, the nuclease recognizes a double-stranded target site, e.g., a double-stranded DNA target site. Target sites for many naturally occurring nucleases, including many naturally occurring DNA restriction nucleases, are well known to those of skill in the art. Often, DNA nucleases, such as EcoRI, HindIII, or BamHI, recognize palindromic double-stranded DNA target sites of 4 to 10 base pairs in length and cleave each of the two DNA strands at specific positions within the target site. Some endonucleases cleave double-stranded nucleic acid target sites symmetrically, e.g., cleave both strands at the same position, resulting in ends containing base-paired nucleotides, also referred to herein as blunt ends. Other endonucleases cleave double-stranded nucleic acid target sites asymmetrically, e.g., cleave each strand at a different position such that the ends contain unpaired nucleotides.Unpaired nucleotides at the ends of double-stranded DNA molecules are also called "overhangs," e.g., "5' overhangs" or "3' overhangs," depending on whether the unpaired nucleotides form the 5' or 3' end of the corresponding DNA strand. Ends of double-stranded DNA molecules that terminate with unpaired nucleotides are also called sticky ends and can "stick" to the ends of other double-stranded DNA molecules that contain complementary unpaired nucleotides. Nuclease proteins typically contain a "binding domain" that mediates the interaction of the protein with the nucleic acid substrate (and in some cases also specifically binds to the target site) and a "cleavage domain" that catalyzes the cleavage of phosphodiester bonds within the nucleic acid backbone. In some embodiments, nuclease proteins can bind to and cleave nucleic acid molecules in monomeric form, while in other embodiments, the nuclease protein must dimerize or otherwise cleave the target nucleic acid molecule. The binding and cleavage domains of naturally occurring nucleases, as well as binding and cleavage domains that can be fused to generate nucleases, are well known to those of skill in the art. For example, zinc finger or transcription activator-like elements can be used as binding domains to specifically bind to a desired target site and fused or conjugated to a cleavage domain, such as the cleavage domain of fokl, to create an engineered nuclease that cleaves the target site.
[0231] Non-limiting examples of exonucleases include exonuclease I, exonuclease II, exonuclease III, exonuclease IV, exonuclease V, exonuclease VII, exonuclease VIII, lambda exonuclease, Xrn1, mung bean nuclease, TREX2, exonuclease T, T7 exonuclease, strandase exonuclease, 3'-5' exophosphodiesterase, and Bal31 nuclease.
[0232] In some embodiments, the fusion protein further comprises a linker between the recombinant protein and the aptamer-binding protein. The linker can be of any length and comprise any amino acid sequence. The linker can be flexible so that it does not constrain either of the two components it links together in any particular orientation. The linker can essentially act as a spacer. In selected embodiments, the linker connects the C-terminus of the recombinant protein to the N-terminus of the aptamer-binding protein. In selected embodiments, the linker comprises the amino acid sequence of a 16-residue XTEN linker, SGSETPGTSESATPES (SEQ ID NO: 15), or a 37-residue EXTEN linker, SASGGSSGGSSGSETPGTSESATPESSGGSSGGSGGS (SEQ ID NO: 148).
[0233] In some embodiments, the fusion protein further comprises a nuclear localization sequence (NLS). The nuclear localization sequence can be located anywhere within the fusion protein (e.g., at the C-terminus of the aptamer-binding protein, the N-terminus of the aptamer-binding protein, or the C-terminus of the recombinant protein). In selected embodiments, the nuclear localization sequence is linked to the C-terminus of the recombinant protein. Several nuclear localization sequences are known in the art (e.g., Lange, A. et al., J Biol Chem. 2007; 282(8): 5101-5105, incorporated herein by reference) and can be used in connection with the present invention. The nuclear localization sequence can be the SV40 NLS, PKKKRKV (SEQ ID NO: 16); Ty1 NLS, NSKKRSLEDNETEIKVSRDTWNTKNMRSLEPPRSKKRIH (SEQ ID NO: 17); c-Myc NLS, PAAKRVKLD (SEQ ID NO: 18); biSV40 NLS, KRTADGSEFESPKKKRKV (SEQ ID NO: 19); and Mut NLS, PEKKRRRPSGSVPVLARPSPPKAGKSSCI (SEQ ID NO: 20). In selected embodiments, the nuclear localization sequence is the SV40 NLS, PKKKRKV (SEQ ID NO: 16).
[0234] The Cas protein and fusion protein are desirably contained alone in a single composition, or in combination with each other and / or with a polynucleotide (e.g., a vector) containing a guide RNA sequence and an aptamer sequence. The Cas protein and / or fusion protein may or may not be physically or chemically bound to the polynucleotide. The Cas protein and / or recombinant protein may be associated with the polynucleotide using any suitable method for protein-protein or protein-virus linkage known in the art.
[0235] The present invention further provides compositions and vectors comprising a polynucleotide comprising a nucleic acid sequence encoding a fusion protein comprising a microbial recombinant protein operably linked to an RNA aptamer-binding protein.
[0236] The composition or vector may further comprise at least one or both of a polynucleotide comprising a nucleic acid sequence encoding a Cas protein and a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence. In some embodiments, the nucleic acid molecule comprising a guide RNA sequence further comprises at least one RNA aptamer sequence. In some embodiments, the polynucleotide comprising a nucleic acid sequence encoding a Cas protein further comprises a sequence encoding at least one peptide aptamer sequence.
[0237] The descriptions of nucleic acid molecules, including guide RNA sequences, aptamer sequences, Cas proteins, recombination proteins and aptamer-binding proteins, provided above in connection with the systems of the present invention are also applicable to the polynucleotides of the recited compositions and vectors.
[0238] A nucleic acid sequence encoding a Cas protein and / or a nucleic acid sequence encoding a fusion protein comprising a recombinant protein operably linked to an aptamer-binding protein can be provided to a cell on the same vector (e.g., in cis) as a nucleic acid molecule comprising a guide RNA sequence and / or an RNA aptamer sequence. In such embodiments, a unidirectional promoter can be used to control the expression of each nucleic acid sequence. In another embodiment, a combination of a bidirectional promoter and a unidirectional promoter can be used to control the expression of multiple nucleic acid sequences.
[0239] In other embodiments, the nucleic acid sequence encoding the Cas protein, the nucleic acid sequence encoding the fusion protein comprising the recombinant protein operably linked to the aptamer-binding protein, and the nucleic acid molecule comprising the guide RNA sequence and / or the RNA aptamer sequence can be provided to the cell on separate vectors (e.g., in trans). Each of the nucleic acid sequences in each of the separate vectors can comprise the same or different expression control sequences. The separate vectors can be provided to the cell simultaneously or sequentially.
[0240] A vector containing a nucleic acid sequence encoding a fusion protein comprising a Cas protein and a recombinant protein operably linked to an aptamer-binding protein can be introduced into a host cell (including any suitable prokaryotic or eukaryotic cell) capable of expressing the polypeptide encoded thereby. Accordingly, the present invention provides isolated cells containing the vectors or nucleic acid sequences disclosed herein. Preferred host cells are those that can be grown easily and reliably, have a reasonably fast growth rate, possess a well-characterized expression system, and can be easily and efficiently transformed or transfected. Examples of suitable prokaryotic cells include, but are not limited to, cells from the genera Bacillus (e.g., Bacillus subtilis and Bacillus brevis), Escherichia (e.g., E. coli), Pseudomonas, Streptomyces, Salmonella, and Envinia. Suitable eukaryotic cells are known in the art and include, for example, yeast, insect, and mammalian cells. Examples of suitable yeast cells include cells from the genera Kluyveromyces, Pichia, Rhino-sporidium, Saccharomyces, and Schizosaccharomyces. Exemplary insect cells include Sf-9 and HIS (Invitrogen, Carlsbad, Calif.), as described, for example, in Kitts et al., Biotechniques, 14:810-817 (1993); Lucklow, Curr. Opin. Biotechnol., 4:564-572 (1993); and Lucklow et al., J. Virol., 67:4566-4579 (1993), which are incorporated herein by reference. Desirably, the host cell is a mammalian cell, and in some embodiments, the host cell is a human cell. Many suitable mammalian and human host cells are known in the art, and many are available from the American Type Culture Collection (ATCC, Manassas, Va.).Examples of suitable mammalian cells include, but are not limited to, Chinese hamster ovary cells (CHO) (ATCC No. CCL61), CHO DHFR cells (Urlaub et al., Proc. Natl. Acad. Sci. USA, 97:4216-4220 (1980)), human embryonic kidney (HEK) 293 or 293T cells (ATCC No. CRL1573), and 3T3 cells (ATCC No. CCL92). Other suitable mammalian cell lines are the monkey COS-1 (ATCC No. CRL1650) and COS-7 cell lines (ATCC No. CRL1651), and the CV-1 cell line (ATCC No. CCL70). Further exemplary mammalian host cells include primate, rodent, and human cell lines, including transformed cell lines. Normal diploid cells, cell lines derived from in vitro culture of primary tissues, and primary explants are also suitable. Other suitable mammalian cell lines include, but are not limited to, mouse neuroblastoma N2A cells, HeLa, HEK, A549, HepG2, mouse L-929 cells, and BHK or HaK hamster cell lines. Methods for selecting suitable mammalian host cells, as well as methods for transforming, culturing, amplifying, screening, and purifying the cells, are known in the art.
[0241] Methods for modifying target DNA The present invention also provides a method for modifying target DNA. In some embodiments, the method modifies genomic DNA sequences in cells, although any desired nucleic acid can be modified. When applied to DNA contained in cells, the method comprises introducing the system, composition, or vector described herein into cells containing the target genomic DNA sequence. The above-described nucleic acid molecules including guide RNA sequences, Cas proteins, recombinant proteins, mobilization systems and polynucleotides encoding them, cells, target genomic DNA sequences, and their components in relation to the system of the present invention are also applicable to the method for modifying target genomic DNA sequences in cells. Depending on the cell type, the system, composition, or vector can be introduced by any method known in the art, including, but not limited to, chemical transfection, electroporation, microinjection, biolistic delivery via a gene gun, or magnetically assisted transfection.
[0242] In certain embodiments, delivery of the editing system or component comprises delivery of a ribonucleoprotein (RNP) complex. According to the present invention, targeting nucleic acids, including but not limited to, gRNA and dgRNA, can be provided in a complex, such as but not limited to, a complex comprising Cas9 or dCas9. In certain embodiments, the RNP complex comprises a guide nucleic acid and a Cas9 fusion protein, such as but not limited to, a complex comprising dCas9-SSAP. In certain embodiments, the RNP complex comprises a guide nucleic acid and a recombinant protein, such as SSAP or SSB, that can be adapted or modified to bind to the guide nucleic acid. In certain embodiments, the guide nucleic acid and the recombinant protein or Cas9 fusion protein comprise a binding element that promotes complex formation. In a non-limiting example, the recombinant protein comprises an MCP domain and the guide RNA comprises an MS2 aptamer, whereby binding of the MS2 aptamer to the MCP domain generates an RNP.
[0243] In some embodiments, the guide RNA and the Cas and / or recombinant protein polypeptide are incubated together to form a ribonucleoprotein (RNP) complex, which is then introduced into the cell, e.g., mixed together in a vessel to form the RNP complex, and then the RNP complex is introduced into the cell. In other embodiments, the Cas polypeptide described herein can be an mRNA encoding the Cas polypeptide, and the Cas mRNA is introduced into the primary cell along with the modified sgRNA as an "All RNA" CRISPR system.
[0244] In some embodiments, RNP complex and donor nucleic acid or vector are simultaneously introduced into cell.In other embodiments, RNP complex and donor nucleic acid or vector are sequentially introduced into primary cell.In some examples, RNP complex is introduced into primary cell before donor.In other examples, donor is introduced into primary cell before RNP complex. For example, the RNP complex can be introduced into the cell about 1 minute, about 2 minutes, about 3 minutes, about 4 minutes, about 5 minutes, about 6 minutes, about 7 minutes, about 8 minutes, about 9 minutes, about 10 minutes, about 11 minutes, about 12 minutes, about 13 minutes, about 14 minutes, about 15 minutes, about 16 minutes, about 17 minutes, about 18 minutes, about 19 minutes, about 20 minutes, about 25 minutes, about 30 minutes, about 35 minutes, about 40 minutes, about 45 minutes, about 50 minutes, about 55 minutes, about 60 minutes, about 90 minutes, about 120 minutes, about 150 minutes, about 180 minutes, about 210 minutes, or about 240 minutes or more before the donor nucleic acid or vector, or vice versa. U.S. Patent No. 11,193,141 describes the introduction of an RNP complex and a homologous donor adeno-associated virus (AAV) vector into cells to mediate targeted integration. The described method can be used with the present invention. U.S. Patent Application Publication No. 2019 / 0093128 describes the introduction into a zygote of a ribonucleoprotein (RNP) containing a class 2 CRISPR / Cas endonuclease complexed with a corresponding CRISPR / Cas guide RNA that hybridizes to a target sequence in the genomic DNA of the zygote.
[0245] Non-limiting examples include the use of (1) purified Cas9 or dCas9 protein; (2) synthetic guide RNA with the MS2 aptamer; (3) purified MCP-SSAP fusion protein; and (4) donor DNA (double-stranded and single-stranded DNA donors for HEK293 and K562, and AAV donors for HSCs) delivered into HEK293, K562, and primary hematopoietic stem cells (mouse and human) for knock-in editing.
[0246] The following table provides exemplary sequences for generating knock-ins containing ALB and AAVS1. The sequences can be used in RNPs, nucleic acids, vectors, etc. for expression. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9]
[0247] When the system described herein is introduced into a cell containing a target genomic DNA sequence, the guide RNA sequence binds to the target genomic DNA sequence in the cell genome, the Cas protein associates with the guide RNA and can induce a double-strand break or a single-strand nick in the target genomic DNA sequence, and the aptamer recruits the recombinant protein to the target genomic DNA sequence through the aptamer-binding protein of the fusion protein, thereby altering the target genomic DNA sequence in the cell. When the composition or vector described herein is introduced into a cell, the nucleic acid molecule containing the guide RNA sequence, the Cas9 protein, and the fusion protein are first expressed in the cell.
[0248] In some embodiments, the cells are in an organism or host, such that introducing the disclosed systems, compositions, or vectors into the cells comprises administration to a subject. This method can include providing or administering to a subject in vivo or by transplantation of ex vivo treated cells, systems, compositions, or vectors of the system.
[0249] A "subject" can be human or non-human, and can include, for example, animal strains or species used as "model systems" for research purposes, such as mouse models as described herein. Similarly, a subject can include either an adult or a juvenile (e.g., a child). Furthermore, a subject can refer to any organism, preferably a mammal (e.g., human or non-human), that can benefit from the administration of the compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates, such as chimpanzees, and other ape and monkey species; livestock such as cows, horses, sheep, goats, and pigs; domestic animals such as rabbits, dogs, and cats; laboratory animals, including rodents such as rats, mice, and guinea pigs; and the like. Examples of non-mammals include, but are not limited to, birds, fish, and the like. In one embodiment of the methods and compositions provided herein, the mammal is a human. Plants include, but are not limited to, sugarcane, corn, wheat, rice, oil palm, potato, soybean, vegetables, cassava, sugar beet, tomato, barley, banana, watermelon, onion, sweet potato, cucumber, apple, seed cotton, orange, and the like.
[0250] As used herein, the terms "providing," "administering," and "introducing" are used interchangeably herein and refer to the placement of a system of the present invention into a subject by a method or route that results in at least partial localization of the system to a desired site. The system may be administered by any suitable route that results in delivery to the desired location in the subject.
[0251] As used herein, the phrase " modify DNA sequence " refers to modify at least one physical characteristic of the DNA sequence of interest.DNA modification includes, for example, single-strand or double-strand DNA break, one or more nucleotide deletion or insertion, and other modifications that affect the structural integrity or nucleotide sequence of DNA sequence.The modification of target sequence in genomic DNA can result in, for example, gene correction, gene replacement, gene tagging, transgene insertion, nucleotide deletion, gene disruption, gene mutation, gene knockdown, etc.
[0252] In some embodiments, the systems and methods described herein can be used to correct one or more defects or mutations in a gene (referred to as "gene correction"). In such cases, the target genomic DNA sequence encodes a defective version of the gene, and the system further includes a donor nucleic acid molecule encoding a wild-type or corrected version of the gene. Thus, in other words, the target genomic DNA sequence is a "disease-associated" gene. The term "disease-associated gene" refers to any gene or polynucleotide whose gene product is expressed at an abnormal level or in an abnormal form in cells obtained from a disease-affected individual compared to tissues or cells obtained from a disease-free individual. A disease-associated gene can be expressed at an abnormally high or low level, and altered expression correlates with the development and / or progression of the disease. A disease-associated gene also refers to a gene whose mutation or genetic variation is directly responsible or that is in linkage disequilibrium with a gene(s) involved in the etiology of the disease. Examples of such "single gene" or "monogenic" disease-causing genes include, but are not limited to, adenosine deaminase, alpha-1 antitrypsin, cystic fibrosis transmembrane conductance regulator (CFTR), beta-hemoglobin (HBB), oculocutaneous albinism II (OCA2), huntingtin (HTT), myotonic dystrophy protein kinase (DMPK), low-density lipoprotein receptor (LDLR), apolipoprotein B (APOB), neurofibromin 1 (NF1), polycystic kidney disease 1 (PKD1), polycystic kidney disease 2 (PKD2), coagulation factor VIII (F8), dystrophin (DMD), phosphate-regulated endopeptidase homolog, X-linked (PHEX), methyl-CpG-binding protein 2 (MECP2), and ubiquitin-specific peptidase 9Y, Y-linked (USP9Y).Other single gene or monogenic diseases are known in the art and are described, for example, in Chial, H. Rare Genetic Disorders: Learning About Genetic Disease Through Gene Mapping, SNPs, and Microarray Data, Nature Education 1(1):192 (2008); Online Mendelian Inheritance in Man (OMIM); and the Human Gene Mutation Database (HGMD), which are incorporated herein by reference.
[0253] The present invention provides for knock-in of large transgenes at therapeutically relevant loci in the human genome. In certain embodiments, the loci provide cell- or tissue-specific expression. In certain embodiments, the invention involves inserting a nucleic acid into the albumin (ALB) locus. The ALB locus provides liver targeting in human hepatocytes and is highly expressed in a liver-specific manner. In certain embodiments, the invention involves inserting a nucleic acid into the AAVS1 locus. The AAVS1 locus is well expressed in certain tissue types, can be used for a wide variety of treatments, and is a safe harbor locus for gene therapy due to its low liver expression. U.S. Patent Application Publication No. 2018 / 0214490 A1 describes gene therapy for lysosomal storage diseases, involving targeting transgenes to safe harbor loci such as the AAVS1, HPRT, and CCR5 genes in human cells and Rosa26 in mouse cells. U.S. Patent No. 9,267,154 describes the integration of exogenous nucleic acid sequences into the PPP1R12C locus, which is widely expressed in most tissues. Cell-specific expression by targeting a transgene (e.g., encoding a chimeric antigen receptor (CAR)) to the T cell receptor alpha constant (TRAC) locus is described. These are exemplary and non-limiting with respect to the loci that can be targeted according to the present invention.
[0254] In another embodiment, the target genomic DNA sequence can contain a gene whose mutation, in combination with mutations in other genes, contributes to a specific disease.Diseases caused by the contribution of multiple genes that lack a simple (e.g., Mendelian) inheritance pattern are referred to in the art as "multifactorial" or "polygenic" diseases.Examples of multifactorial or polygenic diseases include, but are not limited to, asthma, diabetes, epilepsy, hypertension, bipolar disorder, and schizophrenia.Certain developmental abnormalities can also be inherited in a multifactorial or polygenic pattern, including, for example, cleft lip / palate, congenital heart defects, and neural tube defects.
[0255] In another embodiment, the method for modifying target genomic DNA sequence can be used to delete nucleic acid from target sequence in cell by cutting target sequence and allowing cell to repair cut sequence in the absence of exogenously provided donor nucleic acid molecule.The deletion of nucleic acid sequence in this manner can be used in various applications, such as for example, to remove disease-causing trinucleotide repeat sequence in neural cells, to create gene knockout or knockdown, and to generate mutations for disease model in research.
[0256] The term "donor nucleic acid molecule" refers to a nucleotide sequence to be inserted into target DNA (e.g., genomic DNA). As described above, donor DNA can include, for example, a sequence encoding a gene or part of a gene, a tag or localization sequence, or a regulatory element. Donor nucleic acid molecules can be of any length. In some embodiments, donor nucleic acid molecules are 10 to 10,000 nucleotides in length. For example, they are about 100 to 5,000 nucleotides in length, about 200 to 2,000 nucleotides in length, about 500 to 1,000 nucleotides in length, about 500 to 5,000 nucleotides in length, about 1,000 to 5,000 nucleotides in length, or about 1,000 to 10,000 nucleotides in length.
[0257] The disclosed systems and methods overcome challenges encountered during conventional gene editing, including low efficiency and off-target events, particularly with kilobase-scale nucleic acids. In some embodiments, the disclosed systems and methods improve the efficiency of gene editing. For example, the disclosed systems and methods can have efficiencies that are 2- to 10-fold higher than conventional CRISPR-Cas9 systems and methods, as shown in Examples 2, 3, and 5. In some embodiments, the improved efficiency is accompanied by a reduction in off-target events. Off-target events can be reduced by more than 50% compared to conventional CRISPR-Cas9 systems and methods; for example, an approximately 90% reduction in off-target events is shown in Example 3. Another aspect of improving the overall accuracy of gene editing systems is the reduction of on-target insertions / deletions (indels), a by-product of HDR editing. In some embodiments, the disclosed systems and methods reduce on-target indels by more than 90% compared to conventional CRISPR-Cas9 systems and methods, as shown in Example 3.
[0258] The present invention further provides kits that contain one or more reagents or other components that are useful, necessary, or sufficient for carrying out any of the methods described herein.For example, the kits can include CRISPR reagents (Cas proteins, guide RNAs, vectors, compositions, etc.), recombineering reagents (recombinant protein-aptamer binding protein fusion proteins, aptamer sequences, vectors, compositions, etc.), transfection or administration reagents, negative and positive control samples (e.g., cells, template DNA), cells, containers (e.g., microcentrifuge tubes, boxes) that contain one or more components, detectable labels, detection and analysis equipment, software, instructions, etc.
[0259] RNA can be delivered using adeno-associated virus (AAV), lentivirus, adenovirus, or other viral vector types, or a combination thereof. RNA can be packaged into one or more viral vectors. In some embodiments, the viral vector is delivered to the target tissue, for example, by intramuscular injection; in other cases, viral delivery is intravenous, transdermal, intranasal, oral, mucosal, or other delivery methods. Such delivery can be by either a single dose or multiple doses. Those skilled in the art will understand that the actual dosage delivered herein can vary greatly depending on various factors, such as the selected vector, target cells, organisms, or tissues, the systemic symptoms of the subject being treated, the degree of transformation / modification desired, the route of administration, the mode of administration, and the type of transformation / modification desired.
[0260] Such dosages may further contain, for example, carriers (such as water, saline, ethanol, glycerol, lactose, sucrose, calcium phosphate, gelatin, dextran, agar, pectin, peanut oil, sesame oil, etc.), diluents, pharmaceutically acceptable carriers (such as phosphate-buffered saline), pharmaceutically acceptable excipients, and / or other compounds known in the art. Such dosage formulations can be readily ascertained by one skilled in the art. The dosages may further contain one or more pharmaceutically acceptable salts, such as mineral acid salts, e.g., hydrochlorides, hydrobromides, phosphates, sulfates, etc.; and salts of organic acids, e.g., acetates, propionates, malonates, benzoates, etc. Additionally, auxiliary substances, such as wetting or emulsifying agents, pH buffering substances, gels or gelling materials, flavorings, coloring agents, microspheres, polymers, suspending agents, etc., may also be present herein. Additionally, particularly when the dosage form is reconstitutable, one or more other conventional pharmaceutical ingredients may also be present, such as preservatives, humectants, suspending agents, surfactants, antioxidants, anti-caking agents, fillers, chelating agents, coating agents, chemical stabilizers, etc. Suitable exemplary ingredients include microcrystalline cellulose, sodium carboxymethylcellulose, polysorbate 80, phenylethyl alcohol, chlorobutanol, potassium sorbate, sorbic acid, sulfur dioxide, propyl gallate, parabens, ethyl vanillin, glycerin, phenol, parachlorophenol, gelatin, albumin, and combinations thereof. A complete discussion of pharmaceutically acceptable excipients is available in REMINGTON'S PHARMACEUTICAL SCIENCES (Mack Pub. Co., NJ 1991), which is incorporated herein by reference.
[0261] In one embodiment herein, delivery is at least 1 x 10 5 In one embodiment herein, the dose is preferably at least about 1 x 10 adenovirus vector particles (also called particle units, pu). 6 particles (e.g., about 1 × 10 6 ~1×1×10 12particles), more preferably at least about 1×10 10 particles, more preferably at least about 1×10 8 particles (e.g., about 1 × 10 8 ~1×10 11 particles or approximately 1 x 10 8 ~1×10 12 particles), most preferably at least about 1×10 10 particles (e.g., about 1 × 10 9 ~1×10 10 particles or approximately 1 x 10 9 ~1×10 12 particles), or even at least about 1 × 10 10 particles (e.g., about 1 × 10 10 ~1×10 12 particles) of adenoviral vector. Alternatively, the dose is about 1 x 10 14 particles or less, preferably about 1 x 10 13 particles or less, and even more preferably about 1×10 12 particles or less, and even more preferably about 1×10 11 particles or less, most preferably about 1 x 10 10 particles (e.g., about 1 × 10 9 Thus, the dose may be, for example, about 1 x 10 6 Particle unit (pu), approximately 2 x 10 6 pu, approx. 4×10 6 pu, about 1×10 7 pu, approx. 2×10 7 pu, approx. 4×10 7 pu, about 1×10 8 pu, approx. 2×10 8 pu, approx. 4×10 8 pu, about 1×10 9 pu, approx. 2×10 9 pu, approx. 4×10 9 pu, about 1×10 10 pu, approx. 2×10 10 pu, approx. 4×10 10 pu, about 1×10 11 pu, approx. 2×10 11 pu, approx. 4×10 11 pu, about 1×10 12 pu, approx. 2×1012 pu, or approximately 4 × 10 12 The adenovirus vector may contain a single dose of pu. See, for example, U.S. Patent No. 8,454,972 (Nabel et al., issued June 4, 2013) for adenovirus vectors and dosages at col. 29, lines 36-58, which is incorporated herein by reference. In one embodiment herein, the adenovirus is delivered in multiple doses.
[0262] In one embodiment herein, delivery is via AAV. The therapeutically effective dose for in vivo delivery of AAV to humans is about 1×10 10 ~Approx. 1×10 10 The range of effective doses is believed to be about 20 to about 50 ml of saline containing 1 x 10 functional AAV / ml solution. Dosages can be adjusted to balance therapeutic benefit with any side effects. In one embodiment herein, the AAV dose is generally about 1 x 10 5 ~1×10 50 Genomic AAV, approximately 1 × 10 8 ~1×10 20 Genomic AAV, approximately 1 × 10 10 ~Approx. 1×10 16 genome or approximately 1 × 10 11 ~Approx. 1×10 16 The concentration of genomic AAV is in the range of 1 × 10 13 The AAV genome may be delivered in a concentration of about 0.001 ml to about 100 ml, about 0.05 ml to about 50 ml, or about 10 ml to about 25 ml of carrier solution. Other effective dosages can be readily established by those skilled in the art through routine testing to establish a dose-response curve. See, for example, U.S. Patent No. 8,404,658 by Hajjar et al., issued March 26, 2013, col. 27, lines 45-60.
[0263] In one embodiment of the present disclosure, delivery is via a plasmid. In such a plasmid composition, the dosage must be a sufficient amount of plasmid to induce a response. For example, an appropriate amount of plasmid DNA in a plasmid composition may be about 0.1 to about 2 mg, or about 1 μg to about 10 μg.
[0264] The dosages herein are based on an average individual weighing 70 kg. The frequency of administration is within the knowledge of a medical or veterinary practitioner (e.g., a doctor, a veterinarian) or a scientist in the field. The mice used in the experiment weigh approximately 20 g. The dosage administered to a 20 g mouse can be extrapolated to a 70 kg individual.
[0265] Lentiviruses are complex retroviruses that have the ability to infect and express their genes in both mitotic and postmitotic cells. The most commonly known lentivirus is the human immunodeficiency virus (HIV), which uses the envelope glycoproteins of other viruses to target a wide range of cell types.
[0266] Lentivirus can be prepared as follows. After cloning pCasES10 (containing the lentiviral transfer plasmid backbone), low-passage (p=5) HEK293FT cells were seeded in T-75 flasks to 50% confluence the day before transfection in DMEM containing 10% fetal bovine serum and no antibiotics. After 20 hours, the medium was replaced with OptiMEM (serum-free) medium, and transfection was performed 4 hours later. Cells were transfected with 10 μg of the lentiviral transfer plasmid (pCasES10) and the following packaging plasmids: 5 μg of pMD2.G (VSV-g pseudotyped), and 7.5 μg of psPAX2 (gag / pol / rev / tat). Transfections were performed in 4 mL of OptiMEM containing cationic lipid delivery agents (50 μL of Lipofectamine 2000 and 100 μL of Plus Reagent). After 6 hours, the medium was replaced with antibiotic-free DMEM containing 10% fetal bovine serum.
[0267] Lentiviruses can be purified as follows. After 48 hours, viral supernatants were harvested. The supernatants were first cleared of debris and filtered through a 0.45 μm low-protein-binding (PVDF) filter. They were then spun at 24,000 rpm in an ultracentrifuge for 2 hours. The viral pellets were resuspended in 50 μl of DMEM overnight at 4°C. They were then aliquoted and immediately frozen at -80°C.
[0268] In another embodiment, minimal non-primate lentiviral vectors based on equine infectious anemia virus (EIAV) are also contemplated, particularly for ocular gene therapy (see, e.g., Balagaan, J Gene Med 2006;8:275-285, published online November 21, 2005, in Wiley InterScience; available at: interscience.wiley.com.DOI:10.1002 / jgm.845). In another embodiment, RetinoStat®, a lentiviral gene therapy vector based on equine infectious anemia virus that expresses the angiogenesis inhibitor proteins endostein and angiostatin, delivered via subretinal injection to treat web-like age-related macular degeneration, is also contemplated (see, e.g., Binley et al., HUMAN GENE THERAPY 23:980-991 (September 2012)), and may be modified for use in the system of the present invention.
[0269] Lentiviral vectors have been disclosed for the treatment of Parkinson's disease, as well, see, for example, US Patent Publication No. 20120295960 and US Patent Nos. 7,303,910 and 7,351,585. Lentiviral vectors have also been disclosed for the treatment of ocular diseases, as see, for example, US Patent Publication Nos. 20060281180, 20090007284, US20110117189; US20090017543; US20070054961, US20100317109. Lentiviral vectors have also been disclosed for delivery to the brain, see, e.g., U.S. Patent Publication Nos. US20110293571; US20110293571, US20040013648, US20070025970, US20090111106 and U.S. Patent No. 7,259,015.
[0270] Several types of particle delivery systems and / or formulations are known to be useful in a variety of biomedical applications. Generally, particles are defined as small objects that behave as whole units with respect to their transport and properties. Particles are further classified according to their diameter. Coarse particles cover the range of 2,500 to 10,000 nanometers. Fine particles are between 100 and 2,500 nanometers in size. Ultrafine particles or nanoparticles are generally between 1 and 100 nanometers in size. The basis for the 100 nm limit is the fact that the novel properties that distinguish particles from bulk materials typically manifest at a critical length scale below 100 nm.
[0271] As used herein, a particle delivery system / formulation is defined as any biological delivery system / formulation comprising particles according to the present invention. A particle according to the present invention is any entity having a maximum dimension (e.g., diameter) of less than 100 microns (μm). In some embodiments, particles of the present invention have a maximum dimension of less than 10. In some embodiments, particles of the present invention have a maximum dimension of less than 2000 nanometers (nm). In some embodiments, particles of the present invention have a maximum dimension of less than 1000 nanometers (nm). In some embodiments, particles of the present invention have a maximum dimension of less than 900 nm, 800 nm, 700 nm, 600 nm, 500 nm, 400 nm, 300 nm, 200 nm, or 100 nm. Typically, particles of the present invention have a maximum dimension (e.g., diameter) of 500 nm or less. In some embodiments, particles of the present invention have a maximum dimension (e.g., diameter) of 250 nm or less. In some embodiments, particles of the present invention have a maximum dimension (e.g., diameter) of 200 nm or less. In some embodiments, particles of the present invention have a maximum dimension (e.g., diameter) of 150 nm or less. In some embodiments, particles of the present invention have a maximum dimension (e.g., diameter) of 100 nm or less. Smaller particles, for example, having a maximum dimension of 50 nm or less, are used in some embodiments of the present invention. In some embodiments, particles of the present invention have a maximum dimension in the range of 25 nm to 200 nm.
[0272] Particle characterization (e.g., characterizing morphology, dimensions, etc.) is performed using a variety of different techniques. Common techniques are electron microscopy (TEM, SEM), atomic force microscopy (AFM), dynamic light scattering (DLS), X-ray photoelectron spectroscopy (XPS), powder X-ray diffraction (XRD), Fourier transform infrared spectroscopy (FTIR), matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF), ultraviolet-visible spectroscopy, dual-polarization interferometry, and nuclear magnetic resonance (NMR). Characterization (size measurement) may be performed on the native particles (e.g., before loading) or after cargo loading (herein, cargo refers to one or more RNAs and / or vectors encoding same, which may include additional components, carriers, and / or excipients) to provide particles of optimal size for delivery for any in vitro, ex vivo, and / or in vivo application of the present invention. In certain preferred embodiments, particle size (e.g., diameter) characterization is based on measurements using dynamic laser scattering (DLS).
[0273] Particulate delivery systems within the scope of the present invention can be provided in any form, including, but not limited to, solids, semi-solids, emulsions, or colloidal particles. As such, any delivery system described herein (e.g., including, but not limited to, lipid-based systems, liposomes, micelles, microvesicles, exosomes, or gene guns) can be provided as particulate delivery systems within the scope of the present invention.
[0274] In the present invention, it is preferred to have one or more components of the system be delivered using nanoparticles or lipid envelopes.CRISPR enzyme mRNA and guide RNA can be delivered simultaneously using nanoparticles or lipid envelopes.Other delivery systems or vectors can also be used in conjunction with the nanoparticle embodiment of the present invention.
[0275] Generally, "nanoparticle" refers to any particle having a diameter of less than 1000 nm. In certain preferred embodiments, nanoparticles of the present invention have a maximum dimension (e.g., diameter) of 500 nm or less. In other preferred embodiments, nanoparticles of the present invention have a maximum dimension in the range of 25 nm to 200 nm. In other preferred embodiments, nanoparticles of the present invention have a maximum dimension of 100 nm or less. In other preferred embodiments, nanoparticles of the present invention have a maximum dimension in the range of 35 nm to 60 nm.
[0276] The nanoparticles encompassed by the present invention can be provided in different forms, such as solid nanoparticles (e.g., metals such as silver, gold, iron, and titanium), non-metallic, lipid-based solids, polymers, nanoparticle suspensions, or combinations thereof. Metallic, dielectric, and semiconductor nanoparticles, as well as hybrid structures (e.g., core-shell nanoparticles), can be prepared. Nanoparticles made of semiconductor materials can also be labeled quantum dots if they are small enough (typically less than 10 nm) for quantization of electronic energy levels to occur. Such nanoscale particles are used as drug carriers or imaging agents in biomedical applications and can be adapted for similar purposes in the present invention.
[0277] Semi-solid and soft nanoparticles have been produced and are within the scope of the present invention. The prototype nanoparticle of semi-solid nature is the liposome. Various types of liposome nanoparticles are currently used clinically as delivery systems for anti-cancer drugs and vaccines. Nanoparticles with one half hydrophilic and the other half hydrophobic, called Janus particles, are particularly effective in stabilizing emulsions. They can self-assemble at the water / oil interface and act as solid surfactants.
[0278] For example, Su X, Fricke J, Kavanagh DG, and Irvine DJ ("In vitro and in vivo mRNA delivery using lipid-enveloped pH-responsive polymer nanoparticles," Mol Pharm. 2011 Jun. 6;8(3):774-87. doi:10.1021 / mp100390w. Epub 2011 Apr. 1) describe biodegradable core-shell nanoparticles with a poly(β-amino ester) (PBAE) core encased in a phospholipid bilayer shell. These were developed for in vivo mRNA delivery. The pH-responsive PBAE component was selected to promote endosomal disruption, while the lipid surface layer was selected to minimize the toxicity of the polycation core. Therefore, they are preferred for delivering the RNA of the present invention.
[0279] In one embodiment, nanoparticles based on self-assembling bioadhesive polymers are contemplated, which may be applied to oral delivery of peptides, intravenous delivery of peptides, and nasal delivery of peptides, all to the brain. Other embodiments, such as oral absorption and ocular delivery of hydrophobic drugs, are also contemplated. Molecular envelope technology involves engineered polymer envelopes that are protected and delivered to disease sites (e.g., Mazza, M. et al., ACS Nano, 2013.7(2):1016-1026; Siew, A. et al., Mol Pharm, 2012.9(1):14-28; Lalatsa, A. et al., J Contr Rel, 2012.161(2):523-36; Lalatsa, A. et al., Mol Pharm, 2012.9(6):1665-80; Lalatsa, A. et al., Mol Pharm, 2012.9(6):1764-74; Garrett, N.L. et al., J Biophotonics, 2012.5(5-6):458-68; Garrett, N.L. et al., J Raman Spect, 2012.43(5):681-688; Ahmad, S. et al., J Royal Soc Interface 2010.7:S423-33; Uchegbu, IF Expert Opin Drug Deliv, 2006.3(5):629-40; Qu, X. et al., Biomacromolecules, 2006.7(12):3452-9 and Uchegbu, IF et al., Int J Pharm, 2001.224:185-199). Depending on the target tissue, a dose of about 5mg / kg is contemplated, in single or multiple doses.
[0280] In one embodiment, nanoparticles capable of delivering RNA to cancer cells to halt tumor growth, developed by Dan Anderson's laboratory at MIT, may be used and / or adapted to the CRISPR-Cas system of the present invention. In particular, the Anderson laboratory has developed a fully automated combinatorial system for the synthesis, purification, characterization, and formulation of new biomaterials and nanoformulations. See, for example, Alabi et al., Proc Natl Acad Sci USA. 2013 Aug. 6; 110(32):12881-6; Zhang et al., Adv Mater. 2013 Sep. 6; 25(33):4641-5; Jiang et al., Nano Lett. 2013 Mar. 13; 13(3):1059-64; Karagiannis et al., ACS Nano. 2012 Oct. 23; 6(10):8484-7; Whitehead et al., ACS Nano. 2012 Aug. 28; 6(8):6922-9 and Lee et al., Nat Nanotechnol. 2012 Jun. 3; 7(6):389-93.
[0281] US Patent Application No. 20110293703 relates to lipidoid compounds that are particularly useful in the administration of polynucleotides, which can be applied to deliver the CRISPR Cas system of the present invention. In one embodiment, the aminoalcohol lipidoid compounds are combined with a drug to be delivered to a cell or subject to form microparticles, nanoparticles, liposomes, or micelles. The drug delivered by the particles, liposomes, or micelles can be in gas, liquid, or solid form, and the drug can be a polynucleotide, protein, peptide, or small molecule. The aminoalcohol lipidoid compounds can be combined with other aminoalcohol lipidoid compounds, polymers (synthetic or natural), surfactants, cholesterol, carbohydrates, proteins, lipids, etc. to form particles. These particles can then be combined with pharmaceutical excipients as needed to form pharmaceutical compositions.
[0282] US Patent Publication No. 0110293703 also provides a method for preparing aminoalcohol lipidoid compounds. One or more equivalents of an amine are reacted with one or more equivalents of an epoxide-terminated compound under appropriate conditions to form the aminoalcohol lipidoid compounds of the present invention. In certain embodiments, all amino groups of the amine are completely reacted with the epoxide-terminated compound to form tertiary amines. In other embodiments, all amino groups of the amine are not completely reacted with the epoxide-terminated compound to form tertiary amines, thereby resulting in primary or secondary amines in the aminoalcohol lipidoid compound. These primary or secondary amines can be left as is or reacted with another electrophile, such as a different epoxide-terminated compound. As will be understood by those skilled in the art, reacting an amine with less than an excess of an epoxide-terminated compound can result in multiple different aminoalcohol lipidoid compounds with various numbers of tails. Certain amines can be fully functionalized with two epoxide-derived compound tails, while other molecules are not fully functionalized with epoxide-derived compound tails. For example, a diamine or polyamine may contain one, two, three, or four epoxide-derived compound tails from various amino moieties on the molecule, resulting in primary, secondary, and tertiary amines. In certain embodiments, all amino groups are not fully functionalized. In certain embodiments, two of the same type of epoxide-terminated compound are used. In other embodiments, two or more different epoxide-terminated compounds are used. The synthesis of aminoalcohol lipidoid compounds can be performed with or without solvent, and the synthesis can be performed at higher temperatures ranging from 30 to 100°C, preferably around 50 to 90°C. The prepared aminoalcohol lipidoid compounds can be purified as needed. For example, a mixture of aminoalcohol lipidoid compounds can be purified to obtain an aminoalcohol lipidoid compound with a specific number of epoxide-derived compound tails. Alternatively, the mixture can be purified to obtain a specific stereoisomer or positional isomer.The amino alcohol lipidoid compounds may be alkylated with alkyl halides (eg, methyl iodide) or other alkylating agents, and / or acylated.
[0283] U.S. Patent Publication No. 0110293703 also provides libraries of amino alcohol lipidoid compounds prepared by the methods of the invention. These amino alcohol lipidoid compounds can be prepared and / or screened using high-throughput techniques, including liquid handlers, robots, microtiter plates, computers, and the like. In certain embodiments, the amino alcohol lipidoid compounds are screened for their ability to transfect cells with polynucleotides or other agents (e.g., proteins, peptides, small molecules).
[0284] U.S. Patent Application Publication No. 20130302401 describes a class of poly(β-amino alcohols) (PBAAs) prepared using combinatorial polymerization. The PBAAs of the present invention can be used in biotechnology and biomedical applications as coatings (such as coatings for films or multilayer films for medical devices or implants), additives, materials, excipients, anti-biofouling agents, micropatterning agents, and cell encapsulation agents. When used as surface coatings, these PBAAs induced different levels of inflammation both in vitro and in vivo depending on their chemical structure. The great chemical diversity of this class of materials allowed us to identify polymer coatings that inhibit macrophage activation in vitro. Furthermore, these coatings reduced inflammatory cell recruitment and fibrosis after subcutaneous implantation of carboxylated polystyrene microparticles. These polymers can be used to form polyelectrolyte complex capsules for cell encapsulation. The present invention may also have many other biological applications, such as antimicrobial coatings, DNA or siRNA delivery, and stem cell tissue engineering. The teachings of US Patent Application Publication No. 20130302401 can be applied to the present system.
[0285] In another embodiment, lipid nanoparticles (LNPs) are contemplated. In particular, anti-transthyretin small interfering RNA encapsulated in lipid nanoparticles (see, e.g., Coelho et al., N Engl J Med 2013;369:819-29) may be applied to the system of the present invention. Doses of about 0.01 to about 1 mg per kg of body weight administered intravenously are contemplated. Medication to reduce the risk of infusion-related reactions is contemplated, such as dexamethasone, acetaminophen, diphenhydramine or cetirizine, and ranitidine. Multiple doses of about 0.3 mg / kilogram every four weeks for five doses are also contemplated. Lipids include, but are not limited to, DLin-KC2-DMA4, C12-200, and the lipids disteroylphosphatidylcholine, cholesterol, and PEG-DMG, which can be formulated in place of siRNA using a spontaneous vesicle formation procedure (see, e.g., Novobrantseva, Molecular Therapy-Nucleic Acids (2012) 1, e4; doi:10.1038 / mtna.2011.3). The component molar ratio can be approximately 50 / 10 / 38.5 / 1.5 (DLin-KC2-DMA or C12-200 / disteroylphosphatidylcholine / cholesterol / PEG-DMG). The final lipid:siRNA weight ratio can be approximately 12:1 and 9:1 for DLin-KC2-DMA and C12-200 lipid nanoparticles (LNPs), respectively. The formulation can have an average particle size of approximately 80 nm with an entrapment efficiency of over 90%. A dose of 3 mg / kg may be contemplated.
[0286] LNP has been shown to be highly effective in delivering siRNA to the liver (see, e.g., Tabernero et al., Cancer Discovery, April 2013, Vol. 3, No. 4, pages 363-470), and therefore, it is contemplated to deliver CRISPR-Cas to the liver. Approximately four doses of 6 mg / kg LNP (or CRISPR-Cas RNA) every two weeks may be contemplated. Tabernero et al. demonstrated that tumor regression was observed after the first two cycles of LNP administered at 0.7 mg / kg, and by the end of six cycles, the patient achieved a partial response with complete regression of lymph node metastases and substantial shrinkage of the liver tumor. This patient achieved a complete response after 40 doses, and the patient remained in remission and completed treatment after 26 months of administration. Two patients with RCC and extrahepatic disease sites including kidney, lung, and lymph nodes that had progressed after previous treatment with a VEGF pathway inhibitor had stable disease at all sites for approximately 8 to 12 months, and a patient with PNET and liver metastases continued on the extension study for 18 months (36 doses) with stable disease.
[0287] However, the charge of LNPs must be considered. Cationic lipids have been combined with negatively charged lipids to induce a non-bilayer structure that promotes intracellular delivery. Because charged LNPs are rapidly removed from the circulation after intravenous injection, ionizable cationic lipids with pKa values less than 7 have been developed (see, for example, Rosin et al., Molecular Therapy, vol. 19, no. 12, pages 1286-2200, December 2011). Negatively charged polymers, such as siRNA oligonucleotides, can be loaded onto LNPs at low pH values (e.g., pH 4), where ionizable lipids exhibit a positive charge. However, at physiological pH values, LNPs exhibit a low surface charge, which is compatible with longer circulation times. Four ionizable cationic lipids were focused on: 1,2-dilinoleyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxy-keto-N,N-dimethyl-3-aminopropane (DLinKDMA), and 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA). These lipid-containing LNP-siRNA systems have been shown to exhibit significantly different gene silencing properties in hepatocytes in vivo, with efficacy varying according to the following sequence using a Factor VII gene silencing model: DLinKC2-DMA > DLinKDMA > DLinDMA >> DLinDAP (see, for example, Rosin et al., Molecular Therapy, Vol. 19, No. 12, pages 1286-2200, December 2011). For formulations containing DLinKC2-DMA in particular, dosages of 1 μg / ml may be contemplated. LNP preparation and CRISPR Cas encapsulation may be used and / or adapted from Rosin et al., Molecular Therapy, Vol. 19, No. 12, pages 1286-2200, December 2011).Cationic lipids 1,2-dilinoleyl-3-dimethylammonium-propane (DLinDAP), 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (DLinDMA), 1,2-dilinoleyloxyketo-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dilinoleyl-4-(2-dimethylaminoethyl)-[1,3]-dioxolane (DLinKC2-DMA), (3-o-[2''-(methoxypolyethylene glycol 2000) succinoyl]-1,2-dimyristoyl-sn-glycol (PEG-S-DMG), and R-3-[(w-methoxy-poly(ethylene glycol) 2000) carbamoyl]-1,2-dimyristoylpropyl-3-amine (PEG-C-DOMG) were purchased from Tekmira. Cholesterol can be provided by Sigma Pharmaceuticals (Vancouver, Canada) or synthesized. Cholesterol can be purchased from Sigma (St. Louis, Mo.). Specific CRISPR Cas RNAs can be encapsulated in LNPs containing DLinDAP, DLinDMA, DLinK-DMA, and DLinKC2-DMA (cationic lipid:DSPC:CHOL:PEG-DMG or PEG-C-DOMG, molar ratio 40:10:40:10). If necessary, 0.2% SP-DiOC18 (Invitrogen, Burlington, Canada) may be incorporated to assess cellular uptake, intracellular delivery, and biodistribution. Encapsulation may be performed by dissolving a lipid mixture composed of cationic lipid:DSPC:cholesterol:PEG-c-DOMG (40:10:40:10 molar ratio) in ethanol to a final lipid concentration of 10 mmol / L. Multilamellar vesicles may be formed by adding this ethanolic solution of lipids dropwise to 50 mmol / L citric acid, pH 4.0, to produce a final concentration of 30% ethanol (vol / vol). Large unilamellar vesicles may be formed after extruding the multilamellar vesicles through two stacked 80 nm Nuclepore polycarbonate filters using an extruder (Northern Lipids, Vancouver, Canada).Encapsulation can be achieved by adding 2 mg / ml of RNA dissolved in 50 mmol / l citric acid, pH 4.0, containing 30% ethanol (vol / vol) dropwise to the extruded preformed large unilamellar vesicles and incubating them at 31°C for 30 min with constant mixing to a final RNA / lipid weight ratio of 0.06 / 1 (wt / wt). Removal of ethanol and neutralization of the formulation buffer were performed by dialysis against phosphate-buffered saline (PBS) at pH 7.4 for 16 h using a Spectra / Por2 regenerated cellulose dialysis membrane. Nanoparticle size distribution can be determined by dynamic light scattering using a NICOMP 370 particle sizer, vesicle / intensity mode, and Gaussian fitting (Nicomp Particle Sizing, Santa Barbara, Calif.). The particle size of all three LNP systems can be approximately 70 nm in diameter. The siRNA encapsulation efficiency can be determined by removing free siRNA from samples taken before and after dialysis using a VivaPureD MiniH column (Sartorius Stedim Biotech). Encapsulated RNA can be extracted from the eluted nanoparticles and quantified at 260 nm. The cholesterol E enzyme assay from Wako Chemicals USA (Richmond, VA) was used to determine the siRNA-to-lipid ratio by measuring the cholesterol content in the vesicles. PEGylated liposomes (or LNPs) can also be used for delivery.
[0288] The preparation of large LNPs can be used and / or adapted from Rosin et al., Molecular Therapy, vol. 19, no. 12, pages 1286-2200, December 2011. A lipid premix solution (20.4 mg / ml total lipid concentration) can be prepared in ethanol containing DLinKC2-DMA, DSPC, and cholesterol in a molar ratio of 50:10:38.5. Sodium acetate can be added to the lipid premix at a molar ratio of 0.75:1 (sodium acetate:DLinKC2-DMA). The mixture can then be combined with 1.85 volumes of citrate buffer (10 mmol / L, pH 3.0) under vigorous stirring to hydrate the lipids, allowing them to spontaneously form liposomes in an aqueous buffer containing 35% ethanol. The liposome solution can be incubated at 37°C to allow for time-dependent particle size growth. Aliquots can be removed at various times during incubation to examine changes in liposome size by dynamic light scattering (Zetasizer Nano ZS, Malvern Instruments, Worcestershire, UK). Once the desired particle size is achieved, an aqueous PEG-lipid solution (stock = 10 mg / ml PEG-DMG in 35% (vol / vol) ethanol) can be added to the liposome mixture to obtain a final PEG molar concentration of 3.5% of total lipid. Upon addition of the PEG-lipid, the liposomes should retain their size and effectively halt further growth. RNA can then be added to the empty liposomes at an approximately 1:10 (wt:wt) siRNA-to-total lipid ratio, followed by a 30-minute incubation at 37°C to form loaded LNPs. The mixture can then be dialyzed overnight in PBS and filtered through a 0.45 μm syringe filter.
[0289] Spherical Nucleic Acid (SNA™) constructs and other nanoparticles (particularly gold nanoparticles) are also contemplated as a means of delivering CRISPR / Cas systems to intended targets. Significant data demonstrate that AuraSense Therapeutics' Spherical Nucleic Acid (SNA™) constructs, based on nucleic acid-functionalized gold nanoparticles, are superior to alternative platforms based on several critical success factors:
[0290] High in vivo stability: Due to their high density loading, the majority of the cargo (DNA or siRNA) remains bound to the construct inside the cell, conferring nucleic acid stability and resistance to enzymatic degradation.
[0291] Deliverability: For all cell types studied (e.g., neurons, tumor cell lines, etc.), the construct exhibits 99% transfection efficiency without the need for carriers or transfection agents.
[0292] Therapeutic Targeting: The unique target binding affinity and specificity of the construct allows for excellent specificity (i.e., limited off-target effects) for the matched target sequence.
[0293] Superior efficacy: The construct is significantly superior to the leading conventional transfection reagents (Lipofectamine 2000 and Cytofectin).
[0294] Low toxicity: The constructs can enter a variety of cultured cells, primary cells, and tissues without apparent toxicity.
[0295] No significant immune response. The construct induces minimal changes in global gene expression as measured by whole-genome microarray studies and cytokine-specific protein assays.
[0296] Chemical Adaptability: Any number of single or combinatorial agents (e.g., proteins, peptides, small molecules) can be used to adapt the surface of the construct.
[0297] This platform for nucleic acid-based therapeutics may be applicable to numerous disease states, including inflammation and infectious diseases, cancer, skin disorders and cardiovascular disease.
[0298] Citations include: Cutler et al., J. Am. Chem. Soc. 2011 133:9254-9257; Hao et al., Small. 2011 7:3158-3162; Zhang et al., ACS Nano. 2011 5:6962-6970; Cutler et al., J. Am. Chem. Soc. 2012 134:1376-1391; Young et al., Nano Lett. 2012 12:3867-71; Zheng et al., Proc. Natl. Acad. Sci. USA. 2012 109:11975-80; Mirkin, Nanomedicine 2012 7:635-638; Zhang et al., J. Am. Chem. Soc. 2012 134:16488-1691, Weintraub, Nature 2013 495:S14-S16, Choi et al., Proc. Natl. Acad. Sci. USA. 2013 110(19):7625-7630, Jensen et al., Sci. Transl. Med. 5,209ra152(2013) and Mirkin, et al., Small, doi.org / 10.1002 / smll.201302143.
[0299] Self-assembling nanoparticles bearing siRNA can be constructed, for example, as a means to target integrin-expressing tumor angiogenesis using polyethyleneimine (PEI) PEGylated with an Arg-Gly-Asp (RGD) peptide ligand attached to the distal end of polyethylene glycol (PEG). These nanoparticles are used to deliver siRNA inhibiting vascular endothelial growth factor receptor-2 (VEGF R2) expression, thereby resulting in tumor angiogenesis (e.g., Schiffelers et al., Nucleic Acids Research, 2004, Vol. 32, No. 19). Nanoplexes can be prepared by mixing equal amounts of aqueous solutions of cationic polymer and nucleic acid to achieve a net molar excess of ionizable nitrogen (polymer) to phosphate (nucleic acid) ranging from 2 to 6. Electrostatic interactions between the cationic polymer and nucleic acid result in the formation of polyplexes with an average particle size distribution of approximately 100 nm, and are therefore referred to herein as nanoplexes. A dose of approximately 100-200 mg of CRISPR Cas is envisioned for delivery in Schiffelers et al.'s self-assembling nanoparticles.
[0300] The nanoplexes of Bartlett et al. (PNAS, September 25, 2007, Vol. 104, No. 39) may be applied to the present invention. Bartlett et al.'s nanoplexes can be prepared by mixing equal amounts of aqueous solutions of cationic polymer and nucleic acid to obtain a net molar excess of ionizable nitrogen (polymer) to phosphate (nucleic acid) ranging from 2 to 6. The electrostatic interaction between the cationic polymer and nucleic acid results in the formation of polyplexes with an average particle size distribution of approximately 100 nm, and are therefore referred to herein as nanoplexes. Bartlett et al.'s DOTA-siRNA was synthesized as follows: 1,4,7,10-tetraazacyclododecane-1,4,7,10-tetraacetic acid mono(N-hydroxysuccinimide ester) (DOTA-NHSester) was ordered from Macrocyclics (Dallas, Texas). A 100-fold molar excess of DOTA-NHS-ester in carbonate buffer (pH 9) was added to a microcentrifuge tube. The contents were allowed to react by stirring at room temperature for 4 hours. The DOTA-RNA sense conjugate was ethanol precipitated, resuspended in water, and annealed to the unmodified antisense strand to obtain DOTA-siRNA. All liquids were pretreated with Chelex-100 (Bio-Rad, Hercules, Calif.) to remove trace metal contaminants. Tf-targeted and non-targeted siRNA nanoparticles can be formed using cyclodextrin-containing polycations. Typically, nanoparticles were formed in water at a charge ratio of 3 (+ / -) and an siRNA concentration of 0.5 g / liter. 1% of the adamantane-PEG molecules on the surface of the targeted nanoparticles were modified with Tf (adamantane-PEG-Tf). The nanoparticles were suspended in a 5% (wt / vol) glucose carrier solution for injection.
[0301] Davis et al. (Nature, Vol. 464, 15 Apr. 2010) conducted a clinical trial of siRNA using a targeted nanoparticle delivery system (clinical trial registration number NCT00689065). Patients with solid tumors refractory to standard treatment received a dose of targeted nanoparticles via 30-minute intravenous infusion on days 1, 3, 8, and 10 of a 21-day cycle. The nanoparticles consisted of a synthetic delivery system containing: (1) a linear cyclodextrin-based polymer (CDP); (2) a human transferrin protein (TF) targeting ligand displayed on the exterior of the nanoparticle to engage with the TF receptor (TFR) on the surface of cancer cells; (3) a hydrophilic polymer (polyethylene glycol (PEG) used to promote nanoparticle stability in biological fluids); and (4) an siRNA designed to reduce the expression of RRM2 (the sequence used in clinical trials was previously designated siR2B+5). TFR has long been known to be upregulated in malignant cells, and RRM2 is an established anticancer target. These nanoparticles (clinical version designated CALAA-01) have been shown to be well tolerated in a multi-dose study in non-human primates. Although one patient with chronic myeloid leukemia received siRNA via liposomal delivery, Davis et al.'s clinical trial is the first human trial to use a targeted delivery system to deliver siRNA systemically to treat patients with solid tumors. To determine whether a targeted delivery system could provide effective delivery of functional siRNA to human tumors, Davis et al. examined biopsies from three patients from three different dosing cohorts. Patients A, B, and C, all with metastatic melanoma, received 18, 24, and 30 mg m -2The subjects received a CALAA-01 dose of siRNA. Similar doses may be contemplated for the CRISPR-Cas system of the present invention. Delivery of the present invention may be achieved using nanoparticles containing linear cyclodextrin-based polymers (CDPs), human transferrin protein (TF) targeting ligands displayed on the exterior of the nanoparticles to engage with TF receptors (TFRs) on the surface of cancer cells, and / or hydrophilic polymers (e.g., polyethylene glycol (PEG) used to promote nanoparticle stability in biological fluids).
[0302] Delivery or administration according to the present invention can be carried out using liposomes. Liposomes are spherical vesicular structures composed of a single or multiple lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes have attracted considerable attention as drug delivery carriers because they are biocompatible and non-toxic, can deliver both hydrophilic and lipophilic drug molecules, protect their cargo from degradation by plasma enzymes, and transport their payload across biological membranes and the blood-brain barrier (BBB) (for a review, see, for example, Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679).
[0303] Liposomes can be made from several different types of lipids, but phospholipids are most commonly used to produce liposomes as drug carriers.Liposome formation occurs spontaneously when lipid membrane is mixed with aqueous solution, but can also be promoted by applying force in the form of shaking, using a homogenizer, ultrasonicator or extrusion device (for review, see, for example, Spuch and Navarro, Journal of Drug Delivery, vol.2011, Article ID 469679, 12 pages, 2011.doi:10.1155 / 2011 / 469679).
[0304] Some other additives may be added to liposomes to modify their structure and properties.For example, either cholesterol or sphingomyelin may be added to the liposome mixture to help stabilize the liposome structure and prevent leakage of the liposome's internal cargo.In addition, liposomes have been prepared from hydrogenated egg phosphatidylcholine or egg phosphatidylcholine, cholesterol, and dicetyl phosphate, and their average vesicle size has been adjusted to about 50 and 100 nm (for review, see, for example, Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679).
[0305] Conventional liposome formulations are primarily composed of natural phospholipids and lipids, such as 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine, and monosialogangliosides. Because these formulations are composed solely of phospholipids, liposome formulations face many challenges, one of which is instability in plasma. Several attempts have been made to overcome these challenges, specifically in the manipulation of lipid membranes. One of these attempts focused on the manipulation of cholesterol. The addition of cholesterol to conventional formulations reduces the rapid release of encapsulated bioactive compounds into the plasma, or 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) increases stability (for a review, see, e.g., Spuch and Navarro, Journal of Drug Delivery, vol. 2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679).
[0306] In particularly advantageous embodiments, Trojan Horse liposomes (also known as Molecular Trojan Horses) are desirable; protocols can be found at cshprotocols.cshlp.org / content / 2010 / 4 / pdb.prot5407.long. These particles enable transgene delivery to the entire brain after intravascular injection. Neutral lipid particles with specific antibodies conjugated to their surface are believed to enable crossing of the blood-brain barrier via endocytosis. The applicant claims to utilize Trojan Horse liposomes to deliver the CRISPR family of nucleases to the brain via intravascular injection, enabling whole-brain transgenic animals without the need for embryonic manipulation. Approximately 1 to 5 g of nucleic acid molecules, such as DNA and RNA, can be contemplated for in vivo administration in liposomes.
[0307] In another embodiment, the system can be administered in a liposome, such as a stable nucleic acid-lipid particle (SNALP) (see, e.g., Morrissey et al., Nature Biotechnology, Vol. 23, No. 8, August 2005). Daily intravenous injection of about 1, 3, or 5 mg / kg / day of the specific CRISPR Cas targeted in the SNALP is contemplated. Daily treatment may be administered for about 3 days, then weekly for about 5 weeks. In another embodiment, a specific CRISPR Cas-encapsulated SNALP administered by intravenous injection at a dose of 1 or 2.5 mg / kg is also contemplated (see, e.g., Zimmerman et al., Nature Letters, Vol. 441, 4 May 2006). SNALP formulations can contain the lipids 3-N-[(w-methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristyloxy-propylamine (PEG-C-DMA), 1,2-dilinoleyloxy-N,N-dimethyl-3-aminopropane (DLinDMA), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC), and cholesterol in a molar percentage ratio of 2:40:10:48 (see, e.g., Zimmerman et al., Nature Letters, Vol. 441, 4 May 2006).
[0308] In another embodiment, stable nucleic acid-lipid particles (SNALP) have been shown to be effective delivery molecules for highly vascularized HepG2-derived liver tumors, but not for poorly vascularized HCT-116-derived liver tumors (see, e.g., Li, Gene Therapy (2012) 19, 775-780). SNALP liposomes can be prepared by formulating D-Lin-DMA and PEG-C-DMA with distearoylphosphatidylcholine (DSPC), cholesterol, and siRNA using a lipid / siRNA ratio of 25:1 and a cholesterol / D-Lin-DMA / DSPC / PEG-C-DMA molar ratio of 48 / 40 / 10 / 2. The resulting SNALP liposomes are approximately 80-100 nm in size.
[0309] In yet another embodiment, the SNALP can comprise synthetic cholesterol (Sigma-Aldrich, St. Louis, Mo., USA), dipalmitoylphosphatidylcholine (Avanti Polar Lipids, Alabaster, Ala., USA), 3-N-[(w-methoxypoly(ethylene glycol)2000)carbamoyl]-1,2-dimyristyloxypropylamine, and cationic 1,2-dilinoleyloxy-3-N,N-dimethylaminopropane (see, for example, Geisbert et al., Lancet 2010;375:1896-905). For example, a total CRISPR Cas dosage of about 2 mg / kg can be contemplated per dose administered as a bolus intravenous infusion.
[0310] In yet another embodiment, the SNALP may contain synthetic cholesterol (Sigma-Aldrich), 1,2-distearoyl-sn-glycero-3-phosphocholine (DSPC; Avanti Polar Lipids Inc.), PEG-cDMA, and 1,2-dilinoleyloxy-3-(N;N-dimethyl)aminopropane (DLinDMA) (see, e.g., Judge, J. Clin. Invest. 119:661-673 (2009)). Formulations used for in vivo studies may contain a final lipid / RNA mass ratio of about 9:1.
[0311] Other cationic lipids, such as the amino lipid 2,2-dilinoleyl-4-dimethylaminoethyl-[1,3]-dioxolane (DLin-KC2-DMA), can be used to encapsulate CRISPR Cas as well as siRNA (see, e.g., Jayaraman, Angew. Chem. Int. Ed. 2012, 51, 8529-8533). Preformed vesicles with the following lipid composition can also be contemplated: amino lipid, distearoylphosphatidylcholine (DSPC), cholesterol, and (R)-2,3-bis(octadecyloxy)propyl-1-(methoxypoly(ethylene glycol)2000)propylcarbamate (PEG-lipid) in a molar ratio of 40 / 10 / 40 / 10, respectively, and a FVII siRNA / total lipid ratio of approximately 0.05 (w / w). To ensure a narrow particle size distribution in the 70-90 nm range and a low polydispersity index of 0.11-0.04 (n = 56), particles can be extruded up to three times through an 80 nm membrane before adding CRISPR Cas RNA. Particles containing highly potent amino lipids 16 can be used, and the molar ratio of the four lipid components 16, DSPC, cholesterol, and PEG-lipid (50 / 10 / 38.5 / 1.5), can be further optimized to enhance in vivo activity.
[0312] Any suitable CRISPR / Cas gene editing system known in the art can be used in the systems and methods described herein as needed.CRISPR / Cas gene editing technology is described, for example, in U.S. Patent Application Publication No. 2014 / 0068797; U.S. Patent No. 8,697,359; U.S. Patent No. 8,771,945; and U.S. Patent No. 8,945,839; U.S. Patent No. US2010 / 0076057; U.S. Patent No. US2011 / 0189776; U.S. Patent No. US2011 / 0223638; U.S. Patent No. US2013 / 0130248; WO / 2008 / 108989; WO / 2010 / 054108; WO / 2012 / 164565;WO / 2013 / 098244;WO / 2013 / 176772;US20150050699;US20150045546;US20150031134;US20150024500;US201403778 68;US20140357530;US20140349400;US20140335620;US20140335063;US20140315985;US20140310830;US20140310828;US2 0140309487;US20140304853;US20140298547;US20140295556;US20140294773;US20140287938;US20140273234;US2014027 3232;US20140273231;US20140273230;US20140271987;US20140256046;US20140248702;US20140242702;US20140242700;U S20140242699;US20140242664;US20140234972;US20140227787;US20140212869;US20140201857;US20140199767;US20140189896;US20140186958;US20140186919;US20140186843;US20140179770;US20140179006; and US20140170753.
[0313] Although the present invention and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations could be made therein without departing from the spirit and scope of the invention as defined in the appended claims.
[0314] The present invention is further described in the following examples, which are given for purposes of illustration only and are not intended to limit the invention in any way. [Example]
[0315] Example 1 - Materials and Methods RecE / T homology screening The RefSeq non-redundant protein database was downloaded from NCBI on October 29, 2019. Position-specific repeat (PSI)-BLAST was used to search for protein homologs. 1 The database was searched using the E. coli Rac prophage RecT (NP_415865.1) and RecE (NP_415866.1) as queries using . Hits were clustered with CD-HIT2 and representative sequences were identified using MUSCLE. 3 From each cluster, a diverse set of RecET homologs was selected for multiple alignment with . FastTree4 was then used for maximum likelihood tree reconstruction with default parameters. A diverse set of RecET homologs was selected, synthesized by GenScript, and cloned into the pMPH_MCP vector for testing.
[0316] Plasmid construction: pX330, pMPH, and pU6-(BbsI)_CBh-Cas9-T2A-BFP plasmids were obtained from Addgene. Tested effector DNA fragments were ordered from IDT, Genewiz, and GenScript. The fragments were Gibson assembled into the backbone using NEBuilder HiFi DNA Assembly Master Mix (New England BioLabs). All sgRNAs (Table 3) were inserted into the backbone using Golden Gate cloning. All constructs were sequence-confirmed by Sanger sequencing of the prepared plasmids. [Table 3]
[0317] Cell cultures Human Embryonic Kidney (HEK) 293T, HeLa, and HepG2 were maintained in Dulbecco's Modified Eagle Medium (DMEM, Life Technologies) containing 10% fetal bovine serum (FBS, HyClone), 100 U / mL penicillin, and 100 μg / mL streptomycin (Life Technologies) at 37°C and 5% CO .
[0318] hES-H9 cells were maintained in mTeSR1 medium (StemCell Technologies) at 37°C and 5% CO. Culture plates were pre-coated with Matrigel (Corning) 12 hours before use, and cells were supplemented with 10 μM Y27632 (Sigma) for the first 24 hours after passaging. Culture medium was changed every 24 hours.
[0319] Transfection: HEK293T cells were seeded into 96-well plates (Corning) at a density of 30,000 cells / well 12–24 hours prior to transfection and transfected with 250 ng of total DNA per well. One day prior to transfection, HeLa and HepG2 cells were seeded into 48-well plates (Corning) at a density of 50,000 and 30,000 cells / well, respectively, and transfected with 400 ng of total DNA per well. Transfection was performed using Lipofectamine 3000 (Life Technologies) according to the manufacturer's instructions.
[0320] For electroporation experiments involving hES-H9 cells, the P3 Primary Cell 4D-Nucleofector™ X Kit S (Lonza) was used according to the manufacturer's protocol. For each reaction, 300,000 cells were nucleofected with 4 μg of total DNA using the DC100 Nucleofector Program.
[0321] Fluorescence-activated cell sorting (FACS) mKate knock-in efficiency was analyzed using a CytoFLEX flow cytometer (Beckman Coulter; Stanford Stem Cell FACS Core). Seventy-two hours after transfection, cells were washed once with PBS and dissociated with TrypLE Express Enzyme (Thermo Fisher Scientific). The cell suspension was then transferred to a 96-well U-bottom plate (Thermo Fisher Scientific) and centrifuged at 300xG for 5 minutes. After removing the supernatant, the pelleted cells were resuspended in 50 μl of 4% FBS in PBS, and the cells were sorted within 30 minutes of preparation.
[0322] RFLP HEK293T cells were transfected with plasmid DNA and PCR template, and genomic DNA was harvested 72 hours later using QuickExtract DNA Extraction Solution (Biosearch Technologies) according to the manufacturer's protocol. Specific primers outside the homologous arms of the PCR template were used to amplify the target genomic region. PCR products were purified with the Monarch PCR & DNA Cleanup Kit (New England BioLabs). 300 ng of purified product was digested with BsrGI (EMX1, New England BioLabs) or XbaI (VEGFA, NEB), and the digested products were analyzed on a 5% Mini-PROTEAN TBE gel (Bio-Rad).
[0323] Next-generation sequencing library preparation. Genomic DNA was extracted 72 hours after transfection using QuickExtract DNA Extraction Solution (Biosearch Technologies). 200 ng of total DNA was used for NGS library preparation. Specific primers (Table 4) were used for the first PCR reaction to amplify the gene of interest. Illumina adapters and index barcodes were added to the fragments by a second PCR using the primers listed in Table 4. The products of the second PCR were purified by gel electrophoresis on a 2% agarose gel using the Monarch DNA Gel Extraction Kit (NEB). The purified products were quantified using the Qubit dsDNA HS Assay Kit (Thermo Fisher) and sequenced on an Illumina MiSeq according to the manufacturer's instructions. [Table 4-1] [Table 4-2] [Table 4-3]
[0324] High-throughput sequencing data analysis: Analyze the processed (demultiplexed, trimmed, and merged) sequencing reads by aligning the sequenced amplicons against the reference and expected HDR amplicons using CRISPResso2. 5 Editing outcomes were determined using the quantification window. To better capture diverse editing outcomes, the quantification window was extended to 10 bp around the expected cleavage site, but substitutions were ignored to avoid including sequencing errors. Only reads that did not contain mismatches with the expected amplicon were considered for HDR quantification, and reads containing indels that partially matched the expected amplicon were included at the overall reported indel frequency.
[0325] Statistical Analysis Unless otherwise specified, all statistical analyses and comparisons were performed using t-tests with a false discovery rate (FDR) of 1% using the two-step step-up method of Benjamini, Krieger, and Yekutieli (Benjamini, Y. et al., Biometrika 93, 491-507 (2006), incorporated herein by reference). All experiments were performed in triplicate to ensure sufficient statistical power in the analysis unless otherwise specified.
[0326] Determination of editing at predicted Cas9 off-target sites. To assess RecT / RecE off-target editing activity at known Cas9 off-target sites, the same genomic DNA extracts for knock-in analysis were used as templates for PCR amplification of the top predicted off-target sites (highly scored as predicted by CRISPOR, a web-based analysis tool) for EMX1, a VEGFA guide; primer sequences are listed in Table 4.
[0327] iGUIDE Off-Target Analysis Based on the previously developed Guide-seq (Tsai, S. et al., Nat Biotechnol 33, 187-197 (2015), incorporated herein by reference), genome-wide unbiased off-target analysis was performed according to the iGUIDE pipeline (Nobles, CL et al., Genome Biol 20, 14 (2019), incorporated herein by reference). HEK293T cells were transfected in 20 μL of Lonza SF Cell Line Nucleofector Solution on a Lonza Nucleofector 4-D using program DS-150 according to the manufacturer's instructions. 300 ng of gRNA-Cas9 plasmid (or 150 ng of each gRNACas9n plasmid for dual nickases), 150 ng of effector plasmid, and 5 pmol of double-stranded oligonucleotide (dsODN) were transfected. Cells were harvested 72 hours later for genomic DNA using the Agencourt DNAdvance Reagent Kit. 400 ng of purified gDNA was then fragmented to an average of 500 bp and ligated with adapters using the NEBNext Ultra II FS DNA Library Prep Kit according to the manufacturer's instructions. Two rounds of nested anchored PCR were performed from the oligo tags to the ligated adapter sequences to amplify the target DNA. The amplified libraries were purified, size-selected, and sequenced using an Illumina Miseq V2 PE300. Sequencing data were analyzed using the published iGUIDE pipeline, with an additional downsampling step to ensure unbiased comparisons across samples. Example 2
[0328] In contrast to mammals, convenient recombineering editing tools are available for bacteria, such as phage lambda Red and RecE / T. Microbial recombineering involves two major steps: template DNA is chewed back by an exonuclease (Exo), followed by single-strand annealing proteins (SSAPs) that support template-directed homologous recombination repair, optionally facilitated by nuclease inhibitors. We developed a system for RNA-guided targeting of RecE / T recombineering activity and achieved kilobase (kb) human gene editing without DNA cleavage.
[0329] We investigated candidate microbial systems with recombineering activity. Two lines of reasoning guided the search: 1) orthogonality: prioritize proteins with minimal similarity to mammalian repair enzymes; and 2) parsimony: focus on systems with minimal interdependence of components. Three protein families were identified: lambda Red, RecE / T, and the phage T7 gp6 (Exo) and gp2.5 (SSAP) recombination machinery. Based on phylogenetic reconstruction, the RecE / T protein was determined to be the most distant from eukaryotic recombination proteins and one of the most compact (Figure 1). Therefore, we utilized the RecE / T system for downstream analysis.
[0330] We systematically searched the NCBI protein database for RecE / T homologs. To develop a portable tool, we examined their evolutionary relationships and length (Figure 2A). Co-occurrence analysis revealed that most RecE / T systems possess only one of the two proteins (Figure 2B). Because prophage integration can be imprecise, we prioritized the 11% of species that possess both homologs as evidence for intact functionality.
[0331] The top 12 candidates were codon-optimized, and MS2 coat protein (MCP) fusions were constructed to recruit these RecE / T homologs (hereafter referred to as "recombinators") to wild-type Streptococcus pyogenes Cas9 (wtCas9) via the MS2 RNA aptamer. To understand their individual molecular effects as Exo and SSAP, each was tested independently (Figure 2C). Initial results revealed the Escherichia coli RecE / T proteins (simplified as RecE and RecT) as promising candidates, as determined by genome knock-in assays (Figure 2D). While RecT is only 269 amino acids (AA) long, RecE was truncated from AA587 (RecE_587) and the carboxy-terminal domain (RecE_CTD) based on functional studies (Muyrers, J.P., Genes Dev. (2000); 14, 1971-1982, incorporated herein by reference).
[0332] To validate RecE / T recombineering in human cells, we measured homology-directed repair (HDR) at five genomic sites using two templates. While RecE variants (RecE_587, RecE_CTD) showed increased variability in knock-in efficiency, RecT significantly enhanced HDR in all cases, replacing approximately 16 bp sequences of EMX1 and VEGFA and knocking in approximately 1 kb cassettes of HSP90AA1, DYNLT1, and AAVS1 (Figure 3A-E, Figure 4). We validated these results using imaging (Figure 3F) and sequenced the junction sites using Sanger sequencing to confirm accurate insertion (Figure 3G). To test whether these activities were truly sequence-specific, we used a non-recruitment control using PP7 coat protein (PCP), which recognizes the PP7 aptamer but not the MS2 aptamer. RecE was active without recruitment, whereas RecT showed increased efficiency in a recruitment-dependent manner (Figure 3H). Without being bound by theory, this may be explained by the promiscuous exonuclease activity of RecE (Figure 2C). The RecE / T recombineering editing (REDIT) tool is called REDITv1, and REDITv1_RecT is the preferred variant. Example 3
[0333] Three tests were performed on REDITv1 to investigate 1) activity across cell types, 2) optimal HDR template design, and 3) specificity. REDITv1 activity was robust across multiple genomic sites in HEK, A549, HepG2, and HeLa cells (Figures 5A-5C, 6A-6C). Notably, in human embryonic stem cells (hESCs), REDITv1 consistently increased kilobase knock-in efficiency in HSP90AA1 and OCT4, demonstrating up to a 3.5-fold improvement compared to Cas9-HDR (Figures 5D-5E, 6D-6E). Different template designs were also tested. REDITv1 performed efficient kilobase editing using HA lengths as short as 200 bp in total, with longer HA supporting higher efficiency. This achieved up to 10% efficiency (without selection) for kb-scale knock-in, a 5-fold increase over Cas9-HDR and significantly higher than the typical efficiency of 1-2% (Figure 7). Finally, we determined the accuracy of REDITv1 using deep sequencing of predicted off-target sites (OTS) and GUIDE-seq. REDITv1 did not improve off-target effects, but detectable OTS remained at sites previously reported for EMX1 and VEGFA (Figures 5F-5G, Figure 8). Briefly, REDITv1 demonstrated kilobase-scale genome recombineering but maintained off-target issues, with REDITv1_RecT being the most efficient. Example 4
[0334] To mitigate undesired editing, we evaluated a version of REDIT with a non-cleaving Cas9 nickase (Cas9n). A similar strategy was previously employed to address off-target issues (Ran, FA et al., Cell (2013), 154:1380-1389, incorporated herein by reference), but HDR efficiency was low. REDIT was tested to determine whether this system could overcome the limitations of endogenous repair and promote nicking-mediated recombination. Indeed, the nickase version demonstrated higher efficiency, with the best results obtained from Cas9n(D10A) with single and double nicking. This Cas9n(D10A) variant was designated REDITv2N (Figure 9A). A 5%-10% knock-in rate without selection was observed using REDITv2N with double nicking, which is comparable to REDITv1 using wtCas9 (Figure 9A, Figure 10A). Junction sequencing confirmed the accuracy of knock-in for all targets (Figure 11). This result represented a 6- to 10-fold improvement over Cas9n-HDR. Even with single-nicking REDITv2N, we observed an efficiency of approximately 2% for 1 kb knock-in, significantly higher than the 0.46% HDR efficiency reported previously using standard single-nicking Cas9n and a less challenging 12-bp knock-in template (Cong, L. et al., Science. 339, 819-823, incorporated herein by reference) (Figure 9A).
[0335] We investigated the off-target activity of REDITv2N using GUIDE-seq. Results showed minimal off-target cleavage and a 90% reduction in OTS compared to REDITv1 (Figure 9B). Specifically, for the DYNLT1-targeting guide, the most abundant KIF6 OTS was significantly enriched in the REDITv1 group but disappeared when REDITv2N was used (Figure 9C). REDITv2N was highly accurate (Figures 9B-9C, Figure 12).
[0336] Another byproduct of HDR editing is on-target insertion-deletions (indels). They can dramatically reduce gene editing yields, especially for long sequences. We measured indel formation in an EMX1 knock-in experiment using deep sequencing. REDITv2N increased HDR efficiency to the same level as its wtCas9 counterpart (Figure 12C, top) and reduced unwanted on-target indels by 92% (Figure 12C, bottom).
[0337] Using concepts from GUIDE-seq, LAM-PCR, and TLA, we developed an NGS-based assay for identifying genome-wide insertion sites (GIS), or GIS-seq (Figure 30A). Using GIS-seq, we obtained NGS read clusters / peaks representing knock-in insertion sites (Figure 30B), with representative reads from on-target sites shown. GIS-seq was applied to the DYNLT1 and ACTB loci to measure knock-in accuracy. Sequencing results showed that REDIT identified fewer off-target insertion sites compared to Cas9, considering sites with high confidence based on maximum likelihood estimation (Figure 30C). Together, clonal Sanger sequencing of knock-in junctions (Figures 9C and 12), GUIDE-seq analysis (Figure 9B), and GIS-seq results (Figures 30A-30C) demonstrated that REDIT can be an efficient method capable of inserting kilobase-long sequences with few unwanted editing events. Example 5
[0338] REDIT was investigated for its ability to edit long sequences without any nicking / cleavage of the target DNA. Notably, when REDITv2D was constructed using catalytically inactive Cas9 (dCas9), precise genome knock-in of kilobase cassettes was observed in human cells (Figure 9D, top, Figure 13). While REDITv2D was less efficient than REDITv2N, it achieved programmable DNA damage-free editing at the kilobase scale with 1-2% efficiency and no selection (Figure 9D, Figure 10B). It was hypothesized that two processes may contribute to REDITv2D recombineering. One possibility was dCas9 unwinding. Because dCas9 induces sequence-specific formation of loops, if dCas9 can unwind DNA, dual binding with two dCas9s is expected to promote genome accessibility to RecE / T. However, no significant increase was observed when two guide RNAs were delivered (Figure 9D, bottom). Another possibility is that DNA unwinding during the cell cycle allowed RecE / T to access the target region mediated by dCas9 binding. Using different REDIT tools, we performed a 1-kb knock-in under various serum levels (10% typical, 2% low, and no serum). Because serum starvation inhibits cell proliferation, the results showed that the cell cycle positively correlated with REDITv2D recombineering (Figure 9E). Upon serum-free treatment, HDR efficiency was only reduced in the REDITv2D(dCas9) group, whereas REDITv1(wtCas9) and REDITv2N(D10A) were unaffected (Figure 9E, Figure 14), confirming that DNA unwinding allowed RecE / T to access the target region. Example 6
[0339] Microscopy analysis revealed incomplete nuclear targeting of REDITv1, particularly REDITv1_RecT (Figure 15). Therefore, different designs of protein linker and nuclear localization signal (NLS) were tested (Figure 15A). An extended XTEN-linker with a C-terminal SV40-NLS was identified as the preferred configuration, designated REDITv3 (Figure 16). REDITv3 further achieved a 2-3-fold increase in HDR efficiency over REDITv2 across genomic targets and Cas9 variants (wtCas9, Cas9n, dCas9) (Figure 17).
[0340] Finally, we utilized REDITv3 in hESCs to engineer kilobase knock-in alleles in human stem cells. REDITv3N single-nicking and double-nicking designs resulted in 5-fold and 20-fold increases in HDR efficiency over non-recombinant controls, respectively (Figure 9F). Efficacy and fidelity were confirmed via a combination of assays described for previous REDIT versions (Figures 9F-G, Figure 18). Furthermore, REDITv3 works effectively with Staphylococcus aureus Cas9 (SaCas9), a compact CRISPR system suitable for in vivo delivery (Figure 19). Example 7
[0341] To further investigate the RecT and RecE_587 variants, both RecT and RecE_587 were truncated to various lengths, as shown in Figures 20A and 21A, respectively. The resulting efficiencies were measured using mKate knock-in assays with both wild-type SpCas9 and Cas9n(D10A) along with single and double nicking at the DYNLT1 locus (Figures 20B-20C and 21B-21C, respectively). The efficiency of the no-recombination group is shown as a control.
[0342] Both truncated forms of RecT and RecE_587 retained significant recombineering activity when used with different Cas9s. Notably, compared to full-length RecT (1–269 aa), new truncated forms such as RecT (93–264 aa) are over 30% smaller, yet they essentially preserved the full activity of RecT in stimulating recombination in eukaryotic cells. Similarly, compared to full-length RecE (1–280 aa), truncated forms such as RecE_587 (120–221 aa) and RecE_587 (120–209 aa) are over 60% smaller yet still retained high recombineering activity in human cells. These truncated forms demonstrated the potential for further engineering minimal functional recombineering enzymes using RecE and RecT protein variants, but given their small size, they also provide valuable, compact recombination tools for human genome editing that are ideal for in vitro, ex vivo, and in vivo delivery.
[0343] Overall, REDIT harnesses the specificity of CRISPR genome targeting with the efficiency of RecE / RecT recombineering. The disclosed high-efficiency, low-error system represents a powerful addition to the existing CRISPR toolkit. REDITv3N's balanced efficiency and precision make it an attractive therapeutic option for knocking in large cassettes in immune and stem cells. Example 8
[0344] Phylogenetic trees of the reconstructed RecE and RecT eukaryotic recombinases from yeast and humans (Figure 1A and 1B) show the evolutionary distance of the proteins based on sequence homology. The dotted boxes indicate the full-length E. coli RecB and E. coli RecE proteins. The catalytic core domains of E. coli RecB and E. coli RecE proteins (solid boxes) were used for comparison. The MS2-MCP recruitment system was used to measure the gene editing activity of this recombination protein family. Here, an sgRNA with an MS2 stem loop is used together with a recombineering protein fused to the MCP protein via a peptide linker and a nuclear localization signal.
[0345] Three exonuclease proteins were used: exonuclease from phage lambda, the RecE 587 core domain of the E. coli RecE protein, and exonuclease from phage T7 (gene name gp6) (Figure 22A). Gene editing activity was measured using mKate knock-in assays at genomic loci (DYNLT1 and HSP90AA1).
[0346] Similar measurements were performed by testing the genome editing efficiency of three single-stranded DNA annealing proteins (SSAPs) from the same three microorganisms as the exonucleases: the Bet protein from phage lambda, the RecT protein from E. coli, and the SSAP (gene name gp2.5) from phage T7 (Figure 22B).
[0347] These results systematically measured and validated the genome recombineering activity of all three major families of phage / microbial recombination systems in eukaryotic cells (lambda phage exonuclease and beta protein; E. coli prophage RecE and RecT proteins; T7 phage exonuclease gp6 and single-stranded binding gp2.5 protein). All six proteins from the three systems achieved efficient gene editing to knock-in kilobase-long sequences into mammalian genomes across two genomic loci. Overall, exonucleases demonstrated approximately three-fold higher recombination efficiency (up to 4% mKate genome knock-in) when compared to non-recombinator controls. Single-stranded annealing proteins (SSAPs) demonstrated higher activity, exhibiting 4- to 8-fold higher gene editing activity than the control group. This demonstrated the general applicability and validity that microbial recombination proteins in the exonuclease and SSAP families can be engineered via the Cas9-based fusion protein system to achieve highly efficient genome recombination in mammalian cells. Example 9
[0348] To demonstrate the generalizability of the REDIT protein design, we developed and tested an alternative recruitment system. For a more compact REDIT system, we fused the REDIT recombinator protein to the N22 peptide, and simultaneously incorporated the sgRNA, which contained boxB, a short recognition sequence for the N22 peptide, replacing the MCP within the sgRNA (Figure 23A). This boxB-N22 system demonstrated comparable editing efficiency at the two genomic sites tested, as shown in Figures 23B-23E, in a side-by-side comparison with the MS2-MCP recruitment system.
[0349] A REDIT system was developed that uses SunTag mobilization, a protein-based mobilization system (Figures 24A and 27A). Because SunTag is based on a fusion protein design, the sgRNA or guide RNA is the same as in wild-type CRISPR systems. Specifically, a REDIT recombinator protein was fused to an scFV antibody peptide (replacing MCP), and a GCN4 peptide (10 copies of the GCN4 peptide separated by a linker) was fused in tandem to the Cas9 protein. Thus, scFV-REDIT can be recruited to the Cas9 complex via the affinity of GCN4 for scFV.
[0350] Using mKate knock-in experiments (Figures 24B and 27B), the editing efficiency was measured at the DYNLT1 locus and the HSP90AA1 locus, respectively. This SunTag-based REDIT system showed a significant increase in gene editing knock-in efficiency at the DYNLT1 genomic site tested. Furthermore, the SunTag design significantly increased HRD efficiency, approximately 2-fold better than Cas9, but did not achieve as high an increase as the MS2 aptamer. Example 10
[0351] To demonstrate the generalizability of the REDIT protein design and develop a versatile REDIT system applicable to various CRISPR enzymes, a Cpf1 / Cas12a-based REDIT system using the SunTag mobilization design was developed (Figure 25A). Two different Cpf1 / Cas12a proteins were tested using the mKate knock-in assay (Lachnospiraceae bacterium ND2006, LbCpf1, and Acidaminococcus sp. BV3L6) as previously shown (Figure 25B).
[0352] These results demonstrate that recombination proteins (exonucleases and single-stranded annealing proteins) can be engineered using alternative designs, such as the SunTag mobilization system, to perform genome editing in eukaryotic cells. These protein-based mobilization systems do not require the use of RNA aptamers or RNA-binding proteins; instead, they utilize fusion protein domains that directly connect to CRISPR enzymes to recruit REDIT proteins.
[0353] In addition to the flexibility in mobilization system design, these results using Cpf1 / Cas12a-type CRISPR enzymes also demonstrate the general compatibility of REDIT proteins with various CRISPR systems for genome recombination. Cpf1 / Cas12a enzymes have different catalytic residues and DNA recognition mechanisms than Cas9 enzymes. Therefore, REDIT recombination proteins (exonuclease and single-strand annealing protein) can function independently of the specific selection of CRISPR enzyme components (e.g., Cas9, Cpf1 / Cas12a). This demonstrates the generalizability of the REDIT system and opens the possibility of using additional CRISPR enzymes (known and unknown) as components of the REDIT system to achieve precise genome editing in eukaryotic cells. Example 11
[0354] To screen various RecE and RecT proteins across the microbial kingdom, 15 different species of microorganisms with RecE / RecT proteins were selected (Table 5). Each protein was codon-optimized and synthesized. As previously described for the E. coli RecE / RecT-based REDIT system, each protein was fused to the MCP protein via an E-XTEN linker along with an additional nuclear localization signal. Using the mKate knock-in gene editing assay, efficiency was measured at the DYNLT1 locus (Figure 26A, Table 6) and the HSP90AA1 locus (Figure 26B, Table 6). Homologues demonstrated the ability to enable and facilitate high-precision gene editing. [Table 5-1] [Table 5-2] [Table 6-1] [Table 6-2] Example 12
[0355] Next, we evaluated the RecT-based REDIT design by comparing it with three categories of existing HDR enhancement tools (Figures 28A and 28B): a fusion of the DNA repair enzyme CtIP with Cas9 (Cas9-HE), a fusion of Cas9 with the functional domain (amino acids 1 to 110) of the human geminin protein (Cas9-Gem), and nocodazole, a small molecule enhancer of HDR via cell cycle control. Across the endogenous targets tested, the RecT-based REDIT design outperformed the three alternative strategies (Figure 3C). Furthermore, the RecT-based REDIT design may act independently of other approaches and synergize with existing methods. To test this hypothesis, we combined the RecT-based REDIT design (conveniently via the MS2-aptamer) with three different approaches (Figure 28A, right). Indeed, the RecT-based REDIT design could further enhance the HDR-promoting activity of the tested tools (Figure 28C). Example 13
[0356] We quantified the effect of template HA length on the editing efficiency of REDIT when using a standard HDR donor with at least 100 bp of HA on each side (Figure 29A, left). Higher HDR rates were observed for both Cas9 and RecT groups with increasing HA length, with REDIT stimulating HDR more effectively than Cas9 using HA lengths as short as approximately 100 bp on each side. When fed longer templates with a total HA length of 600–800 bp, RecT achieved HDR efficiencies of over 10% for selection-free kb-scale knock-in, significantly higher than the 2–3% efficiency achieved using Cas9 alone. A recent report identified that using donor DNA with shorter HA (typically 10–50 bp) can significantly stimulate knock-in efficiency, thanks to high repair activity from the microhomology-mediated end-joining (MMEJ) pathway. The knock-in efficiency of the REDIT-based method was compared with that of Cas9 using donor DNA with 0 bp (NHEJ-based), 10 bp, or 50 bp (MMEJ-based) HA. The results revealed that short HA donors utilizing the MMEJ mechanism resulted in higher editing efficiency compared to HDR donors (Figure 29A, right). At the same time, REDIT was able to enhance knock-in efficiency as long as HA was present (no effect on the 0 bp NHEJ donor). This effect was particularly significant with the 10 bp donor, which had a significant effect and was selected for further characterization and comparison with HDR donors.
[0357] Knock-in cells were clonally isolated, and the targeted genomic region was amplified using primers that bound completely outside the donor DNA for colony Sanger sequencing (Figure 29B). Junction sequencing analysis (approximately 48 colonies per gene per condition) revealed varying degrees of indels at the 5'- and 3'-knock-in junctions, including single or both junctions (Figure 29C). Overall, HDR donors had better accuracy than MMEJ donors, and REDIT slightly improved knock-in yields compared to Cas9, although junction indels were still observed.
[0358] We further compared the efficiency of REDIT and Cas9 when editing different lengths. For longer edits, a 2 kb knock-in cassette was used (Figure 29D), and for shorter edits, a single-stranded oligo donor (ssODN) was used. When the knock-in sequence length was extended to approximately 2 kb using a dual mKate / GFP template, REDIT maintained its HDR-promoting activity compared to Cas9 across the endogenous targets tested (Figure 29D). For ssODN testing, REDIT and Cas9 were used to introduce 12-16 bp of exogenous sequence at two well-established loci, EMX1 and VEGFA. Because the ssODN templates were short (<100 bp HA on each side), next-generation sequencing (NGS) was used to quantify editing events. Similar levels of indels were observed between Cas9 and REDIT, indicating improved HDR efficiency using REDIT. Example 14
[0359] The sensitivity of REDIT's ability to promote HDR in the presence or absence of two different pharmacological inhibitors, RAD51, B02 and RI-1 (Figure 31A). As expected, RAD51 inhibition significantly reduced HDR efficiency for Cas9-based editing (Figures 31B, 31C, and 32A). Interestingly, RAD51 inhibition only moderately reduced REDIT and REDITdn efficiency, as both REDIT / REDITdn methods maintained significantly higher knock-in efficiency compared to Cas9 / Cas9dn under RAD51 inhibition.
[0360] We also used Mirin, a potent chemical inhibitor of DSB repair that has been shown to prevent MRN complex formation, MRN-dependent ATM activation, and inhibit Mre11 exonuclease activity. When cells were treated with Mirin, only the editing efficiency of the Cas9 control experiment was affected by Mirining treatment, while the REDIT version was essentially the same as the vehicle-treated group across all genomic targets (Figure 32A).
[0361] To test whether cell cycle inhibition affects recombination, cells were chemically synchronized at the G1 / S boundary using double thymidine blockade (DTB). When RI-1 or B02 milling was combined with DTB treatment, the REDIT version maintained higher editing efficiency under DNA repair pathway inhibition but decreased editing efficiency under DTB treatment compared to the Cas9 reference experiment (Figure 32B).
[0362] To validate REDIT in various contexts, we applied it to human embryonic stem cells (hESCs) to test its ability to manipulate long sequences in non-transformed human cells. Using REDIT and REDITdn, robust stimulation of HDR was observed across all three genomic sites (HSP90AA1, ACTB, and OCT4 / POU5F1) (Figures 31D and 31E). Notably, REDIT and REDITdn editing, using donor DNA with 200 bp of HA on each side, achieved efficiencies of up to 5% for kb-scale gene editing without selection, compared to approximately 1% efficiency using non-REDIT methods. Furthermore, REDIT improved knock-in efficiency in A549 (lung-derived), HepG2 (liver-derived), and HeLa (cervix-derived) cells, demonstrating up to approximately 15% kb-scale genome knock-in without selection. This improvement was up to four-fold higher than the Cas9 group, supporting the feasibility of using the REDIT method in different cell types. Example 15
[0363] The in vivo use of dCas9-EcRecT (SAFE-dCas9) was tested using cleavage-free dCas9 editing via hydrodynamic tail vein injection. The gene editing vector and template DNA used are shown in Figure 33A. The gene editing vector (60 μg) and template DNA (60 μg) were injected via hydrodynamic tail vein injection to deliver the components to mice. The successful gene editing of liver hepatocytes was monitored by the expression of the protein encoded by the transgene from the albumin locus. The experimental procedure is outlined in Figure 33B.
[0364] Approximately seven days after injection, the livers of the perfused mice were removed. The liver lobes were homogenized and processed to extract liver genomic DNA from primary hepatocytes. The extracted genomic DNA was used for three different downstream analyses: 1) PCR using knock-in specific primers and agarose gel electrophoresis (Figure 34A); 2) Sanger sequencing of the knock-in PCR products (Figure 34B); and 3) high-throughput deep sequencing of the knock-in junction to confirm and quantify the accuracy of gene editing using SAFE-dCas9 in vivo (Figure 34C). Each downstream analysis confirmed the success of the knock-in.
[0365] Furthermore, we tested its in vivo use using adeno-associated virus (AAV) delivery into the lungs of LTC mice. LTC mice contain three genomic alleles: 1) the Lkb1 (flox / flox) allele, which allows for Lkb1-KO when expressing Cre; 2) the R26 (LSL-TdTom) allele, which allows for detection of AAV-transduced cells via the TdTom red fluorescent protein; and 3) the H11 (LSL-Cas9) allele, which allows for expression of Cas9 in AAV-transduced cells. A schematic diagram of the REDI gene editing vector and the Cas9 control vector is shown in Figure 35A. As shown in Figure 35B, successful gene editing using the gene editing vector results in the Kras allele, which drives tumor growth in the lungs of treated mice.
[0366] Approximately 14 weeks after AAV injection, the perfused mouse lungs were removed. Fixed lung tissue was used for imaging analysis to identify tumor formation from successful gene editing (Figure 35C). Quantification of surface tumor number through imaging analysis showed increased gene editing efficiency and increased total tumor number in REDIT-treated mice (Figure 35C). Escherichia coli RecE amino acid sequence (SEQ ID NO: 1): [ka] Escherichia coli RecE_587 amino acid sequence (SEQ ID NO:2): [ka] Escherichia coli CTD_RecE amino acid sequence (SEQ ID NO: 3): [ka] Pantoea brenneri RecE amino acid sequence (SEQ ID NO: 4): [ka] Plautia stali type F symbiont RecE amino acid sequence (SEQ ID NO: 5): [ka] Providencia sp. MGF014 RecE amino acid sequence (SEQ ID NO: 6): [ka] Shigella sonnei RecE amino acid sequence (SEQ ID NO:7): [ka] Pseudobacteriovorax antillogorgiicola RecE amino acid sequence (SEQ ID NO:8): [ka] Escherichia coli RecT amino acid sequence (SEQ ID NO:9): [ka] Pantoea brenneri RecT amino acid sequence (SEQ ID NO: 10): [ka] Plautia stali type F symbiont RecT amino acid sequence (SEQ ID NO: 11): [ka] Providencia sp. MGF014 RecT amino acid sequence (SEQ ID NO: 12): [ka] Shigella sonnei RecT amino acid sequence (SEQ ID NO: 13): [ka] Pseudobacteriovorax antillogorgiicola RecT amino acid sequence (SEQ ID NO: 14): [ka] SV40 NLS amino acid sequence (SEQ ID NO: 16): PKKKRKV Ty1 NLS amino acid sequence (SEQ ID NO: 17): NSKKRSLEDNETEIKVSRDTWNTKNMRSLEPPRSKKRIH c-Myc NLS amino acid sequence (SEQ ID NO: 18): PAAKRVKLD biSV40 NLS amino acid sequence (SEQ ID NO: 19): KRTADGSEFESPKKKRKV Mut NLS amino acid sequence (SEQ ID NO: 20): PEKKRRRPSGSVPVLARPSPPKAGKSSCI Template DNA sequence (underlined sequences indicate substituted or inserted edited sequences) EMX1 HDR template sequence (SEQ ID NO: 79): [ka] VEGFA HDR template sequence (SEQ ID NO: 80): [ka] DYNLT1 HDR template sequence (SEQ ID NO: 81): [ka] HSP90AA1 HDR template sequence (SEQ ID NO: 82): [ka] AAVS1 HDR template sequence (SEQ ID NO: 83): [ka] OCT4 HDR template sequence (SEQ ID NO: 84): [ka] Pantoea stewartii RecT DNA (SEQ ID NO: 85): [ka] Pantoea stewartii RecE DNA (SEQ ID NO: 86): [ka] Pantoea brenneri RecT DNA (SEQ ID NO: 87): [ka] Pantoea brenneri RecE DNA (SEQ ID NO: 88): [ka] Pantoea dispersa RecT DNA (SEQ ID NO: 89): [ka] Pantoea dispersa RecE DNA (SEQ ID NO: 90): [ka] Plautia stali F-type symbiont RecT DNA (SEQ ID NO: 91): [ka] Plautia stali F-type symbiont RecE DNA (SEQ ID NO: 92): [ka] Providencia stuartii RecT DNA (SEQ ID NO: 93): [ka] Providencia stuartii RecE DNA (SEQ ID NO: 94): [ka] Providencia sp. MGF014 RecT DNA (SEQ ID NO: 95): [ka] Providencia sp. MGF014 RecE DNA (SEQ ID NO: 96): [ka] Shewanella putrefaciens RecT DNA (SEQ ID NO: 97): [ka] Shewanella putrefaciens RecE DNA (SEQ ID NO: 98): [ka] Bacillus sp. MUM 116 RecT DNA (SEQ ID NO: 99): [ka] Bacillus sp. MUM 116 RecE DNA (SEQ ID NO: 100): [ka] Shigella sonnei RecT DNA (SEQ ID NO: 101): [ka] Shigella sonnei RecE DNA (SEQ ID NO: 102): [ka] Salmonella enterica RecT DNA (SEQ ID NO: 103): [ka] Salmonella enterica RecE DNA (SEQ ID NO: 104): [ka] Acetobacter RecT DNA (SEQ ID NO: 105): [ka] Acetobacter RecE DNA (SEQ ID NO: 106): [ka] Salmonella enterica subsp. enterica serovar Javiana str. 10721 RecT DNA (SEQ ID NO: 107): [ka] Salmonella enterica subsp. enterica serovar Javiana str. 10721 RecE DNA (SEQ ID NO: 108): [ka] Pseudobacteriovorax antillogorgiicola RecT DNA (SEQ ID NO: 109): [ka] Pseudobacteriovorax antillogorgiicola RecE DNA (SEQ ID NO: 110): [ka] Photobacterium sp. JCM 19050 RecT DNA (SEQ ID NO: 111): [ka] Photobacterium sp. JCM 19050 RecE DNA (SEQ ID NO: 112): [ka] Providencia alcalifaciens DSM 30120 RecT DNA (SEQ ID NO: 113): [ka] Providencia alcalifaciens DSM 30120 RecE DNA (SEQ ID NO: 114): [ka] Pantoea stewartii RecT protein (SEQ ID NO: 115): [ka] Pantoea stewartii RecE protein (SEQ ID NO: 116): [ka] Pantoea brenneri RecT protein (SEQ ID NO: 117): [ka] Pantoea brenneri RecE protein (SEQ ID NO: 118): [ka] Pantoea dispersa RecT protein (SEQ ID NO: 119): [ka] Pantoea dispersa RecE protein (SEQ ID NO: 120): [ka] Plautia stali type F symbiont RecT protein (SEQ ID NO: 121): [ka] Plautia stali type F symbiont RecE protein (SEQ ID NO: 122): [ka] Providencia stuartii RecT protein (SEQ ID NO: 123): [ka] Providencia stuartii RecE protein (SEQ ID NO: 124): [ka] Providencia sp. MGF014 RecT protein (SEQ ID NO: 125): [ka] Providencia sp. MGF014 RecE protein (SEQ ID NO: 126): [ka] Shewanella putrefaciens RecT protein (SEQ ID NO: 127): [ka] Shewanella putrefaciens RecE protein (SEQ ID NO: 128): [ka] Bacillus sp. MUM 116 RecT protein (SEQ ID NO: 129): [ka] Bacillus sp. MUM 116 RecE protein (SEQ ID NO: 130): [ka] Shigella sonnei RecT protein (SEQ ID NO: 131): [ka] Shigella sonnei RecE protein (SEQ ID NO: 132): [ka] Salmonella enterica RecT protein (SEQ ID NO: 133): [ka] Salmonella enterica RecE protein (SEQ ID NO: 134): [ka] Acetobacter RecT protein (SEQ ID NO: 135): [ka] Acetobacter RecE protein (SEQ ID NO: 136): [ka] Salmonella enterica subsp. enterica serovar Javiana str. 10721 RecT protein (SEQ ID NO: 137): [ka] Salmonella enterica subsp. enterica serovar Javiana str. 10721 RecE protein (SEQ ID NO: 138): [ka] Pseudobacteriovorax antillogorgiicola RecT protein (SEQ ID NO: 139): [ka] Pseudobacteriovorax antillogorgiicola RecE protein (SEQ ID NO: 140): [ka] Photobacterium sp. JCM 19050 RecT protein (SEQ ID NO: 141): [ka] Photobacterium sp. JCM 19050 RecE protein (SEQ ID NO: 142): [ka] Providencia alcalifaciens DSM 30120 RecT protein (SEQ ID NO: 143): [ka] Providencia alcalifaciens DSM 30120 RecE protein (SEQ ID NO: 144): [ka] Mouse albumin knock-in sense template (SEQ ID NO: 160) [ka] Mouse Albumin Knock-in Antisense Template (SEQ ID NO: 161) [ka] (SEQ ID NO: 162) [ka] Example 16
[0367] We predicted the structures of E. coli RecT (EcRecT) alone (Figure 36A) and with single-stranded DNA bound (Figures 36B and 36C). The contact interface is consistent with the truncation data (Example 7, Figure 20A). The predicted interactions of EcRecT SSAP amino acids with DNA are shown in Figures 37A and 37B. Example 17
[0368] From the sequence data, 322 SSAP proteins were identified, synthesized, and screened for activity with Cas9 and dCas9. Gene editing activity is shown in Table 7 below, followed by the amino acid sequences of the proteins. [Table 7-1] [Table 7-2] [Table 7-3] [Table 7-4] [Table 7-5] [Table 7-6] [Table 7-7] [Table 7-8] UPI0000010203 (SEQ ID NO: 172) [ka] UPI00000105D3 (SEQ ID NO: 173) [ka] UPI0000030D3A / HAW2682705.1 RecT [Escherichia coli] (SEQ ID NO: 167) [ka] UPI0000030D3E (SEQ ID NO: 166) [ka] UPI000009AF52 (SEQ ID NO: 174) [ka] UPI000009B019 (SEQ ID NO: 175) [ka] UPI000009B628 (SEQ ID NO: 176) [ka] UPI000009BC15 (SEQ ID NO: 177) [ka] UPI00000B3F97 Bet [Gammaproteobacteria] (SEQ ID NO: 178) [ka] UPI000019AB49 Bet [Escherichia coli] (SEQ ID NO: 179) [ka] UPI000034E66D Bet [Lactococcus phage phiLC3] (SEQ ID NO: 180) [ka] UPI00005F0A78 (SEQ ID NO: 181) [ka] UPI000150D6AC (SEQ ID NO: 182) [ka] UPI0001594E53 (SEQ ID NO: 183) [ka] UPI00015968D7 (SEQ ID NO: 184) [ka] UPI00015C01AE (SEQ ID NO: 185) [ka] UPI00015C02E0 (SEQ ID NO: 186) [ka] UPI00019E1F9A (SEQ ID NO: 187) [ka] UPI0001BEF484 (SEQ ID NO: 188) [ka] UPI0001CE597A CK3_26380 [butyric acid-producing bacteria SS3 / 4] (SEQ ID NO: 189) [ka] UPI0001D2DF22 RecT [Cellulosilyticum lentocellum] (SEQ ID NO: 190) [ka] UPI0001E0C499 (SEQ ID NO: 191) [ka] UPI0001E2AFC1 (SEQ ID NO: 192) [ka] UPI0001E35ACE (SEQ ID NO: 193) [ka] UPI00020BA2E0 (SEQ ID NO: 194) [ka] UPI000212F382 (SEQ ID NO: 195) [ka] UPI00022F8B4D (SEQ ID NO: 196) [ka] UPI0002314B74 (SEQ ID NO: 197) [ka] UPI00025CAD2E (sequence number 198) [ka] UPI00025CF49A (SEQ ID NO: 199) [ka] UPI0002AD92E7 (sequence number 200) [ka] UPI0002B78771 (sequence number 201) [ka] UPI0002B78B34 (sequence number 202) [ka] UPI0002B884F0 / WP_003158887.1 Bet [Pseudomonas aeruginosa] (SEQ ID NO: 203) [ka] UPI0002CB4A67 / WP_010792303.1 Bet [Pseudomonas aeruginosa] (SEQ ID NO: 204) [ka] UPI0002E4C0BF (SEQ ID NO: 205) [ka] UPI0003282677 (SEQ ID NO: 206) [ka] UPI00033853AF (SEQ ID NO: 207) [ka] UPI0003427695 (SEQ ID NO: 208) [ka] UPI000353091F (sequence number 209) [ka] UPI000386D631 (SEQ ID NO: 210) [ka] UPI0003E3D237 (SEQ ID NO: 211) [ka] UPI00044F7143 (SEQ ID NO: 212) [ka] UPI0004995B90 (SEQ ID NO: 213) [ka] UPI00051F5876 (SEQ ID NO: 214) [ka] UPI000588C848 (SEQ ID NO: 215) [ka] UPI000598CD40 (SEQ ID NO: 216) [ka] UPI0005DCEBAD (SEQ ID NO: 217) [ka] UPI0005E4CB74 (SEQ ID NO: 218) [ka] UPI0005FEB4B0 (SEQ ID NO: 219) [ka] UPI00062002D2 (sequence number 220) [ka] UPI00064B44C1 (SEQ ID NO: 221) [ka] UPI00064D5E13 (SEQ ID NO: 222) [ka] UPI00065C2D47 Bet [Pseudomonas phage PS-1] (SEQ ID NO: 223) [ka] UPI00067A7349 RecT [Streptococcus phage APCM01] (SEQ ID NO: 224) [ka] UPI0006CE3F5D (SEQ ID NO: 225) [ka] UPI00078E90BE RecT [Pirellula sp. SH-Sr6A] (SEQ ID NO: 226) [ka] UPI00078EBE91 RecT [Pirellula sp. SH-Sr6A] (SEQ ID NO: 227) [ka] UPI00078ED021 (SEQ ID NO: 228) [ka] UPI000795D815 (SEQ ID NO: 229) [ka] UPI00079B135B (SEQ ID NO: 230) [ka] UPI0007B45EC7 (SEQ ID NO: 231) [ka] UPI0007B642FE (SEQ ID NO: 232) [ka] UPI0007B64693 (SEQ ID NO: 233) [ka] UPI0007BCAEAB (SEQ ID NO: 234) [ka] UPI0007F13B78 (SEQ ID NO: 235) [ka] UPI000865F43D (SEQ ID NO: 236) [ka] UPI000865FB15 (SEQ ID NO: 237) [ka] UPI0008D18539 (SEQ ID NO: 238) [ka] UPI0008D990CB (SEQ ID NO: 239) [ka] UPI0008E12231 (SEQ ID NO: 240) [ka] UPI0008EA8633 (SEQ ID NO: 241) [ka] UPI00091F1EB0 (SEQ ID NO: 242) [ka] UPI000958E115 (SEQ ID NO: 243) [ka] UPI0009805C1D (SEQ ID NO: 244) [ka] UPI0009805F63 (SEQ ID NO: 245) [ka] UPI0009880690 (SEQ ID NO: 246) [ka] UPI0009F5E532 (SEQ ID NO: 247) [ka] UPI0009F8F604 (SEQ ID NO: 248) [ka] UPI000A08A794 (SEQ ID NO: 249) [ka] UPI000B36BD3F (SEQ ID NO: 250) [ka] UPI000B38B374 (SEQ ID NO: 251) [ka] UPI000B49B5D9 (SEQ ID NO: 252) [ka] UPI000B4BEFE6 / WP_088258624.1 Bet [Fimbriiglobus ruber] (SEQ ID NO: 253) [ka] UPI000B5661AA (SEQ ID NO: 254) [ka] UPI000B94B1D1 (SEQ ID NO: 255) [ka] UPI000BD04ECE (SEQ ID NO: 256) [ka] WP_032686941.1 RecT [Raoultella planticola] (SEQ ID NO: 257) [ka] WP_069728515.1 RecT [Pantoea brenneri] (SEQ ID NO: 258) [ka] WP_045958294.1 RecT [Xenorhabdus poinarii] (SEQ ID NO: 259) [ka] WP_102086779.1 RecT [Proteus mirabilis] (SEQ ID NO: 260) [ka] WP_109615067.1 RecT [Edwardsiella piscicida] (SEQ ID NO: 261) [ka] WP_124537594.1 RecT [Morganella morganii] (SEQ ID NO: 262) [ka] WP_006657622.1 RecT [Providencia alcalifaciens] (SEQ ID NO: 263) [ka] WP_109401438.1 RecT [Proteus terrae] (SEQ ID NO: 264) [ka] WP_115149784.1 RecT [Plesiomonas shigelloides] (SEQ ID NO: 265) [ka] WP_034910107.1 RecT [Gilliamella apicola] (SEQ ID NO: 266) [ka] WP_016979878.1 RecT [Pseudomonas fluorescens] (SEQ ID NO: 267) [ka] WP_080977968.1 RecT [Pseudomonas stutzeri] (SEQ ID NO: 268) [ka] KXJ39364.1 AXA67_02205 [Methylothermaceae Bacterium B42] (SEQ ID NO: 269) [ka] WP_106478153.1 RecT [Halomonadaceae bacterium R4HLG17] (SEQ ID NO: 270) [ka] WP_129141488.1 RecT [Halomonas coralii] (SEQ ID NO: 271) [ka] WP_084261900.1 RecT [Zymobacter palmae] (SEQ ID NO: 272) [ka] WP_020007369.1 RecT [Salinicoccus albus] (SEQ ID NO: 273) [ka] WP_131521405.1 RecT [unclassified Lysinibacillus] (SEQ ID NO: 274) [ka] WP_132769795.1 RecT [Tepidibacillus fermentans] (SEQ ID NO: 275) [ka] WP_120191052.1 RecT [Ammoniphilus oxalaticus] (SEQ ID NO: 276) [ka] WP_066790810.1 RecT [Rummeliibacillus stabekisii] (SEQ ID NO: 277) [ka] WP_098408280.1 RecT [Bacillus] (multiple species) (SEQ ID NO: 278) [ka] WP_047150996.1 RecT [Aneurinibacillus tyrosinisolvens] (SEQ ID NO: 279) [ka] WP_018705791.1 RecT [Siminovitchia fordii] (SEQ ID NO: 280) [ka] WP_035430909.1 RecT [Bacillus sp. UNC322MFChir4.1] (SEQ ID NO: 281) [ka] RDC50983.1 RecT [Acinetobacter sp. RIT592] (SEQ ID NO: 282) [ka] WP_150051132.1 RecT [Methylomonas rhizoryzae] (SEQ ID NO: 283) [ka] WP_097006457.1 [Lacrimispora amygdalina] (SEQ ID NO: 284) [ka] WP_087225255.1 RecT [Lachnoclostridium sp. An14] (SEQ ID NO: 285) [ka] WP_002566991.1 RecT [Enterocloster bolteae] (SEQ ID NO: 286) [ka] WP_132412730.1 RecT [Kribbella albertanoniae] (SEQ ID NO: 287) [ka] WP_130067396.1 RecT [Bacillus albus] (SEQ ID NO: 288) [ka] WP_087099033.1 RecT [Bacillus cytotoxicus] (SEQ ID NO: 289) [ka] WP_149216302.1 RecT [Bacillus sp. JAS24-2] (SEQ ID NO: 290) [ka] WP_125141636.1 RecT [Clostridium transplantifaecale] (SEQ ID NO: 291) [ka] WP_120055566.1 RecT [Lachnoclostridium pacaense] (SEQ ID NO: 292) [ka] WP_118246619.1 RecT [Clostridium sp. AM58-1XD] (SEQ ID NO: 293) [ka] WP_025114396.1 RecT [Lysinibacillus fusiformis] (SEQ ID NO: 294) [ka] WP_083048409.1 RecT [Marispirochaeta aestuarii] (SEQ ID NO: 295) [ka] WP_099424140.1 RecT [Solibacillus sp. R5-41] (SEQ ID NO: 296) [ka] WP_076065282.1 RecT [Viridibacillus sp. FSL H8-0123] (SEQ ID NO: 297) [ka] WP_024292388.1 RecT [Lacrimispora indolis] (SEQ ID NO: 298) [ka] WP_009524931.1 RecT [Peptoanaerobacter stomatis] (SEQ ID NO: 299) [ka] WP_015358111.1 RecT [Thermoclostridium stercorarium] (SEQ ID NO: 300) [ka] WP_002595146.1 RecT [Enterocloster clostridioformis] (SEQ ID NO: 301) [ka] WP_100306418.1 RecT [Lacrimispora celerecrescens] (SEQ ID NO: 302) [ka] WP_071062796.1 RecT [Andreesenia angusta] (SEQ ID NO: 303) [ka] SFO83314.1 RecT [Amycolatopsis arida] (SEQ ID NO: 304) [ka] WP_110092637.1 RecT [Corynebacterium striatum] (SEQ ID NO: 305) [ka] WP_129692339.1 RecT [Gottfriedia acidiceleris] (SEQ ID NO: 306) [ka] WP_118016648.1 RecT [Unclassified Coprococcus] (multiple species) (SEQ ID NO: 307) [ka] WP_051200279.1 RecT [Butyrivibrio sp. FCS006] (SEQ ID NO: 308) [ka] WP_107514794.1 RecT [Staphylococcus equorum] (SEQ ID NO: 309) [ka] WP_117624242.1 RecT [Hungatella hathewayi] (SEQ ID NO: 310) [ka] WP_118771779.1 RecT [Roseburia intestinalis] (SEQ ID NO: 311) [ka] WP_107378794.1 RecT [Staphylococcus chromogenes] (SEQ ID NO: 312) [ka] WP_094369469.1 RecT [Romboutsia weinsteinii] (SEQ ID NO: 313) [ka] CDF42377.1 [Roseburia sp. CAG:182] (SEQ ID NO: 314) [ka] WP_123609006.1 RecT [Mobilisporobacter senegalensis] (SEQ ID NO: 315) [ka] WP_115856892.1 RecT [Staphylococcus felis] (SEQ ID NO: 316) [ka] WP_108404827.1 RecT [Corynebacterium liangguodongii] (SEQ ID NO: 317) [ka] WP_021747387.1 RecT [unclassified Oscillibacter] (multiple species) (SEQ ID NO: 318) [ka] WP_103110615.1 RecT [Brevibacillus reuszeri] (SEQ ID NO: 319) [ka] WP_016998679.1 RecT [Mammaliicoccus] (SEQ ID NO: 320) [ka] WP_147540090.1 RecT [Clostridiaceae bacteria] (SEQ ID NO: 321) [ka] WP_019168122.1 RecT [Staphylococcus intermedius] (SEQ ID NO: 322) [ka] WP_148820236.1 RecT [Corynebacterium urealyticum] (SEQ ID NO: 323) [ka] WP_096823857.1 RecT [Staphylococcus nepalensis] (SEQ ID NO: 324) [ka] WP_098170605.1 RecT [Bacillus sp. AFS017336] (SEQ ID NO: 325) [ka] WP_087290962.1 RecT [Pseudoflavonifractor sp. An184] (SEQ ID NO: 326) [ka] WP_051264703.1 RecT [Nakamurella lactea] (SEQ ID NO: 327) [ka] CCZ61365.1 [Clostridium hathewayi CAG:224] (SEQ ID NO: 328) [ka] WP_068720576.1 RecT [Veillonellaceae bacterium DNF00626] (SEQ ID NO: 329) [ka] WP_037404193.1 RecT [Solobacterium moorei] (SEQ ID NO: 330) [ka] WP_027347470.1 RecT [Helcococcus sueciensis] (SEQ ID NO: 331) [ka] WP_072526012.1 RecT [Clostridium sp. Marseille-P3244] (SEQ ID NO: 332) [ka] WP_092453396.1 RecT [Clostridium fimetarium] (SEQ ID NO: 333) [ka] WP_027295741.1 RecT [Robinsoniella sp. KNHs210] (SEQ ID NO: 334) [ka] WP_117768035.1 RecT [Blautia sp. OF03-15BH] (SEQ ID NO: 335) [ka] SCJ42694.1 [Ruminococcus sp.] (SEQ ID NO: 336) [ka] WP_092724975.1 RecT [Romboutsia lituseburensis] (SEQ ID NO: 337) [ka] KKZ74881.1 VO63_05385 [Streptomyces showdoensis] (SEQ ID NO: 338) [ka] WP_055284109.1 RecT [Dorea longicatena] (SEQ ID NO: 339) [ka] SDL28883.1 RecT [Streptomyces indicus] (SEQ ID NO: 340) [ka] WP_145458209.1 RecT [Staphylococcus pettenkoferi] (SEQ ID NO: 341) [ka] WP_117787252.1 RecT [Tyzzerella nexilis] (SEQ ID NO: 342) [ka] WP_073112630.1 RecT [Hespellia stercorisuis] (SEQ ID NO: 343) [ka] CDD36322.1 [Roseburia sp. CAG:309] (SEQ ID NO: 344) [ka] WP_128520904.1 RecT [Absicoccus porci] (SEQ ID NO: 345) [ka] GAK01483.1 RecT [Geomicrobium sp. JCM 19055] (SEQ ID NO: 346) [ka] WP_135329961.1 RecT [Streptomyces sp. MZ04] (SEQ ID NO: 347) [ka] WP_079588582.1 RecT [Acetoanaerobium noterae] (SEQ ID NO: 348) [ka] WP_107635892.1 RecT [Staphylococcus haemolyticus] (SEQ ID NO: 349) [ka] WP_107638953.1 RecT [Staphylococcus hominis] (SEQ ID NO: 350) [ka] SUY49750.1 RecT [Lacrimispora sphenoides] (SEQ ID NO: 351) [ka] CDE68291.1 [Clostridium sp. CAG:277] (SEQ ID NO: 352) [ka] WP_060905391.1 RecT [Streptomyces scabiei] (SEQ ID NO: 353) [ka] WP_146678271.1 RecT [Pirellula sp. SH-Sr6A] (SEQ ID NO: 354) [ka] WP_126032909.1 RecT [Bifidobacterium castoris] (SEQ ID NO: 355) [ka] WP_114599505.1 RecT [Staphylococcus warneri] (SEQ ID NO: 356) [ka] SCQ72869.1 RecT protein [Propionibacterium freudenreichii] (SEQ ID NO: 357) [ka] WP_127100780.1 RecT [Asaia sp. W19] (SEQ ID NO: 358) [ka] EIC09117.1 RecT protein [Microbacterium laevaniformans OR221] (SEQ ID NO: 359) [ka] WP_136046271.1 RecT [Microbacterium sp. K41] (SEQ ID NO: 360) [ka] WP_136309287.1 RecT [Streptococcus pyogenes] (SEQ ID NO: 361) [ka] WP_110990907.1 RecT [Mesotoga sp. TolDC] (SEQ ID NO: 362) [ka] WP_109196224.1 RecT [Streptomyces sp. CS014] (SEQ ID NO: 363) [ka] WP_068202759.1 RecT [Isoptericola dokdonensis] (SEQ ID NO: 364) [ka] WP_114797327.1 RecT [Gaiella occulta] (SEQ ID NO: 365) [ka] PAV10712.1 CBG25_01455 [Arsenophonus sp. ENCA] (SEQ ID NO: 366) [ka] WP_147981944.1 RecT [Streptomyces sp. ms191] (SEQ ID NO: 367) [ka] BAQ93806.1 Phage RecT Family (TIGR00616) [Uncultured Mediterranean Phage uvMED] (SEQ ID NO: 368) [ka] WP_061405262.1 RecT [Streptomyces] (multiple species) (SEQ ID NO: 369) [ka] WP_114014965.1 RecT [Streptomyces reniochalinae] (SEQ ID NO: 370) [ka] WP_027699748.1 RecT [Weissella oryzae] (SEQ ID NO: 371) [ka] SYW13692.1 Phage RecT family protein [Oenococcus oeni] (SEQ ID NO: 372) [ka] WP_141158250.1 RecT [Pseudarthrobacter sp. NIBRBAC000502771] (SEQ ID NO: 373) [ka] TAK04183.1 EPO34_03495 [Patescibacteria group bacteria] (SEQ ID NO: 374) [ka] WP_092601202.1 RecT [Actinopolyspora xinjiangensis] (SEQ ID NO: 375) [ka] WP_067024969.1 RecT [Mycobacterium sp. 1245499.0] (SEQ ID NO: 376) [ka] WP_075737485.1 RecT [Streptomyces acidiscabies] (SEQ ID NO: 377) [ka] AKT73182.1 RecT (prophage-associated) [Yersinia pestis] (SEQ ID NO: 378) [ka] WP_123127078.1 RecT [Rufibacter latericius] (SEQ ID NO: 379) [ka] WP_093587584.1 RecT [Unclassified Streptomyces] (multiple species) (SEQ ID NO: 380) [ka] WP_030975214.1 RecT [Streptomyces sp. NRRL S-1824] (SEQ ID NO: 381) [ka] RKT60104.1 RecT [Agromyces sp. OV415] (SEQ ID NO: 382) [ka] WP_017415747.1 RecT [Clostridium tunisiense] (SEQ ID NO: 383) [ka] RYE05836.1 EOP33_01060 [Rickettsiaceae bacteria] (SEQ ID NO: 384) [ka] WP_052399147.1 RecT [Francisella sp. FSC1006] (SEQ ID NO: 385) [ka] WP_067349107.1 RecT [Streptomyces noursei] (SEQ ID NO: 386) [ka] WP_143887802.1 RecT [Streptococcus lutetiensis] (SEQ ID NO: 387) [ka] WP_073793143.1 RecT [Streptomyces uncialis] (SEQ ID NO: 388) [ka] WP_116200709.1 RecT [Amycolatopsis circi] (SEQ ID NO: 389) [ka] WP_020135111.1 RecT [Streptomyces sp. 351MFTsu5.1] (SEQ ID NO: 390) [ka] WP_099421180.1 RecT [Streptococcus macedonicus] (SEQ ID NO: 391) [ka] WP_141925904.1 RecT [Haloactinospora alba] (SEQ ID NO: 392) [ka] WP_136710836.1 RecT [Clostridium tyrobutyricum] (SEQ ID NO: 393) [ka] WP_132110073.1 RecT [Actinocrispum wychmicini] (SEQ ID NO: 394) [ka] WP_125769509.1 RecT [Companilactobacillus furfuricola] (SEQ ID NO: 395) [ka] WP_004234437.1 RecT [Streptococcus parauberis] (SEQ ID NO: 396) [ka] WP_006845711.1 RecT [Weissella koreensis] (SEQ ID NO: 397) [ka] WP_073846185.1 RecT [Amycolatopsis sp. CB00013] (SEQ ID NO: 398) [ka] WP_142511229.1 RecT [Leuconostoc pseudomesenteroides] (SEQ ID NO: 399) [ka] WP_023055804.1 RecT [Peptoniphilus sp. BV3C26] (SEQ ID NO: 400) [ka] PCR98661.1 RecT [Lactococcus fujiensis JCM 16395] (SEQ ID NO: 401) [ka] WP_106316803.1 RecT [Actinoplanes italicus] (sequence number 402) [ka] WP_013655830.1 RecT [Cellulosilyticum lentocellum] (SEQ ID NO: 403) [ka] WP_148001988.1 RecT [Streptomyces sp. adm13(2018)] (SEQ ID NO: 404) [ka] WP_011988985.1 RecT [Clostridium kluyveri] (SEQ ID NO: 405) [ka] GAC42786.1 recombinant DNA repair protein [Paenibacillus popilliae ATCC 14706] (SEQ ID NO: 406) [ka] OBR91022.1 RecT [Clostridium ragsdalei P11] (SEQ ID NO: 407) [ka] SEI77195.1 RecT [Paenibacillus polymyxa] (SEQ ID NO: 408) [ka] KKT72154.1 RecT [Candidatus Collierbacteria GW2011_GWB1_44_6] (SEQ ID NO: 409) [ka] WP_125777163.1 RecT [Antribacter gilvus] (SEQ ID NO: 410) [ka] WP_130123223.1 RecT [Lactococcus sp. S-13] (SEQ ID NO: 411) [ka] WP_147265819.1 RecT [Nocardia puris] (SEQ ID NO: 412) [ka] TCP18101.1 RecT [Nicoletella semolina] (SEQ ID NO: 413) [ka] OAB27843.1 recombinase [Paenibacillus macquariensis subsp. defensor] (SEQ ID NO: 414) [ka] WP_019417330.1 RecT [Anoxybacillus] (SEQ ID NO: 415) [ka] CDA71469.1 Phage RecT Family [Ruminococcus sp. CAG:579] (SEQ ID NO: 416) [ka] WP_019108121.1 RecT [Peptoniphilus senegalensis] (SEQ ID NO: 417) [ka] AFH22576.1 RecT family protein [environmental Halophage eHP-30] (SEQ ID NO: 418) [ka] WP_138067957.1 RecT [Streptococcus pseudoporcinus] (SEQ ID NO: 419) [ka] WP_072904346.1 RecT [Hathewaya proteolytica] (SEQ ID NO: 420) [ka] GAE17732.1 RecT [Bacteroides pyogenes DSM 20611 = JCM 6294] (SEQ ID NO: 421) [ka] CDF09406.1 [Eubacterium sp. CAG:76] (SEQ ID NO: 422) [ka] WP_099299656.1 RecT [Pediococcus pentosaceus] (SEQ ID NO: 423) [ka] WP_118227047.1 RecT [Bacteroides eggerthii] (SEQ ID NO: 424) [ka] WP_094754495.1 RecT [Criibacterium bergeronii] (SEQ ID NO: 425) [ka] WP_045553720.1 RecT [Listeria] (multiple species) (SEQ ID NO: 426) [ka] WP_106024518.1 RecT [Clostridium thermopalmarium] (SEQ ID NO: 427) [ka] WP_073010654.1 RecT [Virgibacillus chiguensis] (SEQ ID NO: 428) [ka] WP_111921306.1 RecT [Clostridium cochlearium] (SEQ ID NO: 429) [ka] WP_019125538.1 RecT [Peptoniphilus grossensis] (SEQ ID NO: 430) [ka] ERL63827.1 YqaK [Schleiferilactobacillus shenzhenensis LY-73] (SEQ ID NO: 431) [ka] WP_051267408.1 RecT [Gulosibacter molinativorax] (SEQ ID NO: 432) [ka] WP_112330076.1 RecT [Cereibacter johrii] (SEQ ID NO: 433) [ka] WP_063601171.1 RecT [Clostridium coskatii] (SEQ ID NO: 434) [ka] WP_118206945.1 RecT [Bacteroides stercoris] (SEQ ID NO: 435) [ka] WP_099840029.1 RecT [Clostridium combesii] (SEQ ID NO: 436) [ka] WP_069686512.1 RecT [Oceanobacillus sp. E9] (SEQ ID NO: 437) [ka] RMD50745.1 [Candidatus Parcubacteria bacteria] (SEQ ID NO: 438) [ka] WP_061413958.1 RecT [Lactococcus sp. DD01] (SEQ ID NO: 439) [ka] WP_147129628.1 RecT [Nocardia ninae] (SEQ ID NO: 440) [ka] WP_074846740.1 RecT [Clostridium cadaveris] (SEQ ID NO: 441) [ka] WP_038246219.1 RecT [Virgibacillus] (multiple species) (SEQ ID NO: 442) [ka] WP_106064284.1 RecT [Clostridium liquoris] (SEQ ID NO: 443) [ka] WP_028562280.1 RecT [Paenibacillus pinihumi] (SEQ ID NO: 444) [ka] WP_068672306.1 RecT [Oceanobacillus sp. Castelsardo] (SEQ ID NO: 445) [ka] WP_067592792.1 RecT [Nocardia terpenica] (SEQ ID NO: 446) [ka] WP_079708113.1 RecT [Paraliobacillus ryukyuensis] (SEQ ID NO: 447) [ka] OLA20462.1 BHW17_09115 [Dorea sp. 42_8] (SEQ ID NO: 448) [ka] WP_058906805.1 RecT [Lactiplantibacillus plantarum] (SEQ ID NO: 449) [ka] RZT66774.1 RecT [Leucobacter luti] (SEQ ID NO: 450) [ka] WP_087916041.1 RecT [Paenibacillus donghaensis] (SEQ ID NO: 451) [ka] WP_009411480.1 RecT [Capnocytophaga sp. oral taxon 324] (SEQ ID NO: 452) [ka] WP_116232802.1 RecT [Paenibacillus sp. VMFN-D1] (SEQ ID NO: 453) [ka] WP_123849158.1 RecT [Chitinophaga lutea] (SEQ ID NO: 454) [ka] WP_078410260.1 RecT [Priestia abyssalis] (SEQ ID NO: 455) [ka] AAT90028.1 phage recombinant protein [Leifsonia xyli subsp. xyli str. CTCB07] (SEQ ID NO: 456) [ka] WP_080022455.1 RecT [Clostridium thermobutyricum] (SEQ ID NO: 457) [ka] WP_081759639.1 RecT [Clostridium jeddahense] (SEQ ID NO: 458) [ka] WP_089281299.1 RecT [Anaerovirgula multivorans] (SEQ ID NO: 459) [ka] RDI65706.1 Phage RecT family recombinase [Nocardia pseudobrasiliensis] (SEQ ID NO: 460) [ka] WP_076170610.1 RecT [Paenibacillus rhizosphaerae] (SEQ ID NO: 461) [ka] WP_106833617.1 RecT [Brevibacillus porteri] (SEQ ID NO: 462) [ka] RDE19343.1 RecT [Parageobacillus thermoglucosidasius] (SEQ ID NO: 463) [ka] WP_138600901.1 RecT [Pseudoalteromonas] (multispecies) (SEQ ID NO: 464) [ka] WP_082209600.1 RecT [Peptostreptococcaceae bacteria VA2] (SEQ ID NO: 465) [ka] WP_026627303.1 RecT [Dysgonomonas capnocytophagoides] (SEQ ID NO: 466) [ka] WP_109523733.1 RecT [Nocardia aurea] (SEQ ID NO: 467) [ka] GAE09585.1 [Paenibacillus sp. JCM 10914] (SEQ ID NO: 468) [ka] RRG08833.1 RecT [Lactobacillus sp.] (SEQ ID NO: 469) [ka] GEA30849.1 CDIOL_17720 [Clostridium diolis] (SEQ ID NO: 470) [ka] WP_077867213.1 RecT [Clostridium saccharobutylicum] (SEQ ID NO: 471) [ka] WP_132305216.1 RecT [Paenibacillus sp. BK033] (SEQ ID NO: 472) [ka] RPI78794.1 EHM45_05245 [Desulfobacteraceae bacteria] (SEQ ID NO: 473) [ka] WP_051624047.1 RecT [Clostridium akagii] (SEQ ID NO: 474) [ka] WP_081735325.1 RecT [Paenibacillus gorillae] (SEQ ID NO: 475) [ka] WP_084505057.1 RecT [Acetobacterium dehalogenans] (SEQ ID NO: 476) [ka] AGF93134.1 RecT protein [uncultured organism] (SEQ ID NO: 477) [ka] WP_076079849.1 RecT [Paenibacillus sp. FSL R7-0333] (SEQ ID NO: 478) [ka] WP_119800346.1 RecT [Paenibacillus sp. 1011MAR3C5] (SEQ ID NO: 479) [ka] WP_025706233.1 RecT [Paenibacillus graminis] (SEQ ID NO: 480) [ka] OIO76374.1 AUJ88_06865 [Gallionellaceae bacteria CG1_02_56_997] (SEQ ID NO: 481) [ka] WP_131535536.1 RecT [Pedobacter nototheniae] (SEQ ID NO: 482) [ka] WP_028113352.1 RecT [Ferrimonas kyonanensis] (SEQ ID NO: 483) [ka] WP_100916003.1 RecT [Pseudoalteromonas spongiae] (SEQ ID NO: 484) [ka] WP_125711747.1 RecT [Companilactobacillus kedongensis] (SEQ ID NO: 485) [ka] WP_002845682.1 RecT [Peptostreptococcus anaerobius] (SEQ ID NO: 486) [ka] WP_115407185.1 RecT [Shewanella morhuae] (SEQ ID NO: 487) [ka] WP_081955873.1 RecT [Helicobacter trogontum] (SEQ ID NO: 488) [ka] WP_064664300.1 RecT [Pseudoalteromonas sp. MQS005] (SEQ ID NO: 489) [ka] WP_069455496.1 RecT [Shewanella xiamenensis] (SEQ ID NO: 490) [ka] RTL04618.1 EKK58_09925 [Candidatus Dependentiae bacteria] (SEQ ID NO: 491) [ka] Example 18
[0369] Gene editing, as exemplified by the CRISPR-Cas9 system, has become a powerful tool for investigating mechanisms of human health and disease. Cas9 editing induces DNA damage at on-target and off-target sites and can rely on error-prone endogenous DNA repair mechanisms. These characteristics often lead to unwanted mutations and safety concerns, which may be exacerbated when Applicant modifies long sequences. Based on previous research showing that mammalian genomic DNA becomes transiently accessible upon dCas9 DNA unwinding and R-loop formation, Applicant hypothesized that single-strand annealing proteins (SSAPs) could stimulate DNA strand exchange for gene editing when coupled to the dCas9-guide RNA complex. Therefore, Applicant developed a non-cleavage gene editing tool using catalytically inactive dCas9 to knock in long sequences. Applicant's data demonstrated that this dCas9-based editor has very low editing error at the target locus, minimal detectable off-target effects, and higher overall accuracy than Cas9 editors. On the other hand, the dCas9-SSAP editor had comparable efficiency to the Cas9 editor and robust performance across human cell lines and stem cells. This dCas9-SSAP editor was effective in inserting sequences of variable length up to the kilobase scale. In experiments in which applicants chemically inhibited DNA repair enzymes, dCas9-SSAP editing demonstrated remarkable independence from endogenous mammalian repair pathways. To facilitate convenient viral delivery of the dCas9-SSAP editor to challenging cell types, applicants performed truncation and aptamer engineering to minimize its size so that it fits into a single AAV vector for future applications. Overall, this tool opens opportunities for safer genome engineering in mammalian cells.
[0370] Since the initial demonstration of CRISPR-Cas9 gene editing, significant efforts have been made to improve and expand gene editing techniques for studying genome function, modeling biological processes, and gene therapy. New generation gene editing tools, such as base editing and prime editing, have substantially improved the efficiency and fidelity of gene editing and are powerful for modifying relatively short sequences. Most gene editing tools function by cleaving genomic DNA to induce single-strand nicks (SSNs) or double-strand breaks (DSBs), which facilitate targeted editing. These DNA modifications are often repaired by error-prone endogenous pathways such as non-homologous end joining (NHEJ) (12). This process often results in unwanted mutations and off-target effects, which can lead to toxicity and raise safety concerns. Such editing errors and off-target effects become increasingly, and sometimes prohibitively, severe when manipulating long genomic sequences (>=100 bp). These undesirable effects limit the application of gene editing to large-scale genome knock-in or in vivo gene editing. Reference is made to WO2020 / 191241, WO2020 / 191153, WO2020 / 191245, WO2020 / 191243, WO2020 / 191233, WO2020 / 191246, WO2020 / 191249, WO2020 / 191239, WO2020 / 191234, WO2020 / 191242, WO2020 / 191248, WO2020191171 and WO2021 / 226558, which involve what is known as prime editing and biprime editing. Each of WO2020 / 191241, WO2020 / 191153, WO2020 / 191245, WO2020 / 191243, WO2020 / 191233, WO2020 / 191246, WO2020 / 191249, WO2020 / 191239, WO2020 / 191234, WO2020 / 191242, WO2020 / 191248, WO2020191171 and WO2021 / 226558 is incorporated herein by reference.The RTs of WO2020 / 191241, WO2020 / 191153, WO2020 / 191245, WO2020 / 191243, WO2020 / 191233, WO2020 / 191246, WO2020 / 191249, WO2020 / 191239, WO2020 / 191234, WO2020 / 191242, WO2020 / 191248, WO2020191171 and WO2021 / 226558 can be used in the practice of the present invention. The linker or functional linking methods of WO2020 / 191241, WO2020 / 191153, WO2020 / 191245, WO2020 / 191243, WO2020 / 191233, WO2020 / 191246, WO2020 / 191249, WO2020 / 191239, WO2020 / 191234, WO2020 / 191242, WO2020 / 191248, WO2020191171 and WO2021 / 226558 can be used in the practice of the present invention.
[0371] Available CRISPR-based methods for editing long sequences, such as homology-directed repair (HDR) or microhomology-mediated end joining (MMEJ), rely on Cas9 cleavage and often result in random indel formation within the genome. Many recent efforts, such as chemical enhancers, enhancement domain fusions, and modified donor DNA, have enhanced precise long-sequence editing. Nickel-based HDR has been shown to reduce editing errors but may reduce efficiency. Therefore, there remains a need for efficient and safe CRISPR editing tools for long-sequence modification.
[0372] Bacteriophages have evolved enzymes that utilize accessible replicating genomic DNA to perform precise recombination. Applicant reasoned that a key enzyme for microbial recombination, namely single-strand annealing protein (SSAP), could be useful for gene editing in mammalian cells, which does not explicitly cleave DNA and does not rely on the error-prone pathway required for Cas9 editing. Motivated by this hypothesis and Applicant's previous work demonstrating its ability to stimulate genomic recombination, Applicant developed a gene editing tool using inactivated Cas9 (dCas9, or catalytically inactive Cas9) and microbial SSAP. This dCas9 editor uses SSAP for knock-in editing when supplied with donor DNA, without the need for genomic DNA cleavage. Applicant named it the dCas9-SSAP editor (dCas9-SSAP).
[0373] To optimize dCas9-SSAP, Applicant conducted a metagenomic search for SSAPs, focusing on RecT homologs, and identified EcRecT as the most efficient for human genome knock-in. For validation, Applicant conducted a series of genome manipulation and chemical perturbation experiments. Applicant's data showed that dCas9-SSAP had knock-in efficiency comparable to that of the wild-type Cas9 reference and significantly higher than that of the Cas9 nickase editor. For kilobase-scale sequence editing, dCas9-SSAP achieved knock-in efficiencies of up to 12% across multiple genome targets and cell lines without selection. More importantly, Applicant's data demonstrated that this new tool generates nearly zero on-target and off-target errors. In assays for 1 kb sequence knock-in, dCas9-SSAP had less than 0.3% editing errors across all cells, while the Cas9 editor had a similar yield but an additional 10%-16% of incorrectly edited cells. Across the loci tested, dCas9-SSAP had editing accuracy of 90%-99.6%, while the accuracy of the Cas9 editor ranged from 10%-38% (Figure 39F).
[0374] Furthermore, Applicant investigated the mechanism of dCas9-SSAP editing by inhibiting several DNA repair enzymes and performing cell cycle synchronization. These experiments demonstrated that dCas9-SSAP, in contrast to Cas9 editing, is less dependent on endogenous DNA repair pathways. The results of Applicant's cell cycle assays supported the hypothesized mechanism of the dCas9 editor; they are consistent with the known biophysical and biochemical properties of dCas9.
[0375] Finally, to aid in the delivery of dCas9-SSAP for future applications, we optimized its molecular design using structurally guided truncation to obtain a minimized dSaCas9-mSSAP, achieving a greater than 50% reduction in size while maintaining a similar level of efficiency. This minimal dCas9 editor allows for convenient delivery using viral vectors such as adeno-associated virus (AAV), potentially useful for difficult-to-transfect cell types or in vivo applications. Overall, the dCas9-SSAP editor is capable of efficient and precise knock-in genome manipulation. With room for further improvement, it has potential research and therapeutic value as a cut-free gene editing tool for mammalian cells.
[0376] Use of phage SSAPs for dCas9 knock-in gene editing. Most CRISPR-based editors capable of knocking in long sequences require SSNs or DSBs, which can trigger competing, error-prone NHEJ pathways, resulting in variable efficiency and accuracy. In contrast, bacteriophages have evolved DNA-modifying enzymes to integrate into the genome of host bacteria via sequence homology, e.g., lambda Red. Such precise phage integration relies on a key homology-directed step: recombination between genomic DNA and donor DNA is stimulated by an SSAP, e.g., lambda Bet or its functional homolog, RecT. From previous work, Applicant reasoned that phage SSAPs may not rely on DNA cleavage due to their unusual ATP-independent activity, in contrast to the ATP-dependent RAD51 protein in human cells. The high affinity of phage SSAPs for single- and double-stranded DNA may enable attachment to donor templates when multiple SSAPs are recruited to genomic targets via RNA-guided dCas9. The target DNA strand then becomes transiently accessible during dCas9-mediated DNA unwinding and R-loop formation, which can facilitate genome-donor DNA exchange without cleavage.
[0377] Based on this hypothesis, we designed a system to recruit SSAPs to catalytically inactive Cas9 (dCas9) (Figure 38A). Although the dCas9 protein cannot cleave DNA, it retains the ability to unwind the target site and form an R-loop, presumably making the non-target strand accessible for SSAP-stimulated homologous recombination. To test this, we engineered and evaluated three major types of microbial SSAPs: the lambda Bet protein (lambda bet); the Escherichia coli Rac prophage RecT (Rac RecT); and the phage T7 gp2.5 (T7 gp2.5). We recruited these SSAPs to an inactivated version of S. pyogenes Cas9 (dSpCas9, hereafter abbreviated as dCas9) via an RNA aptamer MS2 stem-loop (Figures 38A and 38C). This MS2-aptamer was inserted into the sgRNA scaffold, and candidate SSAPs were fused to the N-terminal MS2 coat protein (MCP), which specifically binds to the MS2 aptamer, thereby allowing multiple SSAPs to form complexes with the dCas9-guide RNA. To measure their gene editing activity in human cells, Applicant generated knock-in donors with an 800-bp transgene encoding a fluorescent protein (FP) cassette flanked by homology arms (HA) that allowed in-frame insertion of FPs into housekeeping genes such as DYNLT1, HSP90AA1, and ACTB (Figure 38B, left). Upon accurate knock-in, Applicant quantified gene editing efficiency by measuring the percentage of FP-expressing cells (Figures 38B-D). Applicant's initial studies identified RecT as having higher knock-in editing activity in human cells compared to other SSAPs, whereas no editing above background was observed with dCas9 alone or non-targeting controls (Figures 38C, D). Applicants verified this knock-in editing using gel electrophoresis and sequencing (Figure 44), providing evidence that coupling SSAP to dCas9 via an RNA aptamer enables knock-in gene editing.
[0378] Development of dCas9-SSAP as a mammalian gene editing tool. Applicant conducted metagenomic mining to identify the best SSAP for mammalian gene editing. Applicant focused on RecT homologs and sought to maximize evolutionary diversity through phylogenetic analysis. Applicant systematically searched the NCBI non-redundant sequence database for RecT homologs and identified 2,071 initial candidates. Applicant then constructed a phylogenetic tree, filtered out proteins with high sequence homology, and subsampled evolutionary branches, resulting in 16 highly diverse SSAP candidates (Figure 44).
[0379] Applicants examined SSAP candidates through knock-in screening and evaluated their editing efficiency across three genomic loci: HSP90AA1, DYNL...
Claims
1. A system comprising: (i) a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence; (ii) a recombinant protein comprising: a recombination protein comprising an exonuclease, a single-stranded DNA annealing protein (SSAP), or a single-stranded DNA binding protein (SSB), or a combination of two or more thereof; or (iii) a nucleic acid molecule(s) encoding or delivering (i) and / or (ii) for in vivo expression in a cell; or (iv) a vector(s) containing the nucleic acid molecule(s) of (iii) for in vivo expression in a cell; and Including, The system does not comprise a Cas protein or a nucleic acid encoding a Cas protein. system.
2. at least one RNA or peptide aptamer; an aptamer-binding protein operably linked to said recombinant protein as part of a fusion protein; 10. The system of claim 1, further comprising a mobilization system comprising:
3. 2. The system of claim 1, wherein the recombinant protein comprises a recombinant protein of Table 12, or a derivative, variant, or functional portion thereof, wherein the recombinant protein, or derivative or variant thereof, comprises an amino acid sequence having at least 70% similarity or identity to an amino acid sequence of Table 12.
4. 2. The system of claim 1, wherein the recombination protein comprises RecE, RecT, or a derivative or variant thereof, and the RecE, RecT, or derivative or variant thereof comprises an amino acid sequence having at least 70% identity or similarity or identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-14.
5. A cell comprising the system of claim 1 or a composition comprising same.
6. 10. The system or composition comprising the system of claim 1 for use in a method for modifying a target genomic DNA sequence in a cell containing the target genomic DNA sequence, the method comprising the step of introducing the system or composition into the cell.
7. The system or composition of claim 6 , wherein said introducing into a cell comprises administering to a subject.
8. A system comprising: (i) nucleic acid polymerase(s); (ii) a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence and an RNA for reverse transcription, or a plurality of nucleic acid molecules comprising a guide RNA sequence complementary to a target DNA sequence and an RNA for reverse transcription; (iii) a recombinant protein, a recombination protein comprising an exonuclease, a single-stranded DNA annealing protein (SSAP), or a single-stranded DNA binding protein (SSB), or a combination of two or more thereof; or (iv) a nucleic acid molecule(s) encoding or delivering (i) and / or (ii) and / or (iii) for in vivo expression in a cell; or (v) a vector(s) containing the nucleic acid molecule(s) of (iv) for in vivo expression in a cell; and A system including:
9. A Cas protein, a nucleic acid molecule(s) encoding or delivering a Cas protein for in vivo expression in a cell; vector(s) containing nucleic acid molecule(s) encoding a Cas protein; and / or It is a mobilization system, at least one RNA or peptide aptamer; an aptamer-binding protein operably linked to said recombinant protein as part of a fusion protein; Mobilization system, including 9. The system of claim 8 further comprising:
10. 9. The system of claim 8, wherein the recombinant protein comprises a recombinant protein of Table 12, or a derivative or variant or functional portion thereof, wherein the recombinant protein, or derivative or variant thereof, comprises an amino acid sequence having at least 70% similarity or identity to an amino acid sequence of Table 12.
11. 9. The system of claim 8, wherein the recombination protein comprises RecE, RecT, or a derivative or variant thereof, and the RecE, RecT, or derivative or variant thereof comprises an amino acid sequence having at least 70% identity or similarity or identity to an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-14.
12. 10. The system of claim 9, wherein the Cas protein is a nickase, is catalytically inactive, or is catalytically inactive.
13. The system of claim 8 , wherein the nucleic acid polymerase comprises reverse transcriptase activity.
14. 10. The system of claim 9, wherein one or more of the nucleic acid polymerase, the Cas protein, the recombination protein, and the aptamer-binding protein are operably linked to each other and comprise a fusion protein.
15. A cell comprising the system of claim 8.
16. 9. The system or composition comprising the system of claim 8 for use in a method for modifying a target genomic DNA sequence in a cell containing the target genomic DNA sequence, the method comprising the step of introducing the system or composition into the cell.
17. 9. The system of claim 8 or a composition comprising the same for modifying a target DNA sequence in a cell.
18. 1. A system or composition for use in a recombinant method, the method comprising providing the system or composition in a cell, the system or composition comprising: (i) a nucleic acid molecule comprising a guide RNA sequence complementary to a target DNA sequence, wherein the target DNA sequence comprises a genomic DNA sequence in the cell; and (ii) a recombinant protein comprising: a recombination protein comprising an exonuclease, a single-stranded DNA annealing protein (SSAP), or a single-stranded DNA binding protein (SSB), or a combination of two or more thereof; or (iii) a nucleic acid molecule(s) encoding or delivering (i) and / or (ii) for in vivo expression in a cell; or (iv) a vector(s) containing the nucleic acid molecule(s) of (iii) for in vivo expression in a cell; and Including, System or composition.
19. Cas protein and / or reverse transcriptase (RT), a nucleic acid molecule(s) encoding or delivering a Cas protein and / or a reverse transcriptase (RT) for in vivo expression in said cell; vector(s) containing nucleic acid molecule(s) encoding a Cas protein and / or a reverse transcriptase (RT), and / or It is a mobilization system, at least one aptamer; an aptamer-binding protein operably linked to said recombinant protein as part of a fusion protein; Mobilization system, including 20. The system or composition of claim 18, further comprising:
20. 19. The system or composition of claim 18, wherein the target DNA sequence comprises the genomic sequence of albumin (ALB), AAVS1, HSP90AA1, DYNLT1, ACTB, BCAP31, HIST1H2BK, CLTA, or RAB11A.