Cas endonucleases and related methods

By developing fusion proteins or conjugates of highly identical Cas endonucleases with heterologous moieties, the problems of editing efficiency and specificity of the CRISPR-Cas system in eukaryotic cells have been solved, enabling effective treatment of genetic diseases.

CN121729487APending Publication Date: 2026-03-24FLAGSHIP ENTREPRENEURSHIP & INNOVATION NO 7 CO LTD
View PDF 257 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

The efficiency and specificity of existing CRISPR-Cas systems in nucleic acid editing in eukaryotic cells need to be improved, and they are difficult to effectively treat genetic diseases.

Method used

Novel Cas endonucleases and their functional fragments, variants, and domains are provided, which can bind heterologous moieties to form fusion proteins or conjugates by Cas endonucleases with high identity to the amino acid sequences shown in Table 1, for editing nucleic acid molecules and treating diseases.

Benefits of technology

It improves the efficiency and specificity of nucleic acid editing, enhances its targeting ability in eukaryotic cells, and can effectively treat genetic diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

Provided herein are Cas endonucleases (and functional fragments, functional variants, and domains thereof), nucleic acid molecules encoding the same, and systems comprising the same. The disclosure further relates to methods of utilizing the Cas endonucleases (or nucleic acid molecules encoding them), including, for example, methods of editing nucleic acid molecules (e.g., genes) and methods of treating diseases (e.g., genetic diseases).
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims priority to Greek Patent Application No. 20230100610, filed July 25, 2023, and U.S. Serial No. 63 / 515,768, filed July 26, 2023, the entire contents of each of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present disclosure relates to Cas endonucleases (and functional fragments, functional variants, and domains thereof), nucleic acid molecules encoding the same, and systems comprising the same. The present disclosure further relates to methods of utilizing the Cas endonucleases (or nucleic acid molecules encoding the same), including, for example, in methods of editing nucleic acid molecules (e.g., genes) and methods of treating diseases (e.g., genetic diseases). BACKGROUND

[0003] The CRISPR (clustered regularly interspaced short palindromic repeat)-Cas (CRISPR-associated protein) system is an adaptive immune system of many prokaryotes (e.g., bacteria and archaea) that functions to prevent infection (e.g., by bacteriophages, viruses, and other foreign genetic elements). A typical naturally occurring CRISPR-Cas system comprises a CRISPR RNA (crRNA), a trans-activating CRISPR RNA (tracrRNA), and a Cas endonuclease, wherein the tracrRNA mediates binding to the Cas endonuclease, the crRNA directs the Cas endonuclease to a target nucleic acid molecule, and the Cas endonuclease mediates cleavage of the target nucleic acid molecule (e.g., viral DNA). CRISPR-Cas systems have been engineered and modified for, e.g., nucleic acid (e.g., gene) editing in eukaryotic cells. SUMMARY

[0004] Provided herein, inter alia, are novel Cas endonucleases and polynucleotides encoding the same; fusions and conjugates comprising the Cas endonucleases; methods of manufacture; pharmaceutical compositions; and methods of use, including, e.g., methods of editing nucleic acid molecules (e.g., genes) and methods of treating diseases (e.g., genetic diseases).

[0005] Accordingly, in one aspect, provided herein is a Cas endonuclease (or functional fragment, functional variant, or domain thereof) comprising an amino acid sequence that is at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of any Cas endonuclease set forth in Table 1 or set forth in any one of SEQ ID NOs: 1-320.

[0006] In some embodiments, the amino acid sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence of any Cas endonuclease set forth in Table 1 or any of SEQ ID NOS: 1-320. In some embodiments, the amino acid sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence of any Cas endonuclease set forth in Table 1 or any of SEQ ID NOS: 1-320.

[0007] In some embodiments, the amino acid sequence of the Cas endonuclease has less than 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% identity to the amino acid sequence of the reference Cas endonuclease set forth in SEQ ID NO: 321. In some embodiments, the amino acid sequence of the Cas endonuclease has less than 90% (e.g., 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 60%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%) and greater than 50% (e.g., 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%) identity to the amino acid sequence of the reference Cas endonuclease set forth in SEQ ID NO: 321. In some embodiments, the amino acid sequence of the Cas endonuclease has less than 90% (e.g., 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%) and greater than 76% (e.g., 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%) identity to the amino acid sequence of the reference Cas endonuclease set forth in SEQ ID NO: 321.

[0008] In some embodiments, the Cas endonuclease has one or more (e.g., 1, 2, 3, 4, 5, and / or 6) of the following properties (or is engineered to have one or more of the following properties): (a) the ability to mediate a double-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule; (b) the ability to mediate a single-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule; (c) the inability to mediate a double-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule; (d) the ability to mediate a single-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule and the inability to mediate a double-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule (i.e., nickase activity); (f) DNA endonuclease activity; and / or (g) RNA-guided DNA endonuclease activity.

[0009] In some embodiments, the amino acid sequence of the Cas endonuclease comprises one or more amino acid variations (e.g., substitutions, deletions, additions). In some embodiments, the one or more amino acid variations (e.g., substitutions, deletions, additions) reduce or eliminate the ability of the Cas endonuclease to mediate a double-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule. In some embodiments, the modified Cas endonuclease comprising the one or more amino acid variations (e.g., substitutions, deletions, additions) has the ability to mediate a single-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule and does not have the ability to mediate a double-stranded break in a target double-stranded nucleic acid (e.g., DNA) molecule (i.e., nickase activity). In some embodiments, the one or more amino acid variations (e.g., substitutions, deletions, additions) change the PAM nucleotide sequence recognized by the Cas endonuclease. In some embodiments, the one or more amino acid variations (e.g., substitutions, deletions, additions) (a) reduce the Cas endonuclease activity of the endonuclease by at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% relative to the endonuclease lacking the one or more amino acid variations (e.g., substitutions, deletions, additions); or (b) enhance the Cas endonuclease activity of the endonuclease by at least 1-fold, 2-fold, 5-fold, 10-fold, or 100-fold relative to the Cas endonuclease lacking the one or more amino acid variations (e.g., substitutions, deletions, additions).

[0010] In some embodiments, the Cas endonuclease further comprises one or more heterologous moieties (e.g., heterologous proteins). In some embodiments, the Cas endonuclease comprises 2, 3, 4, or 5 or more heterologous moieties. In some embodiments, the heterologous moieties are attached to the N-terminus, C-terminus, and / or internally between the N-terminus and C-terminus of the endonuclease. In some embodiments, the heterologous moieties (e.g., heterologous proteins) are directly attached to the endonuclease. In some embodiments, the heterologous moieties (e.g., heterologous proteins) are indirectly attached to the Cas endonuclease. In some embodiments, the heterologous moieties (e.g., heterologous proteins) are indirectly attached to the Cas endonuclease via a linker. In some embodiments, the heterologous moieties are a peptide, a protein, a carbohydrate, a lipid, a polymer, or a small molecule. In some embodiments, the heterologous moieties are a nuclear localization signal (NLS), a tag, and / or a reporter gene.

[0011] In an aspect, provided herein is a conjugate comprising a Cas endonuclease described herein and one or more heterologous moieties.

[0012] In some embodiments, the heterologous moieties are a protein, a peptide, a small molecule, a nucleic acid molecule (e.g., DNA, RNA, DNA / RNA hybrid molecule), a carbohydrate, a lipid, or a synthetic polymer. In some embodiments, the heterologous moieties are operably linked to the N-terminus, C-terminus, and / or internally between the N-terminus and C-terminus of the Cas endonuclease. In some embodiments, the heterologous moieties are directly operably linked to the Cas endonuclease. In some embodiments, the heterologous moieties are indirectly operably linked to the Cas endonuclease. In some embodiments, the heterologous moieties are indirectly operably linked to the Cas endonuclease via a linker.

[0013] In an aspect, provided herein are fusion proteins comprising a Cas endonuclease described herein and one or more heterologous proteins. In some embodiments, the heterologous protein is fused to the N-terminus, the C-terminus, and / or internally between the N-terminus and the C-terminus of the Cas endonuclease. In some embodiments, the heterologous protein is directly fused to the Cas endonuclease. In some embodiments, the heterologous protein is indirectly fused to the Cas endonuclease. In some embodiments, the heterologous protein is indirectly fused to the Cas endonuclease via a peptide linker. In some embodiments, the heterologous protein exhibits polymerase (e.g., reverse transcriptase) activity, nucleobase editing activity (e.g., deaminase activity), methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, or double-stranded DNA cleavage activity and nucleic acid binding activity, or any combination of the foregoing.

[0014] In some embodiments, the heterologous protein is a polymerase. In some embodiments, the polymerase has RNA-dependent DNA polymerase activity. In some embodiments, the polymerase is a reverse transcriptase (or a functional fragment, functional variant, or domain thereof). In some embodiments, the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) is derived from a retrovirus or a retrotransposon. In some embodiments, the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises an amino acid sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of a protein set forth in Table 2 or set forth in any one of SEQ ID NOs: 324-476.

[0015] In some embodiments, the heterologous polypeptide is a nucleobase editor. In some embodiments, the nucleobase editor is a deaminase (or a functional fragment, functional variant, or domain thereof). In some embodiments, the deaminase (or a functional fragment, functional variant, or domain thereof) exhibits adenosine deaminase activity and / or cytidine deaminase activity. In some embodiments, the deaminase (or a functional fragment, functional variant, or domain thereof) comprises an amino acid sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of a protein set forth in Table 3 or set forth in any one of SEQ ID NOs: 477-536. In some embodiments, the nucleobase editor is fused to a base excision repair inhibitor (or a functional fragment or functional variant thereof) (e.g., a uracil glycosylase inhibitor (UGI), a nuclease-dead, inosine-specific nuclease (dISN)).

[0016] In an aspect, provided herein is a nucleic acid molecule encoding a Cas endonuclease described herein, a conjugate described herein, or a fusion protein described herein. In some embodiments, the nucleic acid molecule is a DNA or RNA (e.g., mRNA) molecule. In some embodiments, the nucleic acid molecule is codon-optimized. In some embodiments, the nucleic acid molecule further comprises one or more transcriptional or translational regulatory elements (e.g., a promoter, an enhancer (e.g., a cell- or tissue-specific transcriptional regulatory element)). In some embodiments, the nucleic acid molecule further encodes one or more gRNAs (e.g., crRNAs, tracrRNAs, sgRNAs, template RNAs (e.g., as described herein)).

[0017] In an aspect, provided herein is a vector comprising a nucleic acid molecule described herein. In some embodiments, the vector is a viral vector or a non-viral vector (e.g., a plasmid, a minicircle). In some embodiments, the vector is a viral vector (e.g., an adeno-associated virus (AAV) vector, a lentivirus vector, an adenovirus vector).

[0018] In an aspect, provided herein is a carrier comprising a Cas endonuclease described herein, a conjugate described herein, a fusion protein described herein, a nucleic acid molecule described herein, and / or a vector described herein. In some embodiments, the carrier is a nanoparticle, a polymer, a virus (e.g., a recombinant virus), a virus-like particle, a virosome, a fusosome, a vesicle, or a lipid-based carrier. In some embodiments, the carrier is a recombinant virus (e.g., an adeno-associated virus (AAV), a lentivirus, an adenovirus). In some embodiments, the carrier is a lipid-based carrier. In some embodiments, the lipid-based carrier is a lipid nanoparticle (LNP), a liposome, a lipoplex, a nanoliposome, an exosome, or a micelle. In some embodiments, the carrier further comprises one or more gRNAs (e.g., crRNAs, tracrRNAs, sgRNAs, template RNAs (e.g., as described herein)).

[0019] In an aspect, provided herein is a reaction mixture comprising (a) a cell (e.g., a cell comprising a target nucleic acid molecule) or a target nucleic acid molecule; and (b) a Cas endonuclease described herein, a conjugate described herein, a fusion protein described herein, a nucleic acid molecule described herein, a vector described herein, a carrier described herein, and / or a pharmaceutical composition described herein.

[0020] In an aspect, provided herein is a cell comprising a Cas endonuclease described herein, a conjugate described herein, a fusion protein described herein, a nucleic acid molecule described herein, a vector described herein, a reaction mixture described herein, a carrier described herein, and / or a pharmaceutical composition described herein.

[0021] In an aspect, provided herein is a pharmaceutical composition comprising a Cas endonuclease described herein, a conjugate described herein, a fusion protein described herein, a nucleic acid molecule described herein, a vector described herein, a reaction mixture described herein, a carrier described herein, and / or a cell described herein; and a pharmaceutically acceptable excipient.

[0022] In an aspect, provided herein is a kit comprising a Cas endonuclease described herein, a conjugate described herein, a fusion protein described herein, a nucleic acid molecule described herein, a vector described herein, a reaction mixture described herein, a carrier described herein, a cell described herein, and / or a pharmaceutical composition described herein; and optionally instructions for using any one or more of the foregoing.

[0023] In an aspect, provided herein is a system for modifying a target nucleic acid (e.g., DNA) molecule, the system comprising: (a) a Cas endonuclease described herein, a conjugate described herein, a fusion protein described herein, a nucleic acid molecule described herein, a vector described herein, a carrier described herein, a reaction mixture described herein, a cell described herein, and / or a pharmaceutical composition described herein, and (b) a first gRNA (e.g., crRNA and tracrRNA; sgRNA; pegRNA, a template RNA (e.g., as described herein)) or a nucleic acid (e.g., DNA) molecule encoding the first gRNA (e.g., crRNA and tracrRNA; sgRNA; a template RNA (e.g., as described herein)).

[0024] In some embodiments, the system has one or more of the following features: (a) the Cas endonuclease of the system is capable of binding to the first gRNA; (b) the Cas endonuclease of the system is capable of forming a break in a target nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; (c) the Cas endonuclease of the system is capable of forming a single-strand break in a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; (d) the Cas endonuclease of the system is capable of forming a single-strand break in a modified strand (as defined herein) of a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; (e) the Cas endonuclease of the system is capable of forming a double-strand break in a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; (f) the Cas endonuclease of the system is not capable of forming a double-strand break in a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; (g) the Cas endonuclease of the system is capable of forming a single-strand break in a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule and is not capable of forming a double-strand break in a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; (h) the Cas endonuclease of the system is capable of forming a single-strand break in a modified strand (as defined herein) of a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule and is not capable of forming a double-strand break in a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; and / or (i) the system is capable of editing a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., mediating addition of one or more nucleotides to the target nucleic acid, deletion of one or more nucleotides from the target nucleic acid, or substitution of one or more nucleotides in the target nucleic acid).

[0025] In some embodiments, the system is capable of editing a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., mediating addition of one or more nucleotides to the target nucleic acid, deletion of one or more nucleotides from the target nucleic acid, or substitution of one or more nucleotides in the target nucleic acid).

[0026] In some embodiments, the system is capable of editing a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., mediating addition of one or more nucleotides to the target nucleic acid, deletion of one or more nucleotides from the target nucleic acid, or substitution of one or more nucleotides in the target nucleic acid) with increased efficiency relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., a reference Cas endonuclease set forth in SEQ ID NO: 321)).

[0027] In some embodiments, the system is capable of editing a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., mediating addition of one or more nucleotides to the target nucleic acid, deletion of one or more nucleotides from the target nucleic acid, or substitution of one or more nucleotides in the target nucleic acid) with at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200% increased efficiency relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., a reference Cas endonuclease set forth in SEQ ID NO: 321)).

[0028] In some embodiments, the system is capable of editing a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., mediating addition of one or more nucleotides to the target nucleic acid, deletion of one or more nucleotides from the target nucleic acid, or substitution of one or more nucleotides in the target nucleic acid) with at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100% increased efficiency relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., a reference Cas endonuclease set forth in SEQ ID NO: 321)).

[0029] In some embodiments, the system is capable of editing a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., mediating the addition of one or more nucleotides to the target nucleic acid, the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid) with an efficiency that is increased by about 30-200%, 40-200%, 50-200%, 60-200%, 70-200%, 80-200%, 90-200%, 100-200%, 150-200%, 30-150%, 40-150%, 50-150%, 60-150%, 70-150%, 80-150%, 90-150%, 100-150%, 30-100%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, or 90-100% relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., a reference Cas endonuclease set forth in SEQ ID NO: 321)).

[0030] In some embodiments, the target nucleic acid molecule is a DNA molecule. In some embodiments, the target nucleic acid molecule is a double-stranded DNA (dsDNA) molecule. In some embodiments, a portion of the nucleotide sequence of an unmodified strand (as defined herein) of the target dsDNA molecule is complementary to at least a portion of the nucleotide sequence of the first gRNA. In some embodiments, the target nucleic acid molecule is within a genome of a cell (e.g., a eukaryotic cell) (e.g., within a subject (e.g., a human subject), a plant).

[0031] In some embodiments, (b) comprises the first gRNA (e.g., a crRNA and a tracrRNA; or a template RNA (e.g., as described herein)). In some embodiments, (b) comprises a nucleic acid (e.g., DNA) molecule encoding the first gRNA.

[0032] In some embodiments, at least a portion of the nucleotide sequence of the first gRNA is complementary to a portion of the nucleotide sequence of the target nucleic acid molecule (e.g., a gene). In some embodiments, at least a portion of the nucleotide sequence of the first gRNA is complementary to a portion of the nucleotide sequence of an unmodified strand (as defined herein) of a dsDNA target nucleic acid molecule (e.g., a gene). In some embodiments, at least a portion of the nucleotide sequence of the first gRNA binds to a portion of the nucleotide sequence of an unmodified strand (as defined herein) of a dsDNA target nucleic acid molecule (e.g., a gene).

[0033] In some embodiments, the first gRNA comprises an sgRNA (e.g., a single sgRNA, multiple different sgRNAs). In some embodiments, the first gRNA comprises a crRNA (e.g., a single crRNA, multiple different crRNAs) and a tracrRNA (e.g., a single tracrRNA, multiple different tracrRNAs), wherein the crRNA and the tracrRNA are on separate RNA nucleic acid molecules (or are encoded by separate nucleic acid (e.g., DNA) molecules).

[0034] In some embodiments, the first gRNA comprises a template RNA (e.g., a single template RNA, multiple different template RNAs) comprising (e.g., from 5' to 3') a crRNA, a tracrRNA, a heterologous object sequence, and a 3' target homology domain. In some embodiments, the template RNA further comprises a sequence that binds a polymerase (e.g., a reverse transcriptase). In some embodiments, the template RNA comprises (e.g., from 5' to 3') a crRNA, a tracrRNA, a sequence that binds a polymerase (e.g., a reverse transcriptase), a heterologous object sequence, and a 3' target homology domain.

[0035] In some embodiments, the first gRNA comprises one or more nucleotides comprising one or more chemical modifications (e.g., base, ribose, and / or internucleotide linkage chemical modifications) (i.e., modified nucleotides). In some embodiments, the modified nucleotides comprise 2'-O-methyl (2'-OMe); 2' O-methoxyethyl (2'-O-MOE); 2' deoxy-2'-fluoro (2'-F); 2'-arabino-fluoro (2'-Ara-F); 2'-O-benzyl; 2'-O-methyl-4-pyridine (2-O-methyl-4-pyridine (2'-O-CH2Py(4)); 2'F-4'-Calpha-OMe; or 2',4'-di-Calpha-OMe, 2'-O-methyl-3'-thioPACE, and / or S-restricted ethyl (cEt). In some embodiments, the modified nucleotides comprise chemically modified internucleosidic (or nucleosidic) linkages. In some embodiments, the modified internucleosidic (or nucleosidic) linkages comprise phosphorothioate (e.g., chiral phosphorothioate), phosphorodithioate, phosphotriester, aminoalkylphosphotriester, alkyl (e.g., methyl) phosphonate (e.g., 3'-alkylene phosphonate, chiral phosphonate), phosphinate, phosphoramidate (e.g., 3'-amino amino phosphoramidate, aminoalkyl amino phosphoramidate), thiocarbonyl amino phosphoramidate, thiocarbonyl alkyl phosphonate, thiocarbonyl alkyl phosphoramidate, or boranophosphane.

[0036] In some embodiments, the first gRNA (e.g., template RNA, sgRNA) comprises a nucleic acid molecule comprising a toe-loop, hairpin, stem-loop, pseudoknot (e.g., Mpknotl portion), aptamer, G-quadruplex, tRNA, riboswitch, or ribozyme. In some embodiments, the first gRNA (e.g., template RNA, sgRNA), wherein the nucleic acid molecule is a pseudoknot (e.g., Mpknotl portion).

[0037] In some embodiments, the system further comprises a second gRNA (or a nucleic acid (e.g., DNA) molecule encoding the gRNA) that directs an endonuclease of the system to form a single-strand break in the unedited strand of the target dsDNA molecule. In some embodiments, at least a portion of the nucleotide sequence of the second gRNA is complementary to a portion of the nucleotide sequence of the edited strand (as defined herein) of the dsDNA target nucleic acid molecule. In some embodiments, at least a portion of the nucleotide sequence of the second gRNA binds to a portion of the nucleotide sequence of the edited strand (as defined herein) of the dsDNA target nucleic acid molecule. In some embodiments, the second gRNA is present on the same nucleic acid molecule as the first gRNA (or a nucleic acid (e.g., DNA) molecule encoding the second gRNA is present on the same nucleic acid (e.g., DNA) molecule encoding the first gRNA). In some embodiments, the second gRNA is present on a different nucleic acid molecule than the first gRNA (or a nucleic acid (e.g., DNA) molecule encoding the second gRNA is present on a different nucleic acid (e.g., DNA) molecule encoding the first gRNA).

[0038] In some embodiments, the system further comprises a donor template nucleic acid (e.g., DNA) molecule (e.g., as defined herein).

[0039] In an aspect, provided herein is a system for modifying a dsDNA molecule, comprising: (a) a fusion protein described herein or a nucleic acid molecule (e.g., DNA, RNA molecule) encoding the fusion protein; and (b) a template RNA (e.g., a single template RNA, a plurality of different template RNAs) comprising (e.g., from 5' to 3') a crRNA, a tracrRNA, a heterologous object sequence, and a 3' target homology domain; or a nucleic acid molecule (e.g., DNA molecule) encoding the template RNA.

[0040] In an aspect, provided herein is a nucleic acid molecule encoding a system described herein. In some embodiments, the nucleic acid molecule is a DNA or RNA (e.g., mRNA) molecule. In some embodiments, the nucleic acid molecule is codon-optimized. In some embodiments, the nucleic acid molecule further comprises one or more transcriptional or translational regulatory elements (e.g., a promoter, an enhancer (e.g., a cell- or tissue-specific transcriptional regulatory element)).

[0041] In an aspect, provided herein is a vector comprising a nucleic acid molecule described herein. In some embodiments, the vector is a viral vector or a non-viral vector (e.g., a plasmid, a minicircle). In some embodiments, the vector is a viral vector (e.g., an adeno-associated virus (AAV) vector, a lentivirus vector, an adenovirus vector).

[0042] In an aspect, provided herein is a carrier comprising a system described herein, a nucleic acid molecule described herein, and / or a vector described herein. In some embodiments, the carrier is a nanoparticle, a polymer, a virus (e.g., a recombinant virus), a virus-like particle, a virosome, a fusosome, a vesicle, or a lipid-based carrier. In some embodiments, the carrier is a recombinant virus (e.g., an adeno-associated virus (AAV), a lentivirus, an adenovirus). In some embodiments, the carrier is a nanoparticle. In some embodiments, the carrier is a lipid-based carrier. In some embodiments, the lipid-based carrier is a lipid nanoparticle (LNP), a liposome, a lipoplex, a nanoliposome, an exosome, or a micelle. In some embodiments, the carrier further comprises one or more gRNAs (e.g., crRNAs, tracrRNAs, sgRNAs, template RNAs (e.g., as described herein)).

[0043] In an aspect, provided herein is a reaction mixture comprising (a) a cell (e.g., a cell comprising a target nucleic acid molecule) or a target nucleic acid molecule; and (b) a system described herein, a nucleic acid molecule described herein, a vector described herein, and / or a carrier described herein.

[0044] In an aspect, provided herein is a cell comprising a system described herein, a nucleic acid molecule described herein, a vector described herein, a carrier described herein, and / or a reaction mixture described herein.

[0045] In an aspect, provided herein is a pharmaceutical composition comprising a system described herein, a nucleic acid molecule described herein, a vector described herein, a carrier described herein, a reaction mixture, and / or a cell described herein; and a pharmaceutically acceptable excipient.

[0046] In one aspect, this document provides a kit comprising the systems, nucleic acid molecules, vectors, loading agents, reaction mixtures, cells, and / or pharmaceutical compositions described herein; and optionally, instructions for use with any one or more of the foregoing.

[0047] In one aspect, this document provides a method for delivering a Cas endonuclease, fusion protein, conjugate, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition to a cell, the method comprising introducing the Cas endonuclease, conjugate, fusion protein, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition described herein into the cell, thereby delivering the Cas endonuclease, fusion protein, conjugate, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition to the cell.

[0048] In some embodiments, the cells are in vitro, ex vivo, or in vivo. In some embodiments, the cells are euploid, not immortalized, part of a tissue, part of an organism, primary cells, non-dividing, haploid (e.g., germline cells), non-cancerous polyploid cells, or derived from a subject with a genetic disease. In some embodiments, the cells are in a subject (e.g., a human subject). In some embodiments, the cells are in a human subject.

[0049] In one aspect, this document provides methods for delivering Cas endonucleases, fusion proteins, conjugates, systems, nucleic acid molecules, vectors, loading agents, reaction mixtures, cells, or pharmaceutical compositions to cells, said methods comprising the Cas endonucleases, conjugates, fusion proteins, systems, nucleic acid molecules, vectors, loading agents, reaction mixtures, cells, or pharmaceutical compositions described herein, thereby delivering said Cas endonucleases, fusion proteins, conjugates, systems, nucleic acid molecules, vectors, loading agents, reaction mixtures, cells, or pharmaceutical compositions to said subject (e.g., human subject).

[0050] In one aspect, this document provides a method for cleaving target sites in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))), the method comprising contacting the cell with a Cas endonuclease, a conjugate, a fusion protein, a system, a nucleic acid molecule, a vector, a delivery agent, a reaction mixture, a cell, or a pharmaceutical composition described herein, thereby cleaving the target site in the target nucleic acid (e.g., DNA) molecule.

[0051] In one aspect, this document provides a method for editing target sites in target nucleic acid (e.g., DNA) molecules (e.g., double-stranded target nucleic acid sequences (e.g., dsDNA (e.g., genomic dsDNA))), the method comprising contacting the cell with the Cas endonuclease described herein, the conjugate described herein, the fusion protein described herein, the system described herein, the nucleic acid molecule described herein, the vector described herein, the loading agent described herein, the reaction mixture described herein, the cell described herein, or the pharmaceutical composition described herein, thereby editing the target site in the target nucleic acid (e.g., DNA) molecule.

[0052] In one aspect, this document provides a method for editing target sites in genomic dsDNA of cells, the method comprising contacting the Cas endonuclease described herein, the conjugate described herein, the fusion protein described herein, the system described herein, the nucleic acid molecule described herein, the vector described herein, the loading agent described herein, the reaction mixture described herein, the cell described herein, or the pharmaceutical composition described herein, thereby editing the target sites in the genomic DNA of the cells.

[0053] In some embodiments, the cells are in vitro, ex vivo, or in vivo. In some embodiments, the cells are euploid, not immortalized, part of a tissue, part of an organism, primary cells, non-dividing, haploid (e.g., germline cells), non-cancerous polyploid cells, or derived from a subject with a genetic disease. In some embodiments, the cells are in a subject (e.g., a human subject). In some embodiments, the cells are in a human subject.

[0054] In one aspect, this article provides a method for editing target sites in dsDNA molecules (e.g., genomic dsDNA (e.g., in cells)), the method comprising: contacting the dsDNA molecule with (a) a fusion protein described herein (or a nucleic acid molecule encoding the fusion protein (e.g., DNA, RNA nucleic acid molecule)) and (b) a template RNA (e.g., a single template RNA, multiple different template RNAs) (which contain (e.g., from 5' to 3') crRNA, tracrRNA, a heterologous object sequence and a 3' target homologous domain) to modify the target sites in the dsDNA molecule (or a nucleic acid molecule encoding the template RNA (e.g., DNA nucleic acid molecule)) thereby editing the target sites in the dsDNA molecule (e.g., genomic dsDNA (e.g., in cells)).

[0055] In some embodiments, the nucleic acid molecule is in a cell (e.g., a eukaryotic cell). In some embodiments, the cell is in vitro, ex vivo, or in vivo. In some embodiments, the cell is in a subject (e.g., a human subject). In some embodiments, the cell is in a human subject. In some embodiments, the editing includes adding one or more nucleotides to the target site of the genomic dsDNA in the cell, deleting one or more nucleotides from the target site, or substituting one or more nucleotides in the target site. In some embodiments, the editing includes adding one or more nucleotides to the target site of the target nucleic acid molecule, deleting one or more nucleotides from the target site, or substituting one or more nucleotides in the target site. In some embodiments, the addition comprises adding approximately 1-500, 1-3200, 1-300, 1-200, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 nucleotides to the target site. In some embodiments, the deletion comprises deleting approximately 1-500, 1-3200, 1-300, 1-200, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 nucleotides to the target site.

[0056] In one aspect, this document provides methods for treating, improving, or preventing a disease in a subject (e.g., a human subject) in need, the methods comprising administering the Cas endonuclease described herein, the conjugate described herein, the fusion protein described herein, the system described herein, the nucleic acid molecule described herein, the vector described herein, the loading agent described herein, the reaction mixture described herein, the cell described herein, or the pharmaceutical composition described herein, thereby treating, improving, or preventing the disease in the subject.

[0057] In some embodiments, the disease is associated with a genetic defect. In some embodiments, the gRNA of the system is capable of targeting the endonuclease to the site of the genetic defect. In some embodiments, the genetic defect includes gene duplication, gene deletion, or gene mutation. In some embodiments, the administration results in correction of the genetic defect. In some embodiments, the subject is a human subject.

[0058] In one aspect, this article provides Cas endonucleases, conjugates, fusion proteins, systems, nucleic acid molecules, vectors, loading agents, reaction mixtures, cells, or pharmaceutical compositions for use at target sites in cleavage of target nucleic acid (e.g., DNA) molecules (e.g., double-stranded target nucleic acid sequences (e.g., dsDNA (e.g., genomic dsDNA))) in subjects in need.

[0059] In one aspect, this document provides the use of the Cas endonuclease, conjugate, fusion protein, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition described herein for the manufacture of a medicament for cleaving a target site in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))) in a subject of need.

[0060] In one aspect, this document provides the Cas endonuclease, conjugate, fusion protein, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition described herein for use at a target site in editing a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA)) in a subject in need.

[0061] In one aspect, this document provides the use of the Cas endonuclease, conjugate, fusion protein, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition described herein for the manufacture of a medicament for editing target sites in a target nucleic acid (e.g., DNA) molecule (e.g., double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))) in a subject in need.

[0062] In one aspect, this document provides the Cas endonuclease, conjugate, fusion protein, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition described herein for use as a drug.

[0063] In one aspect, this document provides the Cas endonuclease, conjugate, fusion protein, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition described herein for use in treating a disease (e.g., a disease related to a genetic defect) in a subject in need.

[0064] In one aspect, this document provides the use of the Cas endonuclease, conjugate, fusion protein, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition described herein for the manufacture of a medicament for treating a disease (e.g., a disease related to a genetic defect) in a subject in need. Detailed Implementation

[0065] Typical CRISPR-Cas editing (e.g., gene editing) systems require Cas endonucleases to mediate the cleavage of target nucleic acid molecules. The ability of Cas endonucleases to mediate target cleavage (e.g., in cells) varies depending on factors such as the efficiency of target cleavage, their ability to mediate double-strand and / or single-strand breaks, prototype spacer adjacent motif (PAM) sequence requirements, PAM specificity, etc. Therefore, a diverse set of Cas endonucleases can be used to provide the ability to select the appropriate Cas endonuclease for each specific target nucleic acid molecule; especially considering the incredible diversity of potential target nucleic acid molecules (e.g., the diversity of genes).

[0066] The inventors have in particular discovered novel Cas endonucleases. Therefore, the Cas endonucleases described herein can be used to modify (e.g., cleave) DNA, for example, in nucleic acid editing systems (e.g., CRISPR-Cas systems). Thus, this disclosure particularly provides Cas endonucleases capable of cleaving target nucleic acid molecules (e.g., DNA, genes, genomic DNA) (e.g., in cells, in cells of a subject); and systems and methods utilizing them (e.g., methods for cleaving nucleic acid molecules, methods for editing nucleic acid molecules (e.g., genomic DNA), and methods for treating diseases (e.g., genetic diseases)). Table of contents 4.1 Definition 4.2 Cas endonuclease 4.2.1 Activity of Cas endonuclease 4.2.1.1 Endonuclease activity 4.2.1.2 gRNA binding activity 4.2.1.3 Target nucleic acid molecule binding activity 4.2.1.4 Target nucleic acid editing activity 4.2.1.5 Changes in activity 4.3 Cas endonuclease fusion proteins and conjugates 4.3.1 Heterologous Proteins 4.3.1.1 Polymerase (e.g., reverse transcriptase (RT)) 4.3.1.2 Nucleobase Editor 4.3.2 Connector 4.3.3 Orientation 4.4 Methods for protein preparation 4.5 system 4.5.1 Target nucleic acid molecules 4.5.2gRNA 4.5.2.1 Multiple gRNAs 4.5.2.2 Modified gRNA 4.5.2.2(i) Properties of Modification 4.5.2.2(i)(a) Sugar modification 4.5.2.2(i)(b) Nucleobase modification 4.5.2.2(i)(c) Nucleoside interlinking modification Exemplary combinations modified by 4.5.2.2(i)(d) 4.5.2.2(ii) Position of Modification 4.5.2.3 Methods for preparing gRNA 4.5.3 Systemic nucleic acid editing activity 4.5.4 Methods for evaluating the nucleic acid editing activity of the system 4.5.5 Exemplary System 4.5.5.1 HDR-based editing system 4.5.5.2 RT-based editing system 4.5.5.3 Nucleotide Editor Editing System 4.6 Nucleic Acid Molecules 4.7 Carrier 4.8 Carrier 4.8.1 Lipid-based carriers 4.8.1.1 Cationic lipids (positively charged) and ionizable lipids 4.8.1.2 Non-cationic lipids (e.g., phospholipids) 4.8.1.3 Structural lipids 4.8.1.4 Polymers and polyethylene glycol (PEG)-lipids 4.8.1.5 Percentage of lipid nanoform formulation components 4.9 cells 4.10 Reaction mixture 4.11 Pharmaceutical Composition 4.12 Reagent Kit 4.13 How to use 4.13.1 Delivery Method 4.13.2 Methods for cleaving target nucleic acid molecules 4.13.3 Methods for editing target nucleic acid molecules 4.13.3.1 Methods for editing target nucleic acid molecules using RT-based systems 4.13.3.2 Methods for editing target nucleic acid molecules using HDR-based systems 4.13.3.3 Methods for editing target nucleic acid molecules using a nucleobase editor-based system 4.13.4 Methods for treating, improving, or preventing diseases 4.1 Definition

[0067] The chapter titles used in this article are for organizational purposes and do not limit the topics described.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by one of skill in the art to which the claimed subject matter pertains. It should be understood that the general and detailed descriptions are exemplary and interpretive, and do not limit the claimed subject matter.

[0069] In this application, unless otherwise stated, the use of the singular includes the plural. For example, unless the context otherwise requires, as used in this disclosure, the singular forms “a / an” and “the” include a plural number of indicators. Furthermore, the use of the term “including” and other forms such as “include,” “includes,” and “included” is not restrictive.

[0070] It should be understood that the aspects and embodiments described herein using the language of "comprising" also include similar aspects and embodiments described using the terms "consisting of" and "substantially consisting of".

[0071] The term “and / or” should be considered as a specific disclosure of each of two specified features or components, with or without the other. Thus, as used in phrases such as “A and / or B” herein, “and / or” is intended to include “A and B”, “A or B”, “A” (alone), and “B” (alone). Similarly, as used in phrases such as “A, B, and / or C”, “and / or” is intended to cover each of the following: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0072] As stated herein, unless otherwise indicated, concentration ranges, percentage ranges, ratio ranges, or integer ranges should be understood to include any integer values ​​within the range, and, where appropriate, to include fractions thereof (such as one-tenth and one-hundredth of an integer).

[0073] The term "about" refers to a value or composition within an acceptable margin of error for a particular value or composition, as understood and / or determined by a person skilled in the art, which will depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. When a particular value or composition is provided in this disclosure, unless otherwise stated, "about" should be understood to mean within an acceptable margin of error for that particular value or composition.

[0074] In the context of the description of proteins herein, it should be understood that polynucleotides (e.g., RNA or DNA nucleic acid molecules) encoding said proteins are also provided herein.

[0075] In the context of the description of proteins, nucleic acid molecules, carriers, and loading agents in this article, it should be understood that the article also provides the isolation forms of the proteins, nucleic acid molecules, carriers, and loading agents.

[0076] In the context of the description of proteins, nucleic acid molecules, etc., this article also provides recombinant forms of the proteins, nucleic acid molecules, etc.

[0077] In the context of describing proteins or proteomes, it should be understood that this article provides proteins that contain primary structures and proteins that fold into their three-dimensional structures (i.e., tertiary or quaternary structures).

[0078] As used herein, the term "application" means the physical introduction of an agent (e.g., a therapeutic agent (or a therapeutic agent precursor that is metabolized or altered in the body of a subject to produce a therapeutic agent in vivo)) (e.g., a system containing an endonuclease for introducing a mutation into a target nucleic acid) into a subject using any of the various methods and delivery systems known to those skilled in the art. Application may also be performed, for example, once, multiple times, and / or over one or more extended periods of time. Therapeutic agents include agents whose action is intended to be preventative (i.e., prophylactic), such as agents for modifying target nucleic acids (e.g., a system containing an endonuclease for introducing a mutation into a target nucleic acid).

[0079] As used herein, the term "bicyclic sugar" refers to a modified sugar moiety (e.g., ribose) comprising two rings, wherein the second ring is formed via a bridge connecting two atoms in the first ring, thereby forming a bicyclic structure. In some embodiments, the first ring of the bicyclic sugar moiety is a furanyl sugar moiety. In some embodiments, the furanyl sugar moiety is a ribosyl sugar moiety.

[0080] As used herein, the term “bicyclic nucleoside” (“BNA”) is a nucleoside that contains a bicyclic sugar.

[0081] As used herein, the term “crRNA” refers to an RNA molecule (e.g., a portion of gRNA (e.g., sgRNA)) capable of binding to the prototype spacer in a target nucleic acid (e.g., DNA) molecule.

[0082] As used herein, the term "disease" refers to an abnormal condition that impairs physiological function. The term encompasses any disorder, disease, abnormality, pathology, condition, symptom, or syndrome in which physiological function is impaired, regardless of its etiological nature. The term disease includes infections (e.g., viral, bacterial, fungal, protozoan infections).

[0083] As used herein, the term "donor template nucleic acid molecule" refers to a nucleic acid molecule containing a donor region and two homologous arms, wherein the donor region contains a target nucleic acid sequence (e.g., containing a target nucleotide variation (e.g., substitution, addition, deletion, inversion, etc.)) and each homologous arm contains a nucleotide sequence (also referred to herein as a homologous arm) of a region flanking the target cleavage site of an endonuclease described herein. Each homologous arm is flanking the donor region such that the donor region is between two homologous arms. In some embodiments, the donor template nucleic acid molecule is a donor DNA template nucleic acid molecule. In some embodiments, the donor template nucleic acid molecule is an RNA template molecule. In some embodiments, the donor template nucleic acid molecule is double-stranded. In some embodiments, the donor template nucleic acid molecule is single-stranded. In some embodiments, the donor template nucleic acid molecule can be used in systems described herein (e.g., HDR-based systems described herein), wherein the molecular mechanisms of the cell can utilize the exogenous donor template nucleic acid to repair and / or resolve cleavage sites in target nucleic acid molecules mediated by (e.g., systemic) endonucleases (or functional fragments, functional variants, or domains thereof).

[0084] The terms "DNA" and "polydeoxyribonucleotide" are used interchangeably and refer to a macromolecule comprising multiple deoxyribonucleotides polymerized via phosphodiester bonds. A deoxyribonucleotide is a nucleotide in which the sugar is deoxyribose.

[0085] As used herein, the term "domain" refers to a structure of a biomolecule (e.g., a protein, nucleic acid (e.g., DNA, RNA) molecule) that contributes to a specified function of the biomolecule (e.g., a protein, nucleic acid (e.g., DNA, RNA)). A domain may contain a continuous region (e.g., a continuous sequence) or distinct non-continuous regions (e.g., non-continuous sequences) of the biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, and reverse transcriptase domains; examples of nucleic acid domains are regulatory domains, such as transcription factor-binding domains. In some embodiments, a domain (e.g., a Cas domain) may contain two or more smaller domains (e.g., a DNA-binding domain and an endonuclease domain).

[0086] As used herein, the term “editing” relating to nucleic acid molecules (e.g., target nucleic acid (e.g., DNA)) and molecules (e.g., double-stranded target nucleic acid sequences (e.g., dsDNA (e.g., genomic dsDNA)) refers to the introduction of a variation (as defined herein) into a nucleic acid molecule (also referred to herein as edit). In some embodiments, variation or editing includes substitution, addition, deletion, or inversion.

[0087] As used herein, the term "edited strand" for double-stranded nucleic acid molecules (e.g., dsDNA molecules) refers to the strand of a double-stranded nucleic acid molecule that has been edited by, for example, endonucleases, systems, etc., as described herein. Similarly, as used herein, the term "unedited strand" for double-stranded nucleic acid molecules (e.g., dsDNA molecules) refers to the strand of a double-stranded nucleic acid molecule that has not been edited by, for example, endonucleases, systems, etc., as described herein.

[0088] As used herein, the term "functional fragment" in relation to a protein refers to a segment of a reference protein that retains at least one specific function. A functional fragment of a protein does not need to retain all the functions of the reference protein. In some cases, one or more functions are selectively reduced or eliminated. In some embodiments, the reference protein is a wild-type protein. For example, a functional fragment of a polymerase, reverse transcriptase, or endonuclease may refer to a segment of said protein that retains its activity. In some embodiments, a functional fragment comprises one or more domains (e.g., one, two, three, or more) of the reference protein.

[0089] As used herein, the term "functional variant" for a protein refers to a protein that contains at least one, but no more than 20%, 15%, 12%, 10%, or 8% amino acid variation (e.g., substitution, deletion, or addition) compared to the amino acid sequence of a reference protein, wherein the protein retains at least one specific function of the reference protein. A functional variant of a protein does not need to retain all the functions of the reference protein (e.g., wild-type). In some cases, one or more functions (e.g., endonuclease activity) are selectively altered, reduced, or eliminated. In some embodiments, the reference protein is a wild-type protein. In some embodiments, the functional variant comprises one or more domains (e.g., 1, 2, 3, or more) of the reference protein.

[0090] As used herein, the terms “functional fragments or variants thereof” and the like for pharmaceutical agents (e.g., proteins) should be understood to include functional variants, functional variations, functional fragments and variants.

[0091] As used herein, the term "fusion" and its grammatical equivalents refer to the operative linking of at least a first polypeptide with a second polypeptide, wherein the first and second polypeptides are not naturally found to be operatively linked together. For example, the first and second polypeptides are derived from different proteins and / or from different organisms. The term fusion encompasses both direct linking of at least two polypeptides via peptide bonds and indirect linking via adapters (e.g., peptide adapters).

[0092] As used herein, the term "fusion protein" and its grammatical equivalents refer to a protein comprising at least one polypeptide operatively linked to another polypeptide, wherein the first polypeptide and the second polypeptide are not naturally found to be operatively linked together. For example, the first polypeptide and the second polypeptide of the fusion protein are each derived from different proteins and / or from a heterologous organism. In some embodiments, the first polypeptide and the second polypeptide are distinct. For clarity, it should be understood that neither the first polypeptide nor the second polypeptide needs to be a full-length protein (e.g., a full-length, naturally occurring protein). For example, the first polypeptide and / or the second polypeptide may comprise or be composed of fragments (e.g., functional fragments or domains of a full-length protein (e.g., engineered, naturally occurring)). At least two polypeptides of a fusion protein may be operatively linked directly by peptide bonds; or operatively linked indirectly by a linker (e.g., a peptide linker). Thus, the term fusion polypeptide encompasses embodiments in which polypeptide A is operatively linked directly to polypeptide B by peptide bonds (polypeptide A-polypeptide B), and embodiments in which polypeptide A is operatively linked to polypeptide B by a peptide linker (polypeptide A-peptide linker-polypeptide B).

[0093] As used herein, the term "guide RNA" or "gRNA" refers to an RNA molecule that can associate with an endonuclease (e.g., the endonuclease described herein) to guide the endonuclease (e.g., the endonuclease described herein) to a target nucleic acid molecule (e.g., within a gene (e.g., within a cell)). gRNA requires both crRNA and tracrRNA. As described throughout, crRNA and tracrRNA can be the same larger RNA molecule (e.g., sgRNA) or part of a separate RNA molecule.

[0094] As used herein, when referring to a second element to describe a first element, the term "heterogeneous" means that the first and second elements do not exist in nature in the arrangement described. For example, a protein containing a "heterogeneous portion" refers to a protein linked to a portion (e.g., a small molecule, protein, polynucleotide, carbohydrate, lipid, synthetic polymer (e.g., PEG polymer)) that is not naturally linked to the protein.

[0095] As used herein, the term "heterologous target sequence" refers to an RNA molecule encoding a target nucleic acid (e.g., DNA) sequence (e.g., a gene) for a desired edit (e.g., substitution, addition, or deletion of one or more nucleotides), said target nucleic acid sequence being used as a template strand by a polymerase (e.g., reverse transcriptase) (e.g., described herein) to polymerize the desired nucleic acid sequence (e.g., DNA sequence (e.g., gene sequence)) (i.e., to polymerize a sequence complementary to the editing template). In some embodiments, the editing template is a portion of a template gRNA (e.g., described herein).

[0096] As is clear from this disclosure, but for clarity it should be understood that the use of the term "heterologous protein" (e.g., any heterologous protein described herein) includes both full-length proteins and proteins shorter than full length, including, for example, functional fragments, functional variants, and domains of full-length proteins.

[0097] As used herein, the term “isolated” in relation to biomolecules (e.g., proteins or polynucleotides) refers to biomolecules (e.g., proteins or polynucleotides) that are substantially free of other cellular components associated with them in their natural state.

[0098] As used herein, the term "translatable RNA" refers to any RNA that encodes at least one polypeptide and can be translated to produce the encoded protein in vitro, in vivo, in situ, or ex vivo. Transducible RNA can be mRNA or circular RNA encoding a polypeptide.

[0099] As used herein, the terms “pharmaceutical” and “part” are used interchangeably and refer to any macromolecule or micromolecule that can be operatively linked to another macromolecule or micromolecule (e.g., a protein (e.g., an endonuclease (or a functional fragment, functional variant, or domain thereof)) or a nucleic acid molecule encoding a protein (e.g., an endonuclease)). Exemplary parts include, but are not limited to, small molecules, proteins, polynucleotides (e.g., DNA, RNA), carbohydrates, lipids, and synthetic polymers (e.g., PEG polymers).

[0100] The terms “nucleic acid molecule” and “polynucleotide” are used interchangeably herein and refer to a polymer of DNA or RNA. Nucleic acid molecules can be single-stranded or double-stranded; contain natural, non-natural, or modified nucleotides; and contain natural, non-natural, or modified internucleotide bonds, including aminophosphate bonds or thiophosphate bonds, rather than phosphodiesters found between nucleotides in unmodified nucleic acid molecules. Nucleic acid molecules include, but are not limited to, all nucleic acid molecules obtained by any means available in the art, including but not limited to recombinant means, such as cloning nucleic acid molecules from recombinant libraries or cell genomes using common cloning techniques and polymerase chain reactions, as well as synthetic means. Those skilled in the art will understand that, unless otherwise stated, the nucleic acid sequences shown in this application will enumerate thymidine (T) in representative DNA sequences, but where the sequence represents RNA (e.g., mRNA), thymidine (T) will be replaced with uracil (U). Thus, any RNA polynucleotide encoded by DNA identified by a specific sequence identification number may also contain a corresponding RNA (e.g., mRNA) sequence encoded by said DNA, wherein each thymidine (T) sequence of said DNA is replaced with uracil (U).

[0101] As used herein, the term “nucleobase editor” refers to agents (e.g., biomolecules (e.g., proteins (or functional fragments, functional variants or domains thereof)) that can mediate nucleobase editing activity.

[0102] As used herein, the term "nucleobase editing activity" refers to the ability of an agent (e.g., a biomolecule (e.g., a protein (or a functional fragment, functional variant, or domain thereof))) to chemically alter the nucleobases within a polynucleotide. In some embodiments, nucleobase editing activity is cytidine deaminase activity, for example, converting a target C·G to T·A. In some embodiments, nucleobase editing activity is adenosine deaminase activity, for example, converting A·T to G·C. In some embodiments, nucleobase editing activity is both cytidine deaminase and adenosine deaminase activity, for example, converting A·T to G·C.

[0103] As used herein, the term "operably linked" refers to a bond between two parts in a functional relationship. For example, when peptides are linked (directly or indirectly via peptide linkers), the peptide is operably linked to another peptide such that both peptides are functional (e.g., in-frame fusion proteins containing endonucleases as described herein). Or, for example, transcriptionally regulatory polynucleotides, such as promoters, enhancers, or other expression control elements, are operably linked to polynucleotides encoding proteins to influence the transcription of the polynucleotide encoding said protein. The term "operably linked" also refers to a partial conjugation to, for example, a polynucleotide or peptide (e.g., the conjugation of a PEG polymer to a protein).

[0104] As used herein, the term “PAM” or “prototype spacer adjacent motif” refers to a short nucleic acid molecule (typically about 2-6 base pairs in length) following a nucleic acid region targeted for cleavage by an endonuclease (e.g., the endonuclease described herein, or the endonuclease of the system described herein). In some embodiments, the PAM is required for the endonuclease (e.g., the endonuclease described herein, or the endonuclease of the system described herein) to cleave the target nucleic acid molecule and is typically located near downstream of the cleavage site (e.g., 3-4 nucleotides).

[0105] As used herein, the determination of the “percentage of identity” between two sequences (e.g., proteins (amino acid sequences) or polynucleotides (nucleic acid sequences)) can be accomplished using mathematical algorithms. For example, specific, non-limiting examples of algorithms for comparing two sequences are described in Karlin S and Altschul SF (1990) PNAS [Proceedings of the National Academy of Sciences] 87: 2264-2268, as modified in Karlin S and Altschul SF (1993) PNAS [Proceedings of the National Academy of Sciences] 90: 5873-5877, each of which is incorporated herein by reference in its entirety. Such algorithms are incorporated in the NBLAST and XBLAST procedures of Altschul SF et al., (1990) J Mol Biol [Journal of Molecular Biology] 215: 403, which is incorporated herein by reference in its entirety. BLAST nucleotide searches were performed using the NBLAST nucleotide procedure parameter set (e.g., score = 100, word length = 12) to obtain nucleotide sequences homologous to the nucleic acid molecules described herein. BLAST protein searches were performed using the XBLAST procedure parameter set (e.g., score = 50, word length = 3) to obtain amino acid sequences homologous to the protein molecules described herein. For vacancy alignment comparisons, Gapped BLAST can be used as described in Altschul SF et al., (1997) Nuc Acids Res [Nucleic Acid Research] 25: 3389-3402, which is incorporated herein by reference in its entirety. Alternatively, PSI BLAST can be used for searches to detect long-distance relationships between molecules (as above). When using the BLAST, Gapped BLAST, and PSI Blast programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used (see, for example, the National Center for Biotechnology Information (NCBI) on the World Wide Web ncbi.nlm.nih.gov). Another specific, non-limiting example of a mathematical algorithm for comparing sequences is described in Myers and Miller, 1988, CABIOS [Computer Applications in the Biological Sciences] 4:11-17, which is incorporated herein by reference in its entirety. This algorithm is incorporated into the ALIGN program (version 2.0) and is part of the GCG sequence alignment software package. When comparing amino acid sequences using the ALIGN program, the PAM120 weighted residue table, vacancy length penalty 12, and vacancy penalty 4 can be used. The percentage of identity between two sequences, with or without vacancy, can be determined using techniques similar to those described above. When calculating the percentage of identity, typically only exact matches are calculated.

[0106] As used herein, the term “multiple” means two or more (e.g., three or more, four or more, five or more, six or more, seven or more, nine or more, or ten or more).

[0107] As used herein, the term "pharmaceutical composition" means a composition suitable for administration to animals (e.g., human subjects) and comprising a pharmaceutical agent (e.g., a therapeutic agent) and a pharmaceutically acceptable carrier or diluent. "Pharmaceutically acceptable carrier or diluent" means a substance intended for contact with human and / or non-human animal tissues without excessive toxicity, irritation, anaphylactic response, or other problems or complications, in proportion to a reasonable therapeutic benefit / risk ratio.

[0108] As used herein, “protein” and “polypeptide” refer to a polymer of at least two (e.g., at least five) amino acids linked by peptide bonds. The term “polypeptide” does not refer to a polymer chain of amino acids of a specific length. Shorter amino acid polymers (e.g., about 2-50 amino acids) are generally referred to as peptides in the art, and longer amino acid polymers (e.g., about 50 amino acids) are referred to as polypeptides. However, the terms “peptide” and “polypeptide” are used interchangeably with “protein” herein. In some embodiments, proteins fold into their three-dimensional structures. In the context of proteins as understood herein, it should be understood that this document provides proteins comprising a primary structure, and proteins folded into their three-dimensional structures (i.e., tertiary or quaternary structures).

[0109] As used herein, the term "preventive treatment" refers to treatment administered to subjects for the purpose of reducing the risk of developing pathology in subjects who do not exhibit signs of disease or only exhibit early signs of disease.

[0110] The terms “RNA” and “polynucleotide” are used interchangeably herein and refer to a macromolecule comprising multiple ribonucleotides polymerized via phosphodiester bonds. Ribonucleotides are nucleotides whose sugar is ribose. RNA may contain modified nucleotides; and contains natural, non-natural, or altered internucleotide bonds, such as aminophosphate bonds or thiophosphate bonds, rather than the phosphodiester bonds found between nucleotides in unmodified nucleic acid molecules.

[0111] As used herein, the term "sgRNA" refers to a gRNA molecule that contains both crRNA and tracrRNA. The components of sgRNA can be arranged in any suitable order, and any component can be operatively linked directly or indirectly (e.g., via nucleotide linkers) to one or more adjacent components.

[0112] As used herein, the term "signal peptide" or "signal sequence" refers to a sequence that can guide the transport or localization of proteins (such as endonucleases) to an organelle, cellular compartment, or extracellular export. The term encompasses both signal sequence peptides and nucleic acid sequences encoding signal peptides. Therefore, reference to a signal peptide in the context of nucleic acids refers to a nucleic acid sequence encoding a signal peptide. Exemplary signal sequences include, for example, nuclear localization signals and nuclear export signals.

[0113] As used herein, the term "subject" includes any animal, such as a human or other animal. In some embodiments, the subject is a vertebrate (e.g., a mammal, bird, fish, reptile, or amphibian). In some embodiments, the subject is a human. In some embodiments, the method subject is a non-human mammal. In some embodiments, the subject is a non-human mammal, such as a non-human primate (e.g., monkey, ape), an ungulate (e.g., cattle, buffalo, sheep, goat, pig, camel, llama, alpaca, deer, horse, donkey), a carnivore (e.g., dog, cat), a rodent (e.g., rat, mouse), or a rabbit (e.g., rabbit). In some embodiments, the subject is a bird, such as a member of the bird groups Galliformes (e.g., chicken, turkey, pheasant, quail), Anseriformes (e.g., duck, goose), Paleognathea (e.g., ostrich, emu), Columbiformes (e.g., pigeon, wild pigeon), or Psittaciformes (e.g., parrot).

[0114] As used herein, the term "template RNA" refers to a gRNA molecule comprising crRNA, tracrRNA, a heterologous target sequence, and a 3' target homologous domain. In some embodiments, the template RNA further comprises an RNA sequence that binds a polymerase (e.g., a reverse transcriptase, such as the reverse transcriptase of the fusion protein described herein). The components of the template RNA may be arranged in any suitable order, and any component may be operatively linked directly or indirectly (e.g., via a nucleotide linker) to one or more adjacent components. In some embodiments, the template RNA comprises crRNA, tracrRNA, a heterologous target sequence, and a 3' target homologous domain from 5' to 3'. In some embodiments, the template RNA comprises crRNA, tracrRNA, a sequence that binds a polymerase (e.g., a reverse transcriptase, such as the reverse transcriptase of the fusion protein described herein), a heterologous target sequence, and a 3' target homologous domain from 5' to 3'. In some embodiments, the template RNA is part of a system described herein (e.g., a reverse transcriptase-based system).

[0115] As used herein, the term "therapeuticly effective amount" of a pharmaceutical agent (e.g., a therapeutic agent) means any amount of an agent (e.g., a therapeutic agent) that, when used alone or in combination with another therapeutic agent, improves the condition of a disease, for example, by protecting a subject from the onset of a disease (or infection); improves the symptoms of a disease or infection, for example, by reducing the severity of symptoms, the frequency or duration of symptoms, or increasing the period of no disease or no symptoms; prevents or reduces damage or disability caused by a disease or infection; or promotes the resolution of a disease (or infection). The ability of a therapeutic agent to improve the condition of a disease can be evaluated using a variety of methods known to a skilled technician, such as in human subjects during a clinical trial, in animal model systems where efficacy in humans can be predicted, or by measuring the activity of the agent in an in vitro assay.

[0116] As used herein, the term “tracrRNA” refers to an RNA molecule (e.g., a portion of gRNA, such as sgRNA) that mediates the binding of gRNA to an endonuclease (e.g., the endonuclease described herein).

[0117] As used herein, the terms “treat,” “treating,” “treatment,” etc., refer to reducing or improving a disease and / or one or more symptoms associated with it, or achieving a desired pharmacological and / or physiological effect. It should be understood that, although not excluded, treating a disease does not require the complete elimination of the disease or one or more symptoms associated with it. In some embodiments, the effect is therapeutic, i.e., but not limited to, the effect partially or completely reduces, weakens, eliminates, slows, alleviates, lessens, lowers the intensity of the disease and / or unpleasant symptoms attributable to the disease, or cures the disease and / or unpleasant symptoms attributable to the disease. In some embodiments, the effect is preventative, i.e., the effect protects against or prevents the occurrence or recurrence of the disease. For this purpose, the method disclosed in this invention includes administering a therapeutically effective amount of the composition as described herein.

[0118] As used herein, a “variant” or “mutation” of a nucleic acid molecule (e.g., a nucleic acid molecule encoding an endonuclease as described herein) means a nucleic acid molecule that contains at least one nucleotide substitution, inversion, addition, or deletion compared to a reference nucleic acid molecule. As used herein, the term “variant” or “mutation” of a protein means a peptide or protein (e.g., an endonuclease as described herein) that contains at least one amino acid residue substitution, inversion, addition, or deletion compared to a reference protein.

[0119] As used herein, the term "3' target homologous domain" refers to an RNA molecule capable of hybridizing with the 3' end (3' target sequence) of a single-stranded nucleic acid lobe resulting from an induced single-strand break (i.e., nick) in a target double-stranded nucleic acid (e.g., DNA) molecule (e.g., by an endonuclease (or a fusion protein containing it) as described herein). Hybridization of the 3' target homologous domain with the 3' target sequence produces a double-stranded structure that can be used as a substrate for the polymerization of nucleic acid (e.g., DNA) molecules (e.g., using a heterologous object sequence) by a polymerase (e.g., a reverse transcriptase) (e.g., as described herein). In some embodiments, the 3' target homologous domain is part of a template RNA (e.g., as described herein). 4.2 Cas endonuclease

[0120] This document provides, in particular, Cas endonucleases (and their functional fragments, functional variants, and domains) that can be used to modify (e.g., edit) nucleic acid molecules (e.g., DNA, genes, genomes, e.g., intracellularly, e.g., within the cells of a subject (e.g., mammalian subjects, e.g., human subjects)) (e.g., in vivo, in vitro, or ex vivo). In some embodiments, the Cas endonucleases are not naturally occurring. The amino acid sequences of exemplary Cas endonucleases disclosed herein are shown in Table 1 and SEQ ID NO: 1-320. Table 1. Amino acid sequences of Cas endonucleases.

[0121] In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain) comprises, or is composed of, an amino acid sequence having at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of any polypeptide shown in Table 1 or any of SEQ ID NO: 1-320. In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, an amino acid sequence having at least about 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with, the amino acid sequence of any polypeptide shown in Table 1 or any of SEQ ID NO: 1-320.

[0122] In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain) comprises, or is composed of, an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of the polypeptides shown in Table 1. In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain) comprises, or is composed of, an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of the polypeptide shown in Table 1.

[0123] In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 1, and further comprises one or more but less than 20% (e.g., less than 15%, less than 12%, less than 10%, less than 8%) amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 1, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 1, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 1, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 1, and further comprises, or is composed of, about 1-200, 1-150, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-5, 10-200, 10-150, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 1-040, 10-30, 10-20, 50-200, 50-150, 50-100, 50-90, 50-80, 50-70, or 50-60 amino acid variations (e.g., substitutions, additions, deletions, etc.).

[0124] In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 1, and further comprises one or more, but less than 20% (e.g., less than 15%, less than 12%, less than 10%, less than 8%) amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 1, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 1, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 1, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 1, and further comprises, or is composed of, about 1-200, 1-150, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-5, 10-200, 10-150, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 1-040, 10-30, 10-20, 50-200, 50-150, 50-100, 50-90, 50-80, 50-70, or 50-60 amino acid substitutions.

[0125] In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain) comprises, or is composed of, an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of any one of SEQ ID NO: 1-320. In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain) comprises, or is composed of, an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of any one of SEQ ID NO: 1-320. In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain) comprises, or is composed of, an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of any one of SEQ ID NO: 1-320.

[0126] In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises one or more but less than 20% (e.g., less than 15%, less than 12%, less than 10%, less than 8%) amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises or consists of any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises or consists of about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises or is composed of any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises or is composed of no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises or is composed of any one of the amino acid sequences in SEQ ID NO: 1-320. The amino acid sequence of any one of 1-320 or composed thereof, and further comprising or composed of about 1-200, 1-150, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-5, 10-200, 10-150, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 1-040, 10-30, 10-20, 50-200, 50-150, 50-100, 50-90, 50-80, 50-70 or 50-60 amino acid variations (e.g., substitution, addition, deletion, etc.).

[0127] In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or its functional fragment, functional variant, or domain thereof) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, or 200 amino acid substitutions. In some embodiments, the amino acid sequence of the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 1-320, and further comprises, or is composed of, about 1-200, 1-150, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-5, 10-200, 10-150, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 1-040, 10-30, 10-20, 50-200, 50-150, 50-100, 50-90, 50-80, 50-70, or 50-60 amino acid substitutions.

[0128] In some embodiments, the amino acid sequence of the Cas endonuclease has less than about 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, or 50% identity with the amino acid sequence of a reference Cas endonuclease (e.g., a naturally occurring Cas endonuclease). In some embodiments, the amino acid sequence of the Cas endonuclease has less than 90% (e.g., less than 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%) and greater than 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89% identity with the amino acid sequence of a reference Cas endonuclease (e.g., a reference naturally occurring Cas endonuclease). In some embodiments, the amino acid sequence of the Cas endonuclease has less than about 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, or 50% identity with the amino acid sequence of the reference Cas9 endonuclease.In some embodiments, the amino acid sequence of the Cas endonuclease has less than 90% (e.g., less than 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%) and greater than 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%) identity with the amino acid sequence of the reference Cas9 endonuclease. In some embodiments, the amino acid sequence of the Cas endonuclease has less than about 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, or 50% identity with the amino acid sequence of a reference Cas9 endonuclease containing the amino acid sequence shown in SEQ ID NO: 321. In some embodiments, the amino acid sequence of the Cas endonuclease has less than 90% (e.g., less than 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%) and greater than 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%) identity with the amino acid sequence of a reference Cas9 endonuclease containing the amino acid sequence shown in SEQ ID NO: 321. 4.2.1 Activity of Cas endonuclease

[0129] The Cas endonucleases described herein may have multiple functions, including functionally distinct domains. In some embodiments, the Cas endonucleases exhibit (or are engineered to exhibit) more than one (e.g., two, three, four, five, or more) different functions (e.g., as described herein). In some embodiments, the Cas endonucleases do not exhibit (or are engineered not to exhibit) one or more (e.g., two, three, four, five, or more) different functions (e.g., as described herein). Exemplary functions include, but are not limited to, endonuclease activity (e.g., introducing double-stranded and / or single-stranded breaks in a nucleic acid sequence), RNA (e.g., gRNA) binding activity, target nucleic acid (e.g., DNA) molecule binding activity, and target nucleic acid molecule editing activity (e.g., when provided as part of a suitable system (e.g., the system described herein)). 4.2.1.1 Endonuclease activity

[0130] In some embodiments, the Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (or a conjugate or fusion protein comprising any of the foregoing) comprises any one or more of the following properties (e.g., 1, 2, 3, 4, 5, 6, or more) (or is engineered to have one or more of the following properties): (a) DNA endonuclease activity; (b) RNA endonuclease activity; (c) DNA / RNA hybrid endonuclease activity; (d) RNA-directed DNA endonuclease activity; (e) DNA-directed DNA endonuclease activity; (f) RNA-directed RNA endonuclease activity; (g) DNA-directed RNA endonuclease activity; (h) the ability to mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule; (i) the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule; (j) the inability to mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule; and / or (k) The ability to mediate single-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules, and the inability to mediate double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules (i.e., nicking enzyme activity).

[0131] In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (or a conjugate or fusion protein comprising any of the foregoing) exhibits (or is engineered to exhibit) the ability to mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule. In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (or a conjugate or fusion protein comprising any of the foregoing) exhibits (or is engineered to exhibit) the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule.

[0132] In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (or a conjugate or fusion protein comprising any of the foregoing) exhibits (or is engineered to exhibit) the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule and cannot mediate double-strand breaks in the target double-stranded nucleic acid (e.g., DNA) molecule (i.e., nicking enzyme activity). In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (or a conjugate or fusion protein comprising any of the foregoing) is capable of (or is engineered to be capable of) mediating single-strand breaks at a higher frequency than double-strand breaks in the target double-stranded nucleic acid (e.g., DNA) molecule. In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (or a conjugate or fusion protein comprising any of the foregoing) is capable of (or engineered to be capable of) mediating single-strand breaks at a higher frequency than double-strand breaks in the target double-stranded nucleic acid (e.g., DNA) molecule (e.g., at least 90%, 95%, 96%, 97%, 98%, or 99% of breaks in the target double-stranded nucleic acid (e.g., DNA) molecule are single-strand breaks; or less than 10%, 5%, 4%, 3%, 2%, or 1% of breaks in the target double-stranded nucleic acid (e.g., DNA) molecule are double-strand breaks). In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (or a conjugate or fusion protein comprising any of the foregoing) substantially does not mediate (or is engineered to substantially not mediate) double-strand breaks in the target double-stranded nucleic acid (e.g., DNA) molecule. In some embodiments, Cas endonucleases (or functional fragments, functional variants, or domains thereof) (or conjugates or fusion proteins containing any of the foregoing) do not mediate (or are engineered to not mediate) detectable double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules. 4.2.1.2 gRNA binding activity

[0133] In some embodiments, the Cas endonuclease includes a nucleic acid molecule binding domain. In some embodiments, the Cas endonuclease includes a DNA binding domain. In some embodiments, the Cas endonuclease includes an RNA binding domain. In some embodiments, the Cas endonuclease includes a gRNA binding domain. In some embodiments, the Cas endonuclease is capable of binding the gRNA described herein. In some embodiments, the endonuclease is capable of binding crRNA. In some embodiments, the Cas endonuclease is capable of binding crRNA, which is part of template RNA or sgRNA. It is not intended to be theoretically construed that the binding of the Cas endonuclease to crRNA (e.g., the crRNA of template RNA or sgRNA) promotes the Cas endonuclease's targeting of target nucleic acid molecules (through coordination with tracrRNA (e.g., the tracrRNA of template RNA or sgRNA)). 4.2.1.3 Target nucleic acid molecule binding activity

[0134] In some embodiments, the Cas endonuclease includes a domain capable of binding to a target nucleic acid molecule (e.g., a target double-stranded nucleic acid molecule (e.g., a target dsDNA molecule)). In some embodiments, the Cas endonuclease recognizes a PAM in the target nucleic acid molecule (e.g., a target double-stranded nucleic acid molecule (e.g., a target dsDNA molecule)). In some embodiments, the Cas endonuclease requires the PAM to be present in or adjacent to a target site in the target nucleic acid molecule (e.g., a target double-stranded nucleic acid molecule (e.g., a target dsDNA molecule)) to mediate cleavage of the nucleic acid molecule. In some embodiments, the PAM sequence comprises or is composed of NGG. 4.2.1.4 Target nucleic acid editing activity

[0135] In some embodiments, when provided within a suitable system (e.g., the system described herein (see, for example, §4.5)), the Cas endonuclease can mediate the editing (e.g., addition, deletion, substitution, etc.) of the nucleotide sequence of a target nucleic acid molecule. In some embodiments, the Cas endonuclease exhibits increased editing efficiency relative to a reference Cas endonuclease (e.g., when provided in a suitable system (e.g., the system described herein)). In some embodiments, the Cas endonuclease exhibits an increase in editing efficiency of at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200% or more relative to a reference Cas endonuclease (e.g., when provided in a suitable system (e.g., the system described herein)). In some embodiments, the Cas endonuclease exhibits an increase in editing efficiency of at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100% or more relative to the editing efficiency of a reference Cas endonuclease (e.g., when provided in a suitable system (e.g., the system described herein)). In some embodiments, the Cas endonuclease exhibits an increase in editing efficiency of approximately 30%-200%, 40%-200%, 50%-200%, 60%-200%, 70%-200%, 80%-200%, 90%-200%, 100%-200%, 150%-200%, 30%-150%, 40%-150%, 50%-150%, 60%-150%, 70%-150%, 80%-150%, 90%-150%, 100%-150%, 30%-100%, 40%-100%, 50%-100%, 60%-100%, 70%-100%, 80%-100%, or 90%-100% or more (e.g., when provided in a suitable system, such as the system described herein).

[0136] In some embodiments, the Cas endonuclease exhibits increased editing efficiency relative to the reference Cas endonuclease shown in SEQ ID NO: 321 (e.g., when provided in a suitable system, such as the system described herein). In some embodiments, the Cas endonuclease exhibits increased editing efficiency by at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200% or more relative to the reference Cas endonuclease shown in SEQ ID NO: 321 (e.g., when provided in a suitable system, such as the system described herein). In some embodiments, the Cas endonuclease exhibits an increase in editing efficiency of at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100% or more (e.g., when provided in a suitable system (e.g., the system described herein)) relative to the editing efficiency of the reference Cas endonuclease shown in SEQ ID NO: 321. In some embodiments, relative to SEQ ID NO: 321 The editing efficiency of the reference Cas endonuclease shown in NO:321 is such that the Cas endonuclease exhibits an increase in editing efficiency of approximately 30%-200%, 40%-200%, 50%-200%, 60%-200%, 70%-200%, 80%-200%, 90%-200%, 100%-200%, 150%-200%, 30%-150%, 40%-150%, 50%-150%, 60%-150%, 70%-150%, 80%-150%, 90%-150%, 100%-150%, 30%-100%, 40%-100%, 50%-100%, 60%-100%, 70%-100%, 80%-100%, or 90%-100% or more (e.g., when provided in a suitable system (e.g., the system described herein)). 4.2.1.5 Changes in activity

[0137] In some embodiments, the amino acid sequence of a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of any Cas endonuclease shown in Table 1 or any of SEQ ID NO: 1-320, and further comprises one or more amino acid variations (e.g., substitution, deletion, addition) wherein the one or more amino acid variations (e.g., substitution, deletion, addition) alter the activity of the Cas endonuclease (e.g., the activities described herein (e.g., induction of double-strand breaks, nicking enzyme activity, gRNA binding activity, target nucleic acid binding activity, PAM recognition, etc.)). In some embodiments, the amino acid sequence of a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of any Cas endonuclease shown in Table 1 or any of SEQ ID NO: 1-320, and further comprises one or more amino acid variations (e.g., substitution, deletion, addition) but not exceeding 20%, 15%, 12%, 10%, or 8% of amino acid variations (e.g., substitution, deletion, addition), wherein the one or more amino acid variations (e.g., substitution, deletion, addition) alter the activity of the Cas endonuclease (e.g., the activities described herein (e.g., induction of double-strand breaks, nicking enzyme activity, gRNA binding activity, target nucleic acid binding activity, PAM recognition, etc.)).

[0138] In some embodiments, the one or more amino acid variations (e.g., substitution, deletion, addition) reduce or eliminate the ability of the Cas endonuclease to mediate double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules. In some embodiments, a Cas endonuclease containing the one or more amino acid variations (e.g., substitution, deletion, addition) has the ability to mediate single-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules, but does not have the ability to mediate double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules. In some embodiments, the one or more amino acid variations (e.g., substitution, deletion, addition) alter the PAM nucleotide sequence recognized by the Cas endonuclease. In some embodiments, the one or more amino acid variations (e.g., substitution, deletion, addition) reduce the Cas endonuclease activity of the Cas endonuclease by at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% relative to an Cas endonuclease lacking the one or more amino acid variations (e.g., substitution, deletion, addition). In some embodiments, the one or more amino acid variations (e.g., substitution, deletion, addition) increase the Cas endonuclease activity of the Cas endonuclease by at least 1, 2, 5, 10, or 100 times relative to an Cas endonuclease lacking the one or more amino acid variations (e.g., substitution, deletion, addition). 4.3 Cas endonuclease fusion proteins and conjugates

[0139] In some embodiments, the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (or a nucleic acid molecule encoding the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof)) is operatively linked to a heterologous moiety (e.g., a heterologous protein (e.g., a functional fragment, functional variant, or domain thereof)). Therefore, this document further provides, in particular, fusion proteins comprising a Cas endonuclease (e.g., described herein) (or a functional fragment, functional variant, or domain thereof) and one or more heterologous proteins (or functional fragments, functional variants, or domains thereof). This document further provides, in particular, conjugates comprising a Cas endonuclease (e.g., described herein) (or a functional fragment, functional variant, or domain thereof) (or a nucleic acid molecule encoding the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof)) and one or more heterologous moieties.

[0140] Heterogeneous components include, but are not limited to, proteins, peptides, small molecules, nucleic acid molecules (e.g., DNA, RNA, DNA / RNA hybrid molecules), carbohydrates, lipids, and polymers (e.g., synthetic polymers).

[0141] In some embodiments, an endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively linked to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more heterologous portions. In some embodiments, an endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively linked to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, but not more than 10, heterologous portions. In some embodiments, an endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively linked to no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 heterologous portions. In some embodiments, an endonuclease (or a functional fragment, functional variant thereof) is operatively linked to about 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 heterologous portions. In some embodiments, an endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively linked to about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 heterologous portions. 4.3.1 Heterologous Proteins

[0142] In some embodiments, the heterologous portion is a protein. Therefore, as described above, this document provides fusion proteins comprising a Cas endonuclease (e.g., as described herein) (or a functional fragment, functional variant, or domain thereof) and one or more heterologous proteins. It is clear from this disclosure, but for clarity it should be understood that the use of the term "heterologous protein" (e.g., any heterologous protein described herein) includes full-length proteins as well as functional fragments, functional variants, and domains of, for example, full-length proteins.

[0143] In some embodiments, the fusion protein comprises more than one heterologous protein. In some embodiments, the fusion protein comprises multiple heterologous proteins. In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively ligated to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more heterologous proteins. In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively ligated to at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, but not more than 10, heterologous proteins. In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively ligated to no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 heterologous proteins. In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively ligated to about 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 heterologous proteins (or functional fragments, functional variants, or domains thereof). In some embodiments, a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is operatively ligated to about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 heterologous proteins.

[0144] Exemplary heterologous proteins include, but are not limited to, cell localization signals (e.g., nuclear localization signal peptides, nuclear export signal peptides); detectable proteins (e.g., fluorescent proteins, protein tags (e.g., FLAG tags, HIS tags, HA tags), reporter genes); and enzymes. In some embodiments, the heterologous protein is an enzyme. In some embodiments, the heterologous protein exhibits enzymatic activity.

[0145] In some embodiments, the heterologous protein exhibits one or more of the following: polymerase activity (e.g., reverse transcriptase activity), nucleobase editing activity (e.g., deaminase activity), enzyme activity, epigenetic modification activity, nucleic acid cleavage activity, nucleic acid binding activity, transcriptional regulation activity, methyltransferase activity, demethylase activity (e.g., histone demethylase activity), acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylation activity, deSUMOylation activity, ribosylation activity, deribosylation activity, myristylation activity, demyristylation activity, integrase activity, transposase activity, recombinase activity, ligase activity, helicase activity, or nuclease activity.

[0146] In some embodiments, the heterologous protein exhibits polymerase (e.g., reverse transcriptase) activity, nucleobase modification activity (e.g., deaminase activity), methylase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, or double-stranded DNA cleavage activity and nucleic acid binding activity, or any combination thereof.

[0147] In some embodiments, the heterologous protein is a polymerase (e.g., reverse transcriptase), nucleobase editor (e.g., deaminase), methyltransferase, demethylase (e.g., histone demethylase), acetyltransferase, deacetylase, kinase, phosphatase, ubiquitin ligase, deubiquitinase, adenylate, deadenylate, SUMOylase, deSUMOylase, ribosylase, deribosylase, myristylase, demyristylase, integrase, transposase, recombinase, ligase, helicase, or nuclease, or a functional fragment, functional variant, or domain of any of the foregoing. 4.3.1.1 Polymerase (e.g., reverse transcriptase (RT))

[0148] In some embodiments, the heterologous protein exhibits polymerase (e.g., reverse transcriptase) activity. In some embodiments, the heterologous protein exhibits RNA-dependent DNA polymerase activity. In some embodiments, the heterologous protein exhibits reverse transcriptase activity.

[0149] In some embodiments, the heterologous protein is a polymerase (or a functional fragment, functional variant, or domain thereof). In some embodiments, the polymerase comprises, or is composed of, a catalytic (e.g., polymerase (e.g., reverse transcriptase)) domain of the polymerase (e.g., the polymerase described herein, e.g., reverse transcriptase (RT)). In some embodiments, the polymerase comprises, or is composed of, a catalytic (e.g., polymerase (e.g., reverse transcriptase)) domain of the polymerase (e.g., the polymerase described herein, e.g., reverse transcriptase (RT)) and a nucleic acid (e.g., RNA, DNA) binding domain of the polymerase. In some embodiments, the polymerase comprises, or is composed of, a catalytic (e.g., polymerase (e.g., reverse transcriptase)) domain of the RT (e.g., the polymerase described herein, e.g., reverse transcriptase (RT)). In some embodiments, the polymerase comprises, or is composed of, a catalytic (e.g., polymerase (e.g., reverse transcriptase (RT)) domain of the RT (e.g., the polymerase (e.g., reverse transcriptase (RT)) and an RNA binding domain of the RT.

[0150] In some embodiments, the polymerase comprises an RNase H domain of RT (e.g., RT as described herein). In some embodiments, the polymerase does not contain an RNase H domain of RT (e.g., RT as described herein). In some embodiments, the polymerase comprises a DNA-dependent DNA polymerase domain of RT (e.g., RT as described herein). In some embodiments, the polymerase does not contain a DNA-dependent DNA polymerase domain of RT (e.g., RT as described herein). In some embodiments, the DNA-dependent DNA polymerase domain is the same domain as the reverse transcriptase domain (i.e., the domain has both reverse transcriptase activity and DNA-dependent DNA polymerase activity). In some embodiments, the DNA-dependent DNA polymerase domain is not the same domain as the reverse transcriptase domain.

[0151] In some embodiments, the polymerase comprises, or is composed of, a reverse transcriptase domain of RT (e.g., as described herein), an RNA-binding domain of RT, and an RNase H domain of RT. In some embodiments, the polymerase comprises, or is composed of, a reverse transcriptase domain of RT (e.g., as described herein), and an RNA-binding domain of RT, and does not contain the RNase H domain of RT. In some embodiments, the polymerase comprises, or is composed of, a reverse transcriptase domain of RT (e.g., as described herein), an RNA-binding domain of RT, an RNase H domain of RT, and a DNA-dependent DNA polymerase domain of RT. In some embodiments, the polymerase comprises, or is composed of, a reverse transcriptase domain of RT (e.g., as described herein), an RNA-binding domain of RT, and an RNase H domain of RT, and does not contain the DNA-dependent DNA polymerase domain of RT.

[0152] In some embodiments, the polymerase is RT (or a functional fragment, functional variant, or domain thereof). In some embodiments, RT comprises, or is composed of, the reverse transcriptase domain of RT (e.g., as described herein). In some embodiments, RT comprises the RNA-binding domain of RT. In some embodiments, RT comprises, or is composed of, the RNase domain of RT (e.g., as described herein). In some embodiments, RT does not contain the RNase domain of RT (e.g., as described herein). In some embodiments, RT comprises, or is composed of, the DNA-dependent DNA polymerase domain of RT (e.g., as described herein). In some embodiments, RT does not contain the DNA-dependent DNA polymerase domain of RT (e.g., as described herein). In some embodiments, the DNA-dependent DNA polymerase domain is the same domain as the reverse transcriptase domain (i.e., the domain has both reverse transcriptase activity and DNA-dependent DNA polymerase activity). In some embodiments, the DNA-dependent DNA polymerase domain is not the same domain as the reverse transcriptase domain.

[0153] In some embodiments, RT comprises, or is composed of, a reverse transcriptase domain of RT (e.g., as described herein) and an RNA-binding domain of RT. In some embodiments, RT comprises, a reverse transcriptase domain of RT (e.g., as described herein), an RNA-binding domain of RT, and an RNase domain of RT. In some embodiments, RT comprises, a reverse transcriptase domain of RT (e.g., as described herein), and an RNA-binding domain of RT, and does not contain an RNase domain of RT. In some embodiments, RT comprises, a reverse transcriptase domain of RT (e.g., as described herein), an RNA-binding domain of RT, an RNase domain of RT, and a DNA-dependent DNA polymerase domain of RT. In some embodiments, RT comprises, a reverse transcriptase domain of RT (e.g., as described herein), an RNA-binding domain of RT, and an RNase domain of RT, and does not contain a DNA-dependent DNA polymerase domain of RT. In some embodiments, RT comprises, a reverse transcriptase domain of RT (e.g., as described herein), and an RNA-binding domain of RT, and does not contain an RNase domain of RT or a DNA-dependent DNA polymerase domain of RT.

[0154] Any of the aforementioned domains (e.g., reverse transcriptase domain, RNA-binding domain, RNase domain, DNA-dependent DNA polymerase domain) can be derived from the same or different polymerases (e.g., reverse transcriptase). Any of the aforementioned domains (e.g., reverse transcriptase domain, RNA-binding domain, RNase domain, DNA-dependent DNA polymerase domain) can be derived from a naturally occurring reverse polymerase (e.g., reverse transcriptase) or differ from a naturally occurring polymerase (e.g., reverse transcriptase) (e.g., as defined herein) (e.g., containing one or more amino acid variations). In some embodiments, the RT comprises domains from more than one RT.

[0155] In some embodiments, the RT (or a functional fragment, functional variant, or domain (e.g., a reverse transcriptase domain)) contains a region that specifically recognizes substrate RNA. For example, in some embodiments, the RT (or a functional fragment, functional variant, or domain (e.g., a reverse transcriptase domain)) contains a UTR (e.g., a 3' UTR) that specifically recognizes substrate RNA (e.g., a 3' UTR from a retrotransposon (e.g., a 3' UTR from a non-LTR retrotransposon (e.g., an RLE type, e.g., an R2 retrotransposon)). See, for example, Luan and Eickbush, Mol Cell Biol [Molecular and Cell Biology] 15, 3882-91 (1995)), the entire contents of which are incorporated herein by reference for all purposes. Exemplary 3' UTRs from retrotransposons are described in WO 2021178720 (see, for example, Table 3), the entire contents of which are incorporated herein by reference for all purposes. In some embodiments, RT is a dimer (e.g., homodimer, heterodimer). In some embodiments, RT is a monomer.

[0156] In some embodiments, RT comprises the full-length RT, or is composed of it. In some embodiments, RT comprises functional segments of RT, or is composed of it. In some embodiments, RT comprises functional variations of RT, or is composed of it. In some embodiments, RT comprises both functional segments and functional variations of RT, or is composed of it. In some embodiments, RT comprises one or more domains of RT, or is composed of it. In some embodiments, RT comprises functional segments of one or more domains of RT, or is composed of it. In some embodiments, RT comprises functional variations of one or more domains of RT, or is composed of it. In some embodiments, RT comprises both functional segments and functional variations of one or more domains of RT, or is composed of it.

[0157] In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is a naturally occurring RT. In some embodiments, RT comprises, or is composed of, a functional fragment of a naturally occurring RT. In some embodiments, RT comprises, or is composed of, a functional variant of a naturally occurring RT. In some embodiments, RT comprises, or is composed of, both functional fragments and functional variants of a naturally occurring RT. In some embodiments, RT comprises, or is composed of, one or more domains of a naturally occurring RT. In some embodiments, RT comprises, or is composed of, functional fragments of one or more domains of a naturally occurring RT. In some embodiments, RT comprises, or is composed of, functional variants of one or more domains of a naturally occurring RT. In some embodiments, RT comprises, or is composed of, functional fragments and functional variants of one or more domains of a naturally occurring RT.

[0158] In some embodiments, the RT (or its functional fragment, functional variant, or domain) comprises the amino acid sequence of a naturally occurring RT. In some embodiments, the RT (or its functional fragment, functional variant, or domain) comprises an amino acid sequence that includes at least one amino acid variation relative to the amino acid sequence of a naturally occurring RT. In some embodiments, the amino acid sequence of the RT (or its functional fragment, functional variant, or domain) comprises, or is composed of, an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with, or is composed of, the amino acid sequence of a naturally occurring RT. In some embodiments, the amino acid sequence of the RT (or its functional fragment, functional variant, or domain) comprises, or is composed of, the amino acid sequence of a naturally occurring RT, and further comprises one or more, but less than 15% (e.g., less than 12%, less than 10%, less than 8%), amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of RT (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of naturally occurring RT, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of RT (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of naturally occurring RT, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of RT (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of naturally occurring RT, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of RT (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of naturally occurring RT, and further comprises one or more, but less than 15% (less than 12%, less than 10%, less than 8%) amino acid substitutions. In some embodiments, the amino acid sequence of RT (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of naturally occurring RT, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions.In some embodiments, the amino acid sequence of RT (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of naturally occurring RT, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of RT (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of naturally occurring RT, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions.

[0159] In some embodiments, the amino acid sequence of the RT (or a functional fragment, functional variant, or domain thereof) comprises one or more amino acid variations (e.g., relative to the amino acid sequence of the naturally occurring RT) that provide one or more improved properties (e.g., relative to the amino acid sequence of the naturally occurring RT), including, for example, lower error rate, thermostability, increased sustained synthetic capacity, increased resistance to inhibitors, increased reverse transcriptase rate, increased resistance to modified nucleotides, mediation of the addition of modified DNA nucleotides, proofreading ability, DNA-dependent DNA polymerase activity, or any combination thereof. See, for example, WO 2001068895 and WO2018089860, the entire contents of which are incorporated herein by reference for all purposes.

[0160] Naturally occurring reverse transcriptases (RTs) are known in the art and described herein (see, for example, Table 2). Naturally occurring RTs include, for example, but not limited to, viral (e.g., retroviral) reverse transcriptases, non-LTR retrotransposon reverse transcriptases (e.g., APE type, RLE type), LTR retrotransposon reverse transcriptases, group II intron reverse transcriptases, reverse transcriptases that produce diverse reverse transcription elements, retrotranscriptases, telomerases, and retroplasmid reverse transcriptases. In some embodiments, the RT (or a functional fragment, functional variant, or domain thereof) is a eukaryotic RT or a prokaryotic RT. In some embodiments, the RT (or a functional fragment, functional variant, or domain thereof) is a viral RT or a bacterial RT.

[0161] In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is a retroviral RT. In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is an oncogenic retroviral RT or a foam virus RT. In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is an alpha retroviral RT, beta retroviral RT, delta retroviral RT, ε retroviral RT, gamma retroviral RT, lentivirus RT, bovispumavirus RT, equispumavirus RT, felispumavirus RT, prosimiispumavirus RT, or simiispumavirus RT. In some embodiments, the RT (or a functional fragment, functional variant, or domain thereof) is murine leukemia virus (MLV) RT, Moloney murine leukemia virus (M-MLV) RT, Raul's sarcoma virus (RSV) RT, avian myeloblastoma virus (AMV) RT, human immunodeficiency virus (HIV) RT (e.g., HIV-1 RT, HIV-2 RT), avian leukemia virus RT, mouse mammary tumor virus, feline leukemia virus, bovine leukemia virus (ALV) RT, human t-lymphotropic virus (HTLV) RT (e.g., HTLV-1 RT), simian immunodeficiency virus (SIV) RT, or feline immunodeficiency virus (FIV) RT.

[0162] In some embodiments, RT (or a functional fragment, variant, or domain thereof) is a non-LTR retrotransposon. In some embodiments, RT (or a functional fragment, variant, or domain thereof) is an APE-type non-LTR retrotransposon. In some embodiments, RT (or a functional fragment, variant, or domain thereof) is an APE-type non-LTR retrotransposon from the R1 or Txl clade. In some embodiments, RT (or a functional fragment, variant, or domain thereof) is an RLE-type non-LTR retrotransposon. In some embodiments, RT (or a functional fragment, variant, or domain thereof) is an RLE-type non-LTR retrotransposon from the R2, NeSL, HERO, R4, or CRE clades. In some embodiments, RT (or a functional fragment, variant, or domain thereof) is an R2 RLE-type non-LTR retrotransposon. In some embodiments, the RT (or a functional fragment, functional variant, or domain thereof) is an RT derived from an R2Bm non-LTR retrotransposon, an RT derived from an R2Tg non-LTR retrotransposon, an RT derived from a LINE-1 non-LTR retrotransposon, or an RT derived from a Penelope or Penelope-like element (PLE) non-LTR retrotransposon.

[0163] In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is an LTR retrotransposon (e.g., RT from the Tyl LTR retrotransposon). In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is a group II intron. In some embodiments, RT (or a functional fragment, functional variant thereof) is the group II intron maturation enzyme RT (Marathon RT) from *Eubacterium rectale* (see, e.g., Zhao et al., RNA 24:2 2018, the entire contents of which are incorporated herein by reference for all purposes); the group II intron LtrART; or the thermostable group II intron RT (TGIRT). In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is a diversity-producing retrotranscriptional element (e.g., a diversity-producing retrotranscriptional element from *Bordetella* bacteriophage BPP-1). In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is a reverse transcriptase (e.g., reverse transcriptase from Ec86 (RT86)). In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is a telomerase (e.g., RT from TERT telomerase). In some embodiments, RT (or a functional fragment, functional variant, or domain thereof) is a reverse transcriptase plasmid (e.g., RT from the Mauriceville plasmid).

[0164] The amino acid sequences of exemplary RTs are provided in Table 2 and SEQ ID NO: 324-476. The accession number for each exemplary RT is also provided in Table 2. Table 2. Amino acid sequences of exemplary reverse transcriptases.

[0165] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with, the amino acid sequence of the polypeptide shown in Table 2.

[0166] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 2, and further comprises one or more but less than 15% (e.g., less than 12%, less than 10%, less than 8%) amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 2, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 2, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 2, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitutions, additions, deletions, etc.).

[0167] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 2, and further comprises one or more, but less than 15% (e.g., less than 12%, less than 10%, less than 8%), amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 2, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 2, and further comprises, or is composed of about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 2, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions.

[0168] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment or functional variant thereof) comprises, or is composed of, an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with, the amino acid sequence of any one of SEQ ID NO: 324-476.

[0169] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment or functional variant thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.).

[0170] In some embodiments, the amino acid sequence of the reverse transcriptase (or its functional fragment, functional variant, or domain) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises one or more, but less than 15% (e.g., less than 12%, less than 10%, less than 8%) amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or its functional fragment, functional variant, or domain) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or its functional fragment, functional variant, or domain) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises, or is composed of about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or consists of, any one of the amino acid sequences in SEQ ID NO: 324-476, and further comprises, or consists of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions.

[0171] In some embodiments, RT is the RT (or a functional fragment, functional variant, or domain thereof) described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44) ​​and WO 2023039424 (see, for example, Table 6), the entire contents of which are incorporated herein by reference for all purposes.

[0172] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, an amino acid sequence that is at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44) ​​and WO 2023039424 (see, for example, Table 6).

[0173] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises one or more but less than 15% (e.g., less than 12%, less than 10%, less than 8%) amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitutions, additions, deletions, etc.).

[0174] In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises, or is composed of, one or more amino acid substitutions but less than 15% (e.g., less than 12%, less than 10%, less than 8%). In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide described in WO 2021178720 (see, for example, Tables 1, 2, 3, 30, 41, and 44), and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. 4.3.1.2 Nucleobase Editor

[0175] In some embodiments, the heterologous protein (or a functional fragment, functional variant, or domain thereof) exhibits nucleobase editing activity. In some embodiments, the heterologous protein (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, a nucleobase editing domain of a nucleobase editor (e.g., a domain capable of modifying nucleobases (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA).

[0176] In some embodiments, the heterologous protein is a nucleobase editor (or a functional fragment, functional variant, or domain thereof). In some embodiments, the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, a nucleobase editing domain (e.g., a domain capable of modifying bases (e.g., A, T, C, G, or U) within a nucleic acid molecule (e.g., DNA). In some embodiments, the nucleobase editor is a deaminase (or a functional fragment, functional variant, or domain thereof). In some embodiments, the deaminase is a cytidine deaminase (or a functional fragment, functional variant, or domain thereof). In some embodiments, the deaminase is adenosine deaminase (or a functional fragment, functional variant, or domain thereof).

[0177] In some embodiments, a nucleobase editor comprises a naturally occurring nucleobase editor (e.g., a deaminase) (or a functional fragment, functional variant, or domain thereof). In some embodiments, a nucleobase editor (e.g., a deaminase) comprises a functional fragment of a naturally occurring nucleobase editor. In some embodiments, a nucleobase editor (e.g., a deaminase) comprises a functional variant of a naturally occurring nucleobase editor. In some embodiments, a nucleobase editor (e.g., a deaminase) comprises both functional fragments and variants of a naturally occurring nucleobase editor. In some embodiments, a nucleobase editor (e.g., a deaminase) comprises one or more domains of a naturally occurring nucleobase editor. In some embodiments, a nucleobase editor (e.g., a deaminase) comprises functional fragments of one or more domains of a naturally occurring nucleobase editor. In some embodiments, a nucleobase editor (e.g., a deaminase) comprises functional variants of one or more domains of a naturally occurring nucleobase editor. In some embodiments, a nucleobase editor (e.g., a deaminase) comprises both functional fragments and functional variants of one or more domains of a naturally occurring nucleobase editor.

[0178] In some embodiments, a nucleobase editor (e.g., a deaminase) is a eukaryotic nucleobase editor (or a functional fragment, functional variant, or domain thereof). In some embodiments, a nucleobase editor (e.g., a deaminase) is a prokaryotic nucleobase editor (or a functional fragment, functional variant, or domain thereof). In some embodiments, a nucleobase editor (e.g., a deaminase) is a viral nucleobase editor (or a functional fragment, functional variant, or domain thereof). In some embodiments, a nucleobase editor (e.g., a deaminase) is a bacterial nucleobase editor (or a functional fragment, functional variant, or domain thereof).

[0179] Naturally occurring nucleobase editors (e.g., deaminases (e.g., cytidine deaminase, adenosine deaminase)) are known in the art and described herein (see, for example, Table 3).

[0180] For example, naturally occurring cytidine deaminases include, but are not limited to, the apolipoprotein B mRNA editing complex (APOBEC) family of deaminases and cytidine deaminase 1 (CDA1). The APOBEC family includes, for example, but not limited to, APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now typically referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine or cytosine) deaminases (AID). Cytidine deaminases can be derived from any suitable organism, including, for example, humans, chimpanzees, gorillas, monkeys, cows, dogs, rats, or mice. Exemplary cytidine deaminases are described in WO 2022 / 204268, the entire contents of which are incorporated herein by reference for all purposes.

[0181] Naturally occurring adenosine deaminases include, for example, but not limited to, adenosine deaminase ADAR (e.g., ADAR1, ADAR2), adenosine deaminase ADAT, and TadA (e.g., from *Escherichia coli* (ecTadA)). TadA and its variants are known in the art and described, for example, in WO 2018 / 027078 and WO 2022 / 204268, the entire contents of which are incorporated herein by reference for all purposes. Adenosine deaminases can be derived from any suitable organism (e.g., *Escherichia coli*). In some embodiments, adenosine deaminases are derived from *Escherichia coli*, *Staphylococcus aureus*, *Salmonella typhi*, *Shewanella putrefaciens*, *Haemophilus influenzae*, *Caulobacter crescentus*, or *Bacillus subtilis*. In some embodiments, adenosine deaminases are derived from *Escherichia coli*. In some embodiments, adenosine deaminase is ecTadA. In some embodiments, ecTadA is a variant as described in WO 2018 / 027078 or WO 2022 / 204268, the entire contents of which are incorporated herein by reference for all purposes.

[0182] In some embodiments, the adenosine deaminase is a variant TadA deaminase. In some embodiments, the variant TadA deaminase is the deaminase described in WO 2022 / 204268 (see, for example, Table 3, pp. 91-93), the entire contents of which are incorporated herein by reference for all purposes. In some embodiments, TadA is provided as a monomer or a dimer (e.g., a heterodimer of wild-type Escherichia coli TadA and engineered TadA variants). In some embodiments, the adenosine deaminase is an eighth-generation TadA as described in WO2022 / 204268 (see, for example, Table 4). 8 variants. In some embodiments, the adenosine deaminase is the eighth-generation TadA as shown in WO 2022 / 204268 (see, for example, pages 91-92). Eight variants, all of which are incorporated herein by reference for all purposes.

[0183] Exemplary nucleobase editors are described, for example, in WO 2022 / 204268; WO 2018 / 027078; WO 2017 / 070632; Komor, AC et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature [Nature] 533, 420-424 (2016); Gaudelli, NM et al., “Programmable base editing of A·T to G»C in genomic DNA without DNA cleavage” Nature [Nature] 551, 464-471 (2017); Komor, AC et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A baseeditors with higher efficiency and product purity [Improved base excision repair inhibition and phage Mu Gam protein production with higher efficiency and product purity C:G to T:A base editors]"Science Advances 3:eaao4774 (2017); and Rees, HA et al., "Base editing: precision chemistry on the genome and transcriptome of living cells." Nat Rev Genet. Dec 2018;19(12):770-788. doi: 10.1038 / s41576-018-0059-l, the entire contents of which are hereby incorporated herein by reference for all purposes.

[0184] The amino acid sequences of exemplary nucleobase editors are provided in Table 3. Table 3. Amino acid sequences of exemplary nucleobase editors.

[0185] In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with, the amino acid sequence of the polypeptide shown in Table 3.

[0186] In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 3, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) of amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 3, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 3, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 3, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitutions, additions, deletions, etc.).

[0187] In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 3, and further comprises one or more, but less than 15% (less than 12%, less than 10%, less than 8%), amino acid substitutions. In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 3, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptides shown in Table 3, and further comprises, or is composed of about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, the amino acid sequence of the polypeptide shown in Table 3, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions.

[0188] In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with, the amino acid sequence of any one of SEQ ID NO: 477-536.

[0189] In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) of amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitutions, additions, deletions, etc.). In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises, or is composed of, about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.). In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises, or is composed of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid variations (e.g., substitution, addition, deletion, etc.).

[0190] In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) amino acid substitutions. In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises, or is composed of, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment, functional variant, or domain thereof) comprises, or is composed of any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises, or is composed of about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid substitutions. In some embodiments, the amino acid sequence of the nucleobase editor (or a functional fragment or functional variant thereof) comprises, or consists of, any one of the amino acid sequences in SEQ ID NO: 477-536, and further comprises, or consists of, no more than about 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions.

[0191] The nucleobase editor described herein can be further operably linked (e.g., fused) to another heterologous portion (e.g., a heterologous protein). In some embodiments, the nucleobase editor described herein can be further operably linked (e.g., fused) to another heterologous portion (e.g., a heterologous protein). In some embodiments, the nucleobase editor is fused to a base excision repair inhibitor, such as a glycosylation inhibitor (UGI) domain or a nuclease-dead inosine-specific nuclease (dISN) domain. 4.3.2 Connector

[0192] As described herein, a heterologous portion (e.g., a heterologous protein (e.g., a reverse transcriptase, a nucleobase editor)) can be directly operably linked to or indirectly operably linked to a Cas endonuclease (e.g., as described herein). In some embodiments, the heterologous protein is directly operably linked to a Cas endonuclease (e.g., as described herein). In some embodiments, the heterologous polypeptide is directly operably linked to a Cas endonuclease via a peptide bond (e.g., as described herein). In some embodiments, the heterologous protein is indirectly operably linked to a Cas endonuclease (e.g., as described herein). In some embodiments, the heterologous protein is indirectly operably linked to a Cas endonuclease via a linker (e.g., as described herein).

[0193] In some embodiments, the heterologous protein is operatively ligated to a Cas endonuclease indirectly via a peptide linker (e.g., as described herein). In some embodiments, the peptide linker is one or any combination of a cleavable linker, an uncleavable linker, a flexible linker, a rigid linker, a helical linker, and / or a non-helical linker. In some embodiments, the peptide linker comprises or contains about 2-30, 5-30, 10-30, 15-30, 20-30, 25-30, 2-25, 5-25, 10-25, 15-25, 20-25, 2-20, 5-20, 10-20, 15-20, 2-15, 5-15, 10-15, 2-10, or 5-10 amino acid residues. In some embodiments, the peptide linker comprises, or is composed of, at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid residues. In some embodiments, the linker comprises, or is composed of, about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid residues. In some embodiments, the linker comprises, or consists of, no more than about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acid residues. In some embodiments, the amino acid sequence of the peptide linker comprises, or consists of, glycine amino acid residues, serine amino acid residues, or both glycine amino acid residues and serine amino acid residues. In some embodiments, the amino acid sequence of the peptide linker comprises, or consists of, glycine amino acid residues, serine amino acid residues, and proline amino acid residues.

[0194] The amino acid sequences of exemplary peptide linkers are provided in Table 4. Table 4. Amino acid sequences of exemplary peptide linkers.

[0195] In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, the amino acid sequence of any of the linkers shown in Table 4. In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, the amino acid sequence of any of the linkers shown in Table 4, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) of amino acid variations (e.g., amino acid substitution, deletion, or addition). In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, the amino acid sequence of any of the linkers shown in Table 4, and comprises one, two, or three amino acid variations (e.g., substitution, deletion, addition). In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, the amino acid sequence of any of the linkers shown in Table 4, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) of amino acid substitutions. In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, the amino acid sequence of any of the linkers shown in Table 4, and comprises one, two, or three amino acid substitutions.

[0196] In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 537-658. In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 537-658, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) of amino acid variations (e.g., amino acid substitution, deletion, or addition). In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 537-658, and comprises one, two, or three amino acid variations (e.g., substitution, deletion, addition). In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 537-658, and further comprises one or more but less than 15% (less than 12%, less than 10%, less than 8%) of amino acid substitutions. In some embodiments, the amino acid sequence of the peptide linker comprises, or is composed of, any one of the amino acid sequences in SEQ ID NO: 537-658, including one, two, or three amino acid substitutions.

[0197] In some embodiments, the connector is the connector (or a functional fragment, functional variant, or structural domain thereof) described in WO 2021178720 or WO 2023039424, the entire contents of which are incorporated herein by reference for all purposes. 4.3.3 Orientation

[0198] One or more heterologous portions (e.g., one or more heterologous proteins) and Cas endonucleases (e.g., those described herein) (or functional fragments, functional variants, or domains thereof) may be arranged in any conformation or order, provided that the Cas endonuclease protein (e.g., those described herein) (or functional fragments, functional variants, or domains thereof) maintains its ability to mediate its function, and in embodiments where the heterologous portion (e.g., the heterologous protein) has a specific function, the heterologous portion (e.g., the heterologous protein) may mediate its function.

[0199] In some embodiments, a heterologous portion (e.g., a heterologous protein) is operatively attached to the N-terminus, C-terminus, or interior between the N-terminus and C-terminus of a Cas endonuclease (or a functional fragment, functional variant, or domain thereof). In some embodiments, a heterologous portion (e.g., a heterologous protein) is operatively attached to the C-terminus of a Cas endonuclease (or a functional fragment, functional variant, or domain thereof). In some embodiments, a heterologous portion (e.g., a heterologous protein) is operatively attached to both the N-terminus and C-terminus of the endonuclease (or a functional fragment, functional variant, or domain thereof).

[0200] In some embodiments, the heterologous portion is a heterologous protein (e.g., a polymerase (e.g., reverse transcriptase), nucleobase editor (e.g., deaminase) (e.g., described herein)) that forms a fusion protein with a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein). In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein) and a heterologous protein (e.g., a polymerase (e.g., reverse transcriptase), nucleobase editor (e.g., deaminase) (e.g., described herein)). In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein), a peptide linker (e.g., described herein), and a heterologous protein (e.g., a polymerase (e.g., reverse transcriptase), nucleobase editor (e.g., deaminase) (e.g., described herein)). In this particular orientation, the C-terminus of an endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein) is operatively linked directly or indirectly to the N-terminus of a heterologous protein (e.g., a polymerase (e.g., a reverse transcriptase), a nucleobase editor (e.g., a deaminase) (e.g., described herein)) via a peptide linker (e.g., described herein).

[0201] In some embodiments, the heterologous portion is a heterologous protein (e.g., a polymerase (e.g., reverse transcriptase), nucleobase editor (e.g., deaminase) (e.g., described herein)) that forms a fusion protein with a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein). In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a heterologous protein (e.g., a polymerase (e.g., reverse transcriptase), nucleobase editor (e.g., deaminase) (e.g., described herein)) and a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein). In some embodiments, the fusion protein comprises, from N-terminus to C-terminus: a heterologous protein (e.g., a polymerase (e.g., reverse transcriptase), nucleobase editor (e.g., deaminase) (e.g., described herein)), a peptide linker (e.g., described herein), and a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein). In this particular orientation, the C-terminus of a heterologous protein (e.g., a polymerase (e.g., a reverse transcriptase), a nucleobase editor (e.g., a deaminase) (e.g., described herein)) is operatively linked directly or indirectly to the N-terminus of an endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein) via a peptide linker (e.g., described herein). 4.4 Methods for protein preparation

[0202] The proteins described herein (e.g., Cas endonucleases, fusion proteins, and conjugates) can be produced using standard methods known in the art. For example, each protein can be produced via recombinant technology in host cells (e.g., insect cells, mammalian cells, bacteria) transfected or transduced with a nucleic acid expression vector (e.g., plasmid, viral vector (e.g., baculovirus expression vector)) encoding the protein (e.g., endonuclease, fusion protein, etc.). These general methods are well known in the art. Expression vectors typically contain an expression cassette comprising a nucleic acid sequence capable of inducing expression of a nucleic acid molecule encoding a target protein (e.g., Cas endonuclease, fusion protein, etc.), such as one or more promoters, one or more enhancers, polyadenylation signals, etc. Those skilled in the art will recognize that a variety of promoter and enhancer elements can be used to achieve expression of nucleic acid molecules in host cells. For example, promoters can be constitutive or regulatory and can be obtained from various sources (e.g., viral, prokaryotic, or eukaryotic sources) or artificially designed. Following transfection or transduction, host cells containing an expression vector encoding the target protein are cultured under conditions conducive to the expression of nucleic acid molecules encoding the target protein (e.g., endonucleases, fusion proteins, etc.). Culture media can be obtained from various suppliers, and suitable media can be routinely selected to enable host cells to express the target protein. Host cells can be adherent or suspension cultures, and those skilled in the art can optimize culture methods for specific host cells. For example, suspension cells can be cultured in, for example, a bioreactor in a batch or fed-batch process. The produced protein can be separated from the cell culture by, for example, column chromatography in flow-through or binding-elution mode. Examples include, but are not limited to, ion exchange resins and affinity resins (such as lentil lectin agarose gel) and mixed-mode cation exchange-hydrophobic interaction columns (CEX-HIC). The protein can be concentrated by buffer exchange via ultrafiltration, and the effluent from ultrafiltration can be filtered through a suitable filter (e.g., a 0.22 μm filter). See, for example, Hacker, David (ed.), Recombinant Protein Expression in Mammalian Cells: Methods and Protocols, Humana Press (2018). See also U.S. Patent 5,762,939, the entire contents of which are incorporated herein by reference for all purposes. The proteins described herein (e.g., Cas endonucleases, fusion proteins, and protein conjugates) can be produced synthetically.

[0203] This disclosure provides, in particular, a method for preparing the proteins described herein (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.), the method comprising (a) introducing a nucleic acid molecule encoding the protein (e.g., an endonuclease (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.) into a host cell; (b) culturing the host cell (e.g., under conditions and for a duration sufficient to allow expression of the protein (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.)); and optionally isolating the protein (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.) from the culture medium.

[0204] This disclosure further provides a method for preparing the proteins described herein (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.), the method comprising (a) expressing the proteins (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.) in a recombinant manner; (b) enriching (e.g., purifying) the proteins (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.); (c) evaluating the proteins (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.) for the presence of process impurities or contaminants; and (d) formulating the proteins (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.) into pharmaceutical compositions if the proteins (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), fusion proteins, etc.) meet the threshold specifications for process impurities or contaminants. The process impurities or contaminants being evaluated can be one or more of the following: process-related impurities, such as host cell proteins, host cell DNA, or cell culture components (e.g., inducers, antibiotics, or culture medium components); product-related impurities (e.g., precursors, fragments, aggregates, degradation products); or contaminants, such as endotoxins, bacteria, or viral contaminants. 4.5 system

[0205] This document further provides, in particular, a system comprising a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., described herein) (or a fusion protein or conjugate of any of the foregoing (e.g., described herein)), which is particularly useful for editing nucleic acid molecules (e.g., DNA, genome, genes (e.g., intracellularly, e.g., within the cells of a subject (e.g., a mammalian subject, e.g., a human subject)) (e.g., in vivo, in vitro, or ex vivo). In some embodiments, the system can be used to mediate the addition of one or more nucleotides (e.g., nucleic acid (DNA) molecules) to a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., intracellularly, e.g., within the cells of a subject (e.g., a mammalian subject, e.g., a human subject), the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid.

[0206] Therefore, this document provides systems comprising: (a) (i) the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof); (ii) a fusion protein comprising the Cas endonuclease described herein (or a functional fragment, functional variant thereof); (iii) a conjugate comprising the Cas endonuclease described herein (or a functional fragment, functional variant thereof); (iv) a nucleic acid molecule encoding (a)(i), (a)(ii), and / or (a)(iii) (e.g., a nucleic acid molecule described herein); (v) a vector comprising (a)(iv) (e.g., a vector described herein); (vi) a carrier comprising any one of (a)(i)-(a)(v) (e.g., a carrier described herein); or (vii) a composition comprising any one of (a)(i)-(a)(vi) (e.g., a pharmaceutical composition described herein).

[0207] In some embodiments, the system comprises (a) (i) a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) as described herein; (ii) a fusion protein comprising a Cas endonuclease (or a functional fragment, functional variant thereof) as described herein (e.g., as described herein); (iii) a conjugate comprising a Cas endonuclease (or a functional fragment, functional variant thereof) as described herein (e.g., as described herein); (iv) a nucleic acid molecule encoding (a)(i), (a)(ii), or (a)(iii) (e.g., a nucleic acid molecule described herein); (v) a vector comprising (a)(iv) (e.g., a vector described herein); (vi) a carrier comprising any one of (a)(i)-(a)(v) (e.g., a carrier described herein); or (vii) a composition comprising any one of (a)(i)-(a)(vi) (e.g., a pharmaceutical composition) (e.g., a composition described herein (e.g., a pharmaceutical composition)); and (b) (i) (i) a first gRNA (e.g., crRNA and tracrRNA; sgRNA; template RNA (e.g., as described herein)) or (ii) a nucleic acid (e.g., DNA) molecule that encodes the first gRNA (e.g., crRNA and tracrRNA; sgRNA; template RNA (e.g., as described herein)).

[0208] As described above, the system provided herein can be used in particular for editing nucleic acid molecules (e.g., DNA, genome, genes (e.g., within cells, for example, within the cells of a subject (e.g., mammalian subjects, for example, human subjects)) (e.g., in vivo, in vitro, or ex vivo). In some embodiments, the system provided herein may include one or more of the following features (e.g., any combination or all of them): (a) the system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is capable of binding gRNA (e.g., as described herein); (b) the system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is capable of forming breaks in target nucleic acid (e.g., DNA (e.g., dsDNA)) molecules (e.g., as described herein); (c) the system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is capable of forming single-strand breaks in the edited strand (as defined herein) of target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules (e.g., as described herein); (d) the system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is capable of forming single-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules (e.g., as described herein); (e) The system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is capable of forming double-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules (e.g., as described herein); (f) the system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is unable to form double-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules (e.g., as described herein); (g) the system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is capable of forming single-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules (e.g., as described herein) and is unable to form double-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules (e.g., as described herein) (e.g., exhibiting nicking enzyme activity); (h) The system's Cas endonuclease (or a functional fragment, functional variant, or domain thereof) is capable of forming single-strand breaks in the edited strand (as defined herein) of a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule (e.g., as described herein) and is not capable of forming double-strand breaks in a target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule (e.g., as described herein); and / or (i) the system is capable of mediating the addition of one or more nucleotides to a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., as described herein), the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid. 4.5.1 Target nucleic acid molecules

[0209] As described above, in some embodiments, the system is capable of mediating any of the aforementioned effects in the target nucleic acid molecule (see, for example, § 4.5). In some embodiments, the target nucleic acid molecule is a DNA molecule. In some embodiments, the target nucleic acid molecule is a dsDNA molecule. In some embodiments, a portion of the nucleotide sequence of the unedited strand of the target dsDNA molecule (as defined herein) is complementary to at least a portion of the nucleotide sequence of the system's gRNA (e.g., the gRNA described herein (see, for example, § 4.5.2)).

[0210] In some embodiments, the target nucleic acid molecule is within the genome of a cell (e.g., a eukaryotic cell) (e.g., within a subject (e.g., a human subject)). In some embodiments, the target nucleic acid molecule is a gene (e.g., within a cell (e.g., a eukaryotic cell) (e.g., within a subject (e.g., a human subject)). In some embodiments, the target nucleic acid molecule is within the genome of a cell (e.g., a eukaryotic cell) in vitro, ex vivo, or in vivo. In some embodiments, the target nucleic acid molecule is within the genome of a cell (e.g., a eukaryotic cell) within a subject (e.g., a human subject). 4.5.2gRNA

[0211] In some embodiments, the system comprises guide RNA (gRNA). gRNAs are generally known in the art and are described herein. See, for example, Nishimasu et al., Cell 156, pp. 935-949 (2014), the entire contents of which are incorporated herein by reference for all purposes. As described above, gRNAs include RNAs comprising crRNA and tracrRNA; sgRNA; and template RNA (e.g., as described herein). In some embodiments, the system comprises a nucleic acid (e.g., DNA) molecule encoding any one or more of the aforementioned gRNAs (e.g., crRNA and tracrRNA; sgRNA; template RNA (e.g., as described herein)). Where gRNAs are described herein, this disclosure further covers nucleic acid (e.g., DNA) molecules encoding gRNAs.

[0212] In some embodiments, at least a portion of the nucleotide sequence of the gRNA is complementary to a portion of the nucleotide sequence of a target nucleic acid molecule (e.g., as described herein). In some embodiments, at least a portion of the nucleotide sequence of the gRNA is complementary to a portion of the nucleotide sequence of the unedited strand (as defined herein) of a double-stranded nucleic acid (e.g., dsDNA) target nucleic acid molecule (e.g., as described herein). In some embodiments, at least a portion of the nucleotide sequence of the gRNA binds to a portion of the nucleotide sequence of the edited strand (as defined herein) of a double-stranded nucleic acid (e.g., dsDNA) target nucleic acid molecule (e.g., as described herein).

[0213] In some embodiments, the system comprises crRNA and tracrRNA (or multiple different crRNAs and multiple different tracrRNAs), wherein the crRNA and tracrRNA are on separate RNA molecules. In some embodiments, the system comprises a nucleic acid molecule encoding crRNA and a separate nucleic acid molecule encoding tracrRNA. In some embodiments, the system comprises multiple nucleic acid molecules, each encoding a different crRNA; and multiple nucleic acid molecules, each encoding a tracrRNA (wherein each encoded tracrRNA may be the same or different).

[0214] In some embodiments, the system comprises sgRNA (or multiple different sgRNAs). In some embodiments, the system comprises a nucleic acid (e.g., DNA) molecule encoding the sgRNA. In some embodiments, the system comprises multiple nucleic acid molecules, each encoding a different sgRNA. In some embodiments, the crRNA of each of the multiple sgRNAs is different. In some embodiments, the tracrRNA of each of the multiple sgRNAs is different. In some embodiments, the tracrRNA of each of the multiple sgRNAs is the same. In some embodiments, the crRNA of each of the multiple sgRNAs is different, and the tracrRNA of each of the multiple sgRNAs is the same.

[0215] In some embodiments, the system comprises a template RNA (e.g., a single template RNA, multiple different template RNAs) or a nucleic acid (e.g., DNA) molecule encoding the template RNA (or multiple nucleic acid (e.g., DNA) molecules, each encoding a different template RNA). In some embodiments, the template RNA comprises crRNA, tracrRNA, a heterologous object sequence, and a 3' target homologous domain from 5' to 3'. In some embodiments, the template RNA further comprises a sequence of a polymerase (e.g., a reverse transcriptase, such as the reverse transcriptase of the fusion protein described herein). In some embodiments, the template RNA comprises crRNA, tracrRNA, a sequence of a polymerase (e.g., a reverse transcriptase, such as the reverse transcriptase of the fusion protein described herein), a heterologous object sequence, and a 3' target homologous domain. In some embodiments, the template RNA comprises crRNA, tracrRNA, a sequence of a polymerase (e.g., a reverse transcriptase, such as the reverse transcriptase of the fusion protein described herein), a heterologous object sequence, and a 3' target homologous domain from 5' to 3'.

[0216] In some embodiments, the gRNA (e.g., template RNA) comprises a nucleic acid molecule containing a toe loop, hairpin, stem loop, pseudoknot (e.g., the Mpknot1 moiety), aptamer, G-quadruplex, tRNA, riboswitch, or ribozyme. In some embodiments, the gRNA (e.g., template RNA) comprises a nucleic acid molecule containing a pseudoknot (e.g., the Mpknot1 moiety). In some embodiments, one or more 3' hairpin elements of the gRNA may be removed, for example, as described in WO 2018106727, the entire contents of which are incorporated herein by reference for all purposes. In some embodiments, the gRNA may contain additional hairpin structures, for example, as described in Kocak et al., Nat Biotechnol [Nature Biotechnology] 37(6):657-666 (2019), the entire contents of which are incorporated herein by reference for all purposes. Secondary structures in gRNA (e.g., hairpins) can be predicted on a computer using software tools, such as RNAstructure tools, which are available at ma.urmc.rochester.edu / RNAstructureWeb (Bellaousov et al., Nucleic Acids Res [Nucleic Acids Research] 41: W471-W474 (2013); which is incorporated herein by reference in its entirety).

[0217] Custom gRNA generators and algorithms are commercially available for use in gRNA design. 4.5.2.1 Multiple gRNAs

[0218] In some embodiments, the system comprises multiple gRNAs (e.g., multiple sgRNAs, multiple template RNAs). In some embodiments, the system comprises multiple nucleic acid molecules, each encoding a gRNA (e.g., sgRNA, template RNA).

[0219] In some embodiments, the system comprises a first gRNA (e.g., sgRNA, template RNA) and a second gRNA (e.g., sgRNA, template RNA). In some embodiments, the first gRNA is sgRNA, and the second gRNA is sgRNA. In some embodiments, the first gRNA is sgRNA, and the second gRNA is sgRNA, wherein the nucleotide sequences of the crRNA of the first gRNA and the second gRNA are different. In some embodiments, the first gRNA is template RNA, and the second gRNA is sgRNA. In some embodiments, the first gRNA is template RNA, and the second gRNA is sgRNA, wherein the nucleotide sequences of the crRNA of the first gRNA and the second gRNA are different.

[0220] In some embodiments, a second gRNA (e.g., sgRNA) is capable of guiding a systemic endonuclease (e.g., as described herein) to form a single-strand break in the unedited strand of a target double-stranded nucleic acid (e.g., dsDNA) molecule. In some embodiments, at least a portion of the nucleotide sequence of the second gRNA (e.g., sgRNA) is complementary to a portion of the nucleotide sequence of the edited strand (as defined herein) of the double-stranded nucleic acid (e.g., dsDNA) molecule. In some embodiments, at least a portion of the nucleotide sequence of the second gRNA (e.g., sgRNA) binds to a portion of the nucleotide sequence of the edited strand (as defined herein) of the double-stranded nucleic acid (e.g., dsDNA) molecule.

[0221] In some embodiments, the second gRNA (e.g., sgRNA) is present on the same nucleic acid molecule as the first gRNA (or the nucleic acid (e.g., DNA) molecule encoding the second gRNA is present on the same nucleic acid (e.g., DNA) molecule encoding the first gRNA). In some embodiments, the second gRNA (e.g., sgRNA) is present on a different nucleic acid molecule than the first gRNA (or the nucleic acid (e.g., DNA) molecule encoding the second gRNA is present on a different nucleic acid (e.g., DNA) molecule encoding the first gRNA). 4.5.2.2 Modified gRNA

[0222] In some embodiments, the gRNA (e.g., the gRNA of the system described herein) comprises one or more modified nucleotides (as defined herein) (referred to as modified gRNA). Compared to a corresponding unmodified gRNA, the modified gRNA may have one or more distinct (e.g., improved) properties (e.g., one or more improved properties in vivo). For example, in some embodiments, the modified gRNA (e.g., terminally modified gRNA) may exhibit increased stability in cells (e.g., in vitro, in vivo, or extracellularly) (e.g., compared to unmodified gRNA). In some embodiments, the modified gRNA (e.g., terminally modified gRNA) may exhibit increased stability in vivo (e.g., compared to unmodified gRNA). In some embodiments, the system described herein utilizing modified gRNA exhibits increased nucleic acid (e.g., gene) editing efficiency (e.g., compared to a system containing unmodified gRNA). In some embodiments, the system described herein utilizing modified gRNA exhibits increased on-target nucleic acid (e.g., gene) editing (e.g., compared to a system containing unmodified gRNA). In some embodiments, the system described herein utilizing modified gRNA exhibits reduced off-target nucleic acid (e.g., gene) editing (e.g., compared to a system containing unmodified gRNA). In some embodiments, the systems described herein that utilize modified gRNA exhibit increased affinity for DNA molecule editing (e.g., the system's gRNA exhibits increased affinity for DNA molecules) (e.g., compared to systems containing unmodified gRNA).

[0223] Modified gRNAs can be selected and tested using methods known in the art. For example, structure-guided and systematic approaches (e.g., as described in Mir, A., Alterman, JF, Hassler, MR et al., Heavily and fully modified RNAs guide efficient SpyCas9-mediated genome editing. Nat Commun [Nature Communications] 9,2641 (2018). https: / / doi.org / 10.1038 / s41467-018-05073-z; the entire contents of which are incorporated herein by reference for all purposes) can be used to discover and select gRNA modifications.

[0224] gRNA modification is known in the art and is described herein. See, for example, Allen Daniel et al., Using Synthetically Engineered Guide RNAs to Enhance CRISPR Genome Editing Systems in Mammalian Cells, Frontiers in Genome Editing, Vol. 2 (Article 617910) (2021) DOI=10.3389 / fgeed.2020.617910; and Hendel A, Bak RO, Clark JT et al., Chemically modified guide RNAs enhance CRISPR-Cas genomeediting in human primary cells. Nat Biotechnol. 2015;33(9):985-989. doi:10.1038 / nbt.3290; the entire contents of each are incorporated herein by reference for all purposes.

[0225] The exemplary modifications described herein are primarily relating to gRNA. It should be understood that corresponding modifications can be made to the DNA molecule encoding gRNA. Such corresponding DNA modifications are known in the art and readily identified by those skilled in the art. Therefore, modifications to "gRNA" also include corresponding modifications to the DNA molecule encoding gRNA. (i) Properties of the modification

[0226] Nucleotide modification may include modification of any one or more nucleosides and / or inter-nucleoside bonds. Nucleoside modification includes modification of the sugar (e.g., ribose) moiety and / or nucleotides. In some embodiments, the modified gRNA comprises one or more nucleotides containing a modified sugar (e.g., ribose) moiety. In some embodiments, the modified gRNA comprises one or more nucleotides containing a modified nucleotide. In some embodiments, the modified gRNA comprises one or more nucleotides containing a modified inter-nucleoside bond. In some embodiments, the modified gRNA comprises one or more nucleotides containing one, two, or three of a modified sugar (e.g., ribose) moiety, a modified nucleotide, and / or a modified inter-nucleoside bond. In some embodiments, the modified gRNA comprises one or more nucleotides containing a modified sugar (e.g., ribose) moiety and a modified inter-nucleoside bond.

[0227] Exemplary nucleoside modifications are described below and are also known in the art, see, for example, WO 2018107028A1 (see, for example, Table 4 (as identified by SEQ ID NO)); US 20190316121; Hendel A, Bak RO, Clark JT et al., Chemically modified guide RNAs enhance CRISPR-Cas genomeediting in human primary cells. Nat Biotechnol. [Nature Biotechnology] 33(9):985-989 (2015) doi:10.1038 / nbt.3290; Mir et al., Nat Commun [Nature Communications] 9:2641 (2018) (see, for example, Supplementary Table 1); Allen D, Rosenberg M and Hendel A (2021) Using Synthetically Engineered Guide RNAs to Enhance CRISPR Genome Editing Systems in Mammalian Cells. [Enhancing CRISPR Genome Editing Systems in Mammalian Cells with Synthetically Engineered Guide RNA] Front. GenomeEd. [Frontiers in Genome Editing] 2:617910. doi: 10.3389 / fgeed.2020.617910; all the contents of each are incorporated herein by reference for all purposes. (a) Sugar modification

[0228] In some embodiments, the modified gRNA comprises one or more nucleosides containing a modified sugar (e.g., ribose) moiety.

[0229] The modified ribose moiety may contain substituents at any one or more positions (including, for example, positions 2', 4', and / or 5') of the sugar (e.g., ribose). In some embodiments, the modified sugar (e.g., ribose) contains a substituent at the 2' position of the sugar (e.g., ribose). In some embodiments, the modified sugar (e.g., ribose) contains a substituent at the 4' position of the sugar (e.g., ribose). In some embodiments, the modified sugar (e.g., ribose) contains a substituent at the 5' position of the sugar (e.g., ribose).

[0230] In some embodiments, the gRNA contains any one or more of the following substituents (e.g., at any position of the sugar (e.g., ribose) (e.g., at position 2'): a group for improving the stability of the gRNA, a group for improving the pharmacokinetic properties of the gRNA, a group for improving the pharmacodynamic properties of the gRNA, an RNA cleaving group, a reporter group, an intercalator, or other substituents with similar properties.

[0231] Exemplary substituents include, for example, but not limited to, substitution by any of the following (e.g., at any position in a sugar (e.g., ribose)): OH; F; O-alkyl, S-alkyl, or N-alkyl; O-alkenyl, S-alkenyl, or N-alkenyl; O-ynyl, S-ynyl, or N-ynyl; or O-alkyl-O-alkyl, wherein the alkyl, alkenyl, and ynyl groups may be substituted or unsubstituted C1 to C2 groups. 10 Alkyl or C2 to C 10 Alkenyl and alkynyl groups. Other exemplary substitutions (e.g., at any position in the sugar (e.g., ribose) – e.g., at position 2') include, for example, but not limited to, substitutions via any of the following: O[(CH2)] n O]m, CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON[(CH2) n CH3)]2, where n and m are 1 to about 10.

[0232] In some embodiments, the modified ribose comprises any one or more of the following modifications: 2'-O-methyl (2'-OMe); 2'-O-methoxyethyl (2'-O-MOE); 2'-deoxy-2'-fluorine (2'-F); 2'-arabinose-fluorine (2'-Ara-F); 2'-O-benzyl; 2'-O-methyl-4-pyridine (2'-O-methyl-4-pyridine (2'-O-CH2Py(4)); 2'F-4'-Cα-OMe; or 2',4'-di-Cα-OMe.

[0233] In some embodiments, the gRNA contains any of the following substituents at the 2'-position of the sugar (e.g., ribose): C1 to C2. 10 Lower alkyl, substituted lower alkyl, alkylaryl, aralkyl, O-alkylaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocyclic alkyl, heterocyclic alkylaryl, aminoalkylamino, polyalkylamino or substituted silyl. In some embodiments, gRNA comprises 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl) or 2'-MOE) (see, for example, Martin et al., Helv. Chim. Acta [Swiss Chimica Journal], 1995, 78:486-504, the entire contents of which are incorporated herein by reference for all purposes) (i.e., alkoxy-alkoxy). In some embodiments, the gRNA comprises a 2'-dimethylaminoethoxy group, i.e., an O(CH2)2ON(CH3)2 group, also known as 2'-DMAOE; a 2'-dimethylaminoethoxyethoxy group (also known in the art as 2'-O-dimethylaminoethoxyethyl or 2'-DMAEOE), i.e., 2'-O-CH2-O-CH2-N(CH3)2; a 5'-Me-2'-F nucleotide, a 5'-Me-2'-OMe nucleotide, or a 5'-Me-2'-deoxynucleotide (both the R and S isomers of these three families); a 2'-alkoxyalkyl group; and 2'-NMA (N-methylacetamide). • Non-bicyclic sugar modification

[0234] In some embodiments, the modified sugar (e.g., ribose) moiety comprises a non-bicyclic modified sugar (e.g., ribose) moiety. In some embodiments, the modified sugar (e.g., ribose) moiety comprises a furanyl ring containing one or more substituents, none of which bridge the two atoms of the furanyl ring to form a bicyclic structure. In some embodiments, one or more unbridging substituents of the non-bicyclic modified ribose moiety are branched. Such unbridging substituents can be at any position on the furanyl ring, including but not limited to substituents at the 2', 4', and / or 5' positions.

[0235] In some embodiments, the non-bicyclic modified sugar (e.g., ribose) portion contains a substituent at the 2'-position of the sugar (e.g., ribose). Examples of suitable 2'-substituents for non-bicyclic modified ribose moieties include, but are not limited to: 2'-O-methyl (2'-OMe), 2'-O-methoxyethyl (2'-O-MOE), 2'-deoxy-2'-fluoro (2'-F), 2'-arabinose-fluorine (2'-Ara-F), 2'-O-benzyl, 2'-O-methyl-4-pyridine (2'-O-methyl-4-pyridine (2'-O-CH2Py(4))) and 2'-ON-alkylacetamides (e.g., 2'-ON-methylacetamide (“NMA”), 2'-ON-dimethylacetamide, 2'-ON-ethylacetamide, and 2'-ON-propylacetamide). See, for example, US 6,147,200, Prakash et al., 2003, Org. Lett. [Organic Chemistry Communications], 5, 403-6, the entire contents of which are incorporated herein by reference for all purposes.

[0236] In some embodiments, the 2'-substituent is halogenated, allyl, amino, azide, SH, CN, OCN, CF3, OCF3, or O-C1-C. 10 Alkoxy, O-C1-C 10 Substituted alkoxy groups, O-C1-C 10 Alkyl, O-C1-C 10 Substituted alkyl, S-alkyl, N(R) m )-alkyl, O-alkenyl, S-alkenyl, N(R m )-Alkenyl, O-alkynyl, S-alkynyl, N(R m )-alkynyl, O-alkylene-O-alkyl, alkynyl, alkylaryl, aralkyl, O-alkylaryl, O-aralkylyl, O(CH2)2SCH3, O(CH2)2ON(Rm)(Rn) or OCH2C(=O)-N(Rm)(Rn), wherein each R m and R n It is independently of H, amino protecting group, or substituted or unsubstituted C1-C10 Alkyl, or a 2'-substituent described in any of the following: Cook et al., US 6,531,584; Cook et al., US 5,859,221; and Cook et al., US 6,005,087, the entire contents of which are incorporated herein by reference for all purposes. In some embodiments, these 2'-substituents may be further substituted by one or more substituents independently selected from: hydroxyl, amino, alkoxy, carboxyl, benzyl, phenyl, nitro (NO2), thiol, thioalkoxy, thioalkyl, halogen, alkyl, aryl, alkenyl, and alkynyl.

[0237] In some embodiments, the 2'-substituted non-bicyclic modified nucleoside comprises a sugar (e.g., ribose) moiety containing a non-bridging 2'-substituent selected from the following: F, NH2, N3, OCF3, OCH3, O(CH2)3NH2, CH2CH=CH2, OCH2CH=CH2, OCH2CH2OCH3, O(CH2)2SCH3, O(CH2)2ON(R) m (R) n O(CH2)2O(CH2)2N(CH3)2 and N-substituted acetamides (OCH2C(=O)-N(R) m (R) n ), where each R m and R n It is independently of H, amino protecting group, or substituted or unsubstituted C1-C 10 Alkyl group. In some embodiments, the 2'-substituted non-bicyclic modified nucleoside comprises a sugar (e.g., ribose) moiety containing a non-bridging 2'-substituent selected from the following: F, OCF, OCH3, OCH2CH2OCH3, O(CH2)2SCH3, O(CH2)2ON(CH3)2, O(CH2)2O(CH2)2N(CH3)2, and OCH2C(=O)-N(H)CH3 (“NMA”). In some embodiments, the 2'-substituted non-bicyclic modified nucleoside comprises a sugar (e.g., ribose) moiety containing a non-bridging 2'-substituent selected from the following: F, OCH3, OCH2CH2OCH3, and OCH2C(=O)-N(H)CH3.

[0238] In some embodiments, the non-bicyclic modified sugar (e.g., ribose) moiety includes a substituent at the 3'-position of the sugar (e.g., ribose). Examples of suitable substituents for the 3'-position of the modified sugar (e.g., ribose) moiety include, but are not limited to, alkoxy (e.g., methoxy) and alkyl (e.g., methyl, ethyl).

[0239] In some embodiments, the non-bicyclic modified sugar (e.g., ribose) moiety includes a substituent at the 4'-position of the sugar (e.g., ribose). Examples of suitable substituents for the 4'-position of the non-bicyclic modified sugar (e.g., ribose) moiety include, but are not limited to, alkoxy (e.g., methoxy), alkyl, and those described in Manoharan et al., WO 2015 / 106128.

[0240] In some embodiments, the non-bicyclic modified sugar (e.g., ribose) moiety includes a substituent at the 5'-position of the sugar (e.g., ribose). Examples of suitable substituents for the 5'-position of the modified sugar (e.g., ribose) moiety include, but are not limited to, vinyl (e.g., 5'-vinyl), alkoxy (e.g., methoxy (e.g., 5'-methoxy)), and alkyl (e.g., methyl (R or S) (e.g., 5'-methyl (R or S)), ethyl).

[0241] In some embodiments, the non-bicyclic modified sugar (e.g., ribose) moiety comprises more than one non-bridging sugar substituent, such as a 2'-F-5'-methyl sugar (e.g., ribose) moiety, and modified sugar (e.g., ribose) moiety and modified nucleoside as described in Migawa et al., WO 2008 / 101157 and Rajeev et al., US 2013 / 0203836, the entire contents of which are incorporated herein by reference for all purposes.

[0242] In some embodiments, the modified furanyl sugar (e.g., ribose) moiety and the nucleoside incorporated into such modified furanyl sugar (e.g., ribose) moiety are further defined by isomer configuration. For example, the 2'-deoxyfuranyl sugar (e.g., ribose) moiety can be in seven isomer configurations other than the naturally occurring β-D-deoxyribosyl configuration. Such modified sugar (e.g., ribose) moiety is described, for example, in WO 2019 / 157531, the entire contents of which are incorporated herein by reference for all purposes.

[0243] In some embodiments, the sugar (e.g., ribose) modification comprises an unlocking nucleotide (UNA). An UNA is an unlocking acyclic nucleic acid in which any bond of the sugar has been removed, forming an unlocking sugar (e.g., ribose) residue. For example, in some embodiments, the bond between C1' and C4' (i.e., the covalent carbon-oxygen-carbon bond between the C1' and C4' carbons) has been removed. In some embodiments, the C2'-C3' bond of the sugar (e.g., ribose) (i.e., the covalent carbon-carbon bond between the C2' and C3' carbons) has been removed. See, for example, Nuc. Acids Symp. Series, 52, 133-134 (2008) and Fluiter et al., Mol. Biosyst., 2009, 10, 1039, the entire contents of which are incorporated herein by reference. UNA and methods of preparation are known in the art. See, for example, U.S. Patent No. 8,314,227; and US 2013 / 0096289; US 2013 / 0011922; and US 2011 / 0313020, the entire contents of which are hereby incorporated herein by reference. • Bicyclic sugar modification

[0244] In some embodiments, the modified sugar (e.g., ribose) moiety comprises a substituent that bridges two atoms of the furanose ring to form a second ring, thereby producing a bicyclic sugar (e.g., ribose) moiety. In some embodiments, the bicyclic sugar (e.g., ribose) moiety comprises a bridge between the 4' and 2' furanose ring atoms. Examples of such 4' to 2' bridging sugar substituents include, but are not limited to: 4'-CH2-2', 4'-(CH2)2-2', 4'-(CH2)3-2', 4'-CH2-O-2' (“LNA”), 4'-CH2-S-2', 4'-(CH2)2-O-2' (“ENA”), 4'-CH(CH3)-O-2' (referred to as “restricted ethyl” or “cEt”), 4'-CH2-O-CH2-2', 4'-CH2-N(R)-2', 4'-CH(CH2OCH3)-O-2' (“restricted MOE” or “cMOE”) and their analogues (see, for example, Seth et al., US 7,399,845; Bhat et al., US 7,569,686; Swayze et al., US 7,741,457; and Swayze et al., US 7,741,457). 8,022,193), 4'-C(CH3)(CH3)-O-2' and its analogues (see, e.g., Seth et al., US 8,278,283), 4'-CH2-N(OCH3)-2' and its analogues (see, e.g., Prakash et al., US 8,278,425), 4'-CH2-ON(CH3)-2' (see, e.g., Allenson et al., US 7,696,345 and Allenson et al., US 8,124,745), 4'-CH2-C(H)(CH3)-2' (see, e.g., Zhou et al., J.Org. Chem. [Journal of Organic Chemistry], 2QQ9, 74, 118-134), 4'-CH2-C(=CH2)-2' and its analogues (see, e.g., Seth et al., US 8,278,426), 4'-C(R a R b )-N(R)-O-2'、4'-C(R a R b )-ON(R)-2', 4'-CH2-ON(R)-2' and 4'-CH2-N(R)-O-2', wherein each R, R a and R b Independently, it is H, protecting group, or C1-C 12Alkyl groups (see, for example, Imanishi et al., US 7,427,672). The entire contents of all the foregoing references are incorporated herein by reference for all purposes.

[0245] In some embodiments, such 4' to 2' bridges independently comprise 1 to 4 linked groups independently selected from: -[C(R a (R) b )]n-、-[C(R a (R) b )]nO-、-C(R a )=C(R b )-、-C(R a )=N-、-C(=NR a )-, -C(=O)-, -C(=S)-, -O-, -Si(R a )2-, -S(=O)X- and -N(R a -; where: x is 0, 1, or 2; n is 1, 2, 3, or 4; each R a and R b Independently, it is H, protecting group, hydroxyl group, C1-C 12 Alkyl, substituted C1-C 12 Alkyl, C2-C 12 Alkenyl, substituted C2-C 12 alkenyl, C2-C 12 Alkyne group, substituted C2-C 12 alkynyl group, C5-C 20 Aryl, substituted C5-C 20 Aryl, heterocyclic radical, substituted heterocyclic radical, heteroaryl, substituted heteroaryl, C5-C7 alicyclic radical, substituted C5-C7 alicyclic radical, halogen, OJ1, NJ1J2, SJ1, N3, COOJ1, acyl (C(=O)-H), substituted acyl, CN, sulfonyl (S(=O)2-J1) or sulfoxyl (S(=O)-J1); and each J1 and J2 is independently H, C1-C 12 Alkyl, substituted C1-C 12 Alkyl, C2-C 12 Alkenyl, substituted C2-C 12 alkenyl, C2-C 12 Alkyne group, substituted C2-C 12 alkynyl group, C5-C 20 Aryl, substituted C5-C 20 Aryl, acyl (C(=O)-H), substituted acyl, heterocyclic radical, substituted heterocyclic radical, C1-C 12Aminoalkyl, substituted C1-C 12 Aminoalkyl or protecting group.

[0246] The other bicyclic sugar moiety is known in the art; see, for example: Freier et al., Nucleic Acids Research, 1997, 25(22), 4429-4443; Albaek et al., J. Org. Chem., 2006, 71, 7731-7740; Singh et al., Chem. Commun., 1998, 4, 455-456; Koshkin et al., Tetrahedron, 1998, 54, 3607-3630; Kumar et al., Bioorg. Med. Chem. Lett., 1998, 8, 2219-2222; Singh et al., J. Org. Chem., 1998, 63, 10035-10039; Srivastava et al., J. Am. Chem. Soc. [Journal of the American Chemical Society], 2007, 129, 8362-8379; Wengel et al., US 7,053,207; Imanshi et al., US 6,268,490; Imanshi et al., US 6,770,748; Imanshi et al., US RE44,779; Wengel et al., US 6,794,499; Wengel et al., US 6,670,461; Wengel et al., US 7,034,133; Wengel et al., US 8,080,644; Wengel et al., US 8,034,909; Wengel et al., US 8,153,365; Wengel et al., US 7,572,582; Ramasamy et al., US 6,525,191; Torsten et al., WO 2004 / 106356; Wengel et al., WO 1999 / 014226; Seth et al., WO2007 / 134181; Seth et al., US 7,547,684; Seth et al., US 7,666,854; Seth et al., US 8,088,746; Seth et al., US 7,750,131; Seth et al., US 8,030,467; Seth et al., US 8,268,980; Seth et al., US 8,546,556; Seth et al., US 8,530,640; Migawa et al., US9,012,421; Seth et al., US 8,501,805; and Allenson et al., US 2008 / 0039618 and Migawa et al., US 2015 / 0191727. The entire contents of all the foregoing references are incorporated herein by reference for all purposes.

[0247] In some embodiments, the modified sugar (e.g., ribose) comprises a restricted ethyl nucleotide containing a 4'-CH(CH3)-O-2' bridge. In some embodiments, the restricted ethyl nucleotide is in an S conformation (S-cEt). In some embodiments, the modified sugar (e.g., ribose) comprises a conformation-restricted nucleotide (CRN). A CRN is a nucleotide analog having a linker connecting the C2' and C4' carbons of the ribose or the C3 and -C5' carbons of the ribose. Representative disclosures teaching the preparation of certain CRNs described above include, but are not limited to, US 2013 / 0190383 and WO 2013 / 036868, the entire contents of which are hereby incorporated herein by reference.

[0248] In some embodiments, the bicyclic sugar moiety and the nucleotide incorporated into such bicyclic sugar moiety are further defined by isomer configuration. For example, LNA nucleotides (described herein) can be in the α-L configuration or the β-D configuration. In this document, the general description of bicyclic nucleotides includes both isomer configurations. Any of the aforementioned bicyclic nucleotides having one or more stereochemical sugar configurations can be prepared, including, for example, α-L-furanose and β-D-furanose (see, for example, WO 99 / 14226, the entire contents of which are incorporated herein by reference for all purposes).

[0249] Other representative U.S. patents and U.S. patent disclosures teaching the preparation of bicyclic nucleotides (e.g., locked nucleic acids) include, but are not limited to, the following: U.S. Patent Nos. 6,268,490, 6,525,191, 6,670,461, 6,770,748, 6,794,499, 6,998,484, 7,053,207, 7,034,133, 7,084,125, 7,399,845, 7,427,672, 7,569,686, 7,741,457, 8,022,193, 8,030,467, 8,278,425, 8,278,426, and 8,278,283. The entire contents of US2008 / 0039618 and US2009 / 0012281 are hereby incorporated herein by reference. (b) Nucleobase modification

[0250] In some embodiments, the modified gRNA comprises one or more nucleotides containing modified nucleobases.

[0251] As used herein, “unmodified” nucleobases refer to the purine bases adenine (A) and guanine (G), and the pyrimidine bases thymine (T), cytosine (C), and uracil (U). Modified nucleobases include other synthetic and natural nucleobases.

[0252] Modified nucleobases include, but are not limited to, 5-substituted pyrimidines, 6-azapyrimidines, pyrimidines substituted with alkyl or alkynyl groups, purines substituted with alkyl groups, and purines substituted with N-2, N-6, and O-6 groups. In some embodiments, the modified nucleobases are selected from: 5-methylcytosine, 2-aminopropyladenine, 5-hydroxymethylcytosine, xanthine, hypoxanthine, deoxythymidine (dT), 2-aminoadenine, 6-N-methylguanine, 6-N-methyladenine, 2-propyladenine, 2-thiouracil, 2-thiothymidine and 2-thiocytosine, 5-propynyl (-C=C-CH3)uracil, 5-propynylcytosine, 6-azauracil, 6-azacytosine, 6-azathymidine, 5-ribosyluracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-thiol, 8- Thioalkyl, 8-hydroxy, 8-aza and other 8-substituted purines, 5-halogenated (especially 5-bromo), 5-trifluoromethyl, 5-halogenated uracil and 5-halogenated cytosine, 7-methylguanine, 7-methyladenine, 2-F-adenine, 2-aminoadenine, 7-deadenine, 7-deadenine, 3-deadenine, 3-deadenine, 6-N-benzoyladenine, 2-N-isobutyrylguanine, 4-N-benzoylcytosine, 4-N-benzoyluracil, 5-methyl4-N-benzoylcytosine, 5-methyl4-N-benzoyluracil, universal base, hydrophobic base, promiscuous base, size-expanded base and fluorinated base. Other modified nucleobases include tricyclic pyrimidines, such as 1,3-diazaphenoxazin-2-one, 1,3-diazaphenthiazin-2-one, and 9-(2-aminoethoxy)-1,3-diazaphenoxazin-2-one (G-clamp). Modified nucleobases may also include nucleobases in which the purine or pyrimidine base is substituted by another heterocyclic ring, for example, 7-deadenine, 7-deadenanine, 2-aminopyridine, and 2-pyridone. Other nucleobases include those disclosed in Merigan et al., US 3,687,808; The Concise Encyclopedia of Polymer Science and Engineering, edited by Kroschwitz, JI, John Wiley & Sons, 1990, 858-859; Englisch et al., Angewandte Chemie, International Edition, 1991, 30, 613, the entire contents of which are incorporated herein by reference for all purposes.

[0253] In some embodiments, the modified nucleobases include pseudouridine, 2'-thiouridine (s2U), N6'-methyladenosine, and 5'-methylcytidine (m). 5 C), 5'-fluoro-2'-deoxyuridine, N-ethylpiperidine-7-EAA-triazole modified adenine, N-ethylpiperidine-6'-triazole modified adenine, 6-phenylpyrrolo-cytosine (PhpC), 2',4'-difluorotoluylribonucleoside (rF), or 5'-nitroindole. In some embodiments, the modified nucleobases comprise 5-substituted pyrimidines; 6-azapyrimidines; or N-2, N-6, and O-6 substituted purines (including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine). It has been shown that 5-methylcytosine substitution increases the stability of nucleic acid duplexes by 0.6°C–1.2°C (Sanghvi, YS, Crooke, ST, and Lebleu, B., eds., dsRNA Research and Applications, CRC Press, Boca Raton, 1993, pp. 276–278), and is an exemplary base substitution, especially when combined with 2'-O-methoxyethyl sugar modification.

[0254] Representative U.S. patents and published applications teaching the preparation of certain nucleosides in the modified nucleosides described above, as well as other modified nucleosides, include, but are not limited to, U.S. patent numbers 3,687,808, 4,845,205, 5,130,30, 5,134,066, 5,175,273, 5,367,066, 5,432,272, 5,457,187, 5,459,255, 5,484,908, 5,502,177, 5,525,711, 5,552,540, 5,587,469, 5,594,121, 5,596,091, 5,614,617, 5,681,941, and 5,750. 692, 6,015,886, 6,147,200, 6,166,197, 6,222,025, 6,235,887, 6,380,368, 6,528,640, 6,639,062, 6,617,438, 7,045,610, 7,427,672, 7,495,088 ,5,130,302, 5,134,066, 5,175,273, 5,367,066, 5,432,272, 5,434,257, 5,457,187, 5,459,255, 5,484,908, 5,502,177, 5,525,711, 5,552,540, US The contents of 5,587,469, 5,594,121, 5,596,091, 5,614,617, 5,645,985, 5,681,941, 5,811,534, 5,750,692, 5,948,903, 5,587,470, 5,457,191, 5,763,588, 5,830,653, 5,808,027, 6,166,199 and 6,005,096, in their entirety, are hereby incorporated herein by reference for all purposes. (c) Nucleoside interlinking modification

[0255] In some embodiments, the modified gRNA comprises one or more modified nucleoside-to-nucleotide linkages. Compared to naturally occurring phosphate-ester linkages, modified nucleoside-to-nucleotide linkages can be used to alter (typically increase) the nuclease resistance of agents (e.g., as described herein).

[0256] Naturally occurring nucleoside linkages between RNA and DNA are 3' to 5' phosphodiester linkages. In some embodiments, modified nucleoside linkages contain normal 3'-5' linkages. In some embodiments, modified nucleoside linkages contain 2'-5' linkages. In some embodiments, modified nucleoside linkages have inverted polarity, wherein adjacent nucleoside unit pairs are linked, for example, 3'-5' to 5'-3' or 2'-5' to 5'-2'.

[0257] The two main categories of modified nucleoside linkages can be defined by the presence or absence of a phosphorus atom. • Modified phosphorus-containing nucleotide interlinking

[0258] In some embodiments, the modified nucleoside-to-nucleotide linkages contain a phosphorus atom. Representative modified phosphorus-containing nucleoside-to-nucleotide linkages include, but are not limited to, thiophosphates (PS (Rp isomers or Sp isomers)) (e.g., 5'-thiophosphates) (e.g., chiral thiophosphates), phosphate triesters, aminophosphates (e.g., 3'-aminoaminophosphates and aminoalkylaminophosphates), chiral thiophosphates, dithiophosphates (PS2), aminoalkyl phosphate triesters, methyl and other alkyl phosphonates (e.g., methylphosphonates (MP), 3'-alkylenephosphonates), methoxypropylphosphonates (MOP), 5'-(E)vinylphosphonates, 5'-methylphosphonates, (S)-5'-C-methylphosphonates, hypophosphonates, thiocarbonylaminophosphates, thiocarbonylalkylphosphonates, thiocarbonylalkyl phosphate triesters, boranophosphates, hypophosphonates, and peptide nucleic acids (PNAs).

[0259] Methods for preparing polynucleotides containing one or more modified phosphorus-containing nucleosides linked together are known in the art. See, for example, U.S. Patent Nos. 3,687,808, 4,469,863, 4,476,301, 5,023,243, 5,177,195, 5,188,897, 5,264,423, 5,276,019, 5,278,302, 5,286,717, 5,321,131, and 5,399. 676, 5,405,939, 5,453,496, 5,455,233, 5,466,677, 5,476,925, 5,519,126, 5,536,821, 5,541,316, 5,550,111, 5,563,253, 5,571,799, 5,587,361, 5,6 U.S. Patent RE39464, the entire contents of which are hereby incorporated herein by reference for all purposes. • Modified non-phosphorus nucleotide interlinking

[0260] In some embodiments, the modified nucleoside interlinks do not contain phosphorus atoms. The modified nucleoside interlinks that do not contain phosphorus atoms have a backbone formed by short-chain alkyl or cycloalkyl nucleoside interlinks, mixed heteroatom and alkyl or cycloalkyl nucleoside interlinks, or one or more short-chain heteroatom nucleoside interlinks or heterocyclic nucleoside interlinks. These include those with morpholino groups (partially formed from the sugar moiety of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formyl and thioformyl backbones; methyleneformyl and thioformyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and other backbones having mixed N, O, S, and CH2 components.

[0261] Representative non-phosphorus nucleoside linkers include, but are not limited to, methylene methyl imino (-CH2-N(CH3)-O-CH2-), thiodiester, thiocarbonyl carbamate (-OC(=O)(NH)-S-); siloxane (-O-SiH2-O-); and N,N′-dimethylhydrazine (-CH2-N(CH3)-N(CH3)-).

[0262] Methods for preparing modified nucleoside-linked polynucleotides containing phosphorus atoms are known in the art. See, for example, U.S. Patent Nos. 5,034,506, 5,166,315, 5,185,444, 5,214,134, 5,216,141, 5,235,033, 5,64,562, 5,264,564, 5,405,938, 5,434,257, 5,466,677, 5,470,967, and 5,489,677. The contents of 5,541,307, 5,561,225, 5,596,086, 5,602,240, 5,608,046, 5,610,289, 5,618,704, 5,623,070, 5,663,312, 5,633,360, 5,677,437 and 5,677,439, in their entirety, are hereby incorporated herein by reference. (d) Exemplary combinations of modifications

[0263] As described above, the exemplary modifications listed can be used in any non-exclusive combination. For example, exemplary combinations of modifications include 2'-O-Me 3'-thiophosphate (MS) nucleotide; 2'-O-MOE 3'-thiophosphate nucleotide; 2'-F 3'-thiophosphate nucleotide; 2'-O-Me 3'-thioPACE (MSP) nucleotide; and 2'-deoxy 3'-thiophosphate nucleotide. (ii) Modified position

[0264] The modified nucleotides can be located at any suitable position in the whole gRNA (e.g., the ends of the full-length gRNA (e.g., 5' end, 3' end, or 5' and 3' end residues); any domain of the gRNA (e.g., sgRNA or crRNA or tracrRNA of the template RNA); internal residues of the full-length gRNA; etc.).

[0265] In some embodiments, the ends of the gRNA (e.g., the 5' end, the 3' end, or both 5' and 3' end residues) are modified. In some embodiments, modification of the terminal residues reduces the degradation of the gRNA by exonucleases (e.g., in cells). In some embodiments, modification of the terminal residues increases the stability of the gRNA (e.g., in cells, in vitro, or in vivo). In some embodiments, the 5' end of the gRNA contains one or more modified nucleotides. In some embodiments, 1, 2, 3, 4, or 5 nucleotides at the 5' end are modified. In some embodiments, the 3' end of the gRNA contains one or more modified nucleotides. In some embodiments, 1, 2, 3, 4, or 5 nucleotides at the 3' end are modified. In some embodiments, both the 3' and 5' ends of the gRNA contain one or more modified nucleotides. In some embodiments, 1, 2, 3, 4, or 5 nucleotides at the 3' end are modified, and 1, 2, 3, 4, or 5 nucleotides at the 5' end are modified.

[0266] In some embodiments, one or more internal (i.e., non-terminal) nucleotides of the gRNA are modified. In some embodiments, modification of internal residues reduces the degradation of the gRNA by endonucleases (e.g., in cells). In some embodiments, modification of internal residues increases the stability of the gRNA (e.g., in cells, in vitro, or in vivo). In some embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more of the internal nucleotides of the gRNA are modified.

[0267] In some embodiments, one or more nucleotides of the crRNA (e.g., the crRNA of the template RNA's sgRNA) are modified. In some embodiments, one or more nucleotides of the seed region, PAM distal region, and / or tracrRNA binding region of the crRNA (e.g., the crRNA of the template RNA's sgRNA) are modified. In some embodiments, the 3' and / or 5' nucleotides of the crRNA are modified. In some embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides of the crRNA (e.g., the crRNA of the template RNA's sgRNA) are modified. In some embodiments, one or more nucleotides of the tracrRNA (e.g., the tracrRNA of the template RNA's sgRNA) are modified. In some embodiments, one or more nucleotides of the tracrRNA (e.g., the tracrRNA of the template RNA's sgRNA) that do not interact with a Cas endonuclease (e.g., the Cas endonuclease described herein) are modified. 4.5.2.3 Methods for preparing gRNA

[0268] gRNA can be generated according to standard nucleic acid synthesis methods known in the art as described herein (see, for example, §4.6).

[0269] The generation of multi-domain gRNAs (e.g., sgRNAs, template gRNAs) can be achieved by assembling two or more (e.g., two, three, four, five, six, seven, eight, nine, ten, or more) RNA segments together. For example, these gRNAs can be generated by contacting two or more linear RNA segments together, allowing the 5' end of the first RNA segment to be covalently linked to the 3' end of the second RNA segment. A linker molecule can be contacted with a third RNA segment, allowing the 5' end of the linker molecule to be covalently linked to the 3' end of the third RNA segment. The method may further include linking a fourth, fifth, or additional RNA segment to an elongated molecule. In some cases, this form of assembly can allow for rapid and efficient assembly of gRNA molecules (e.g., multi-domain gRNAs (e.g., sgRNAs, template gRNAs)). See, for example, US 20160102322 A1 (e.g., Figure 10) and WO 2021178720, the entire contents of which are incorporated herein by reference for all purposes.

[0270] In some embodiments, RNA segments can be generated through chemical synthesis. In some embodiments, RNA segments can be generated through in vitro transcription of a nucleic acid template, for example by providing an RNA polymerase that acts on a homologous promoter of a DNA template to generate RNA transcripts. In some embodiments, in vitro transcription is performed using, for example, T7, T3, or SP6 RNA polymerases or derivatives thereof, which act on DNA (e.g., dsDNA, ssDNA, linear DNA, plasmid DNA, linear DNA amplicon, linearized plasmid DNA) that, for example, encodes an RNA segment, for example, under the transcriptional control of a homologous promoter (e.g., T7, T3, or SP6 promoter). In some embodiments, a combination of chemical synthesis and in vitro transcription is used to generate RNA segments for assembly. In some embodiments, in vitro transcription may be more suitable for generating longer RNA molecules (compared to chemical synthesis). In some embodiments, the reaction temperature for in vitro transcription can be reduced, for example, below 37°C (e.g., between 0°C-10°C, 10°C-20°C, or 20°C-30°C) to obtain a higher proportion of full-length transcripts (Krieg Nucleic Acids Res [Nucleic Acids Research] 18:6463 (1990)). In some embodiments, long template RNAs, such as template RNAs greater than 5 kb, are synthesized using protocols that improve long transcript synthesis, such as using T7 RiboMAX Express, thereby generating 27 kb transcripts in vitro (see, for example, Thiel et al. J Gen Virol [Journal of General Virology] 82(6):1273-1281 (2001), the entire contents of which are incorporated herein by reference for all purposes). In some embodiments, modifications to the RNA molecule as described herein can be incorporated during RNA segment synthesis (e.g., by including modified nucleotides or alternative binding chemicals), after the synthesis of RNA segments by chemical or enzymatic processes, after the assembly of one or more RNA segments, or a combination thereof.

[0271] Other exemplary methods that can be used to ligate RNA segments are click chemistry, for example, as described in US7375234, US 7070941, US 20130046084, and US 20160102322 A (the entire contents of each of these are incorporated herein by reference for all purposes). Any click reaction can be used to ligate RNA segments (e.g., Cu-azide-alkyne, strain-promoted azide-alkyne, Staudinger ligation, tetrazine ligation, photoinduced tetrazolium-olefin, thiol-olefin, NHS ester, epoxide, isocyanate, and aldehyde-aminooxy). In some embodiments, using click chemistry to ligate RNA molecules is advantageous because click chemistry is rapid, modular, efficient, generally does not produce toxic waste products, can be carried out using water as a solvent, and / or can be configured to have stereospecificity. 4.5.3 Systemic nucleic acid editing activity

[0272] As described above, the system described herein can be used in particular for editing (e.g., adding, deleting, or substituting one or more nucleotides) target nucleic acid molecules (e.g., DNA, genome, genes (e.g., within cells, such as within the cells of a subject (e.g., mammalian subjects, such as human subjects)) (e.g., in vivo, in vitro, or ex vivo).

[0273] In some embodiments, the system (e.g., the system described herein containing the Cas endonuclease) exhibits increased editing efficiency relative to a reference system containing a reference Cas endonuclease. In some embodiments, the system (e.g., the system described herein containing the Cas endonuclease) exhibits an increase in editing efficiency of at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more relative to a reference system containing a reference Cas endonuclease. In some embodiments, the system (e.g., the system described herein containing the Cas endonuclease) exhibits an increase in editing efficiency of at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, or more relative to a reference system containing a reference Cas endonuclease. In some embodiments, the systems described herein (e.g., systems described herein containing the Cas endonuclease described herein) exhibit an increase in editing efficiency of approximately 30%-200%, 40%-200%, 50%-200%, 60%-200%, 70%-200%, 80%-200%, 90%-200%, 100%-200%, 150%-200%, 30%-150%, 40%-150%, 50%-150%, 60%-150%, 70%-150%, 80%-150%, 90%-150%, 100%-150%, 30%-100%, 40%-100%, 50%-100%, 60%-100%, 70%-100%, 80%-100%, or 90%-100% or more, relative to the editing efficiency of a reference system containing a reference Cas endonuclease.

[0274] In some embodiments, the system (e.g., the system described herein containing the Cas endonuclease) exhibits increased editing efficiency relative to the editing efficiency of the reference Cas endonuclease shown in SEQ ID NO: 321. In some embodiments, the system (e.g., the system described herein containing the Cas endonuclease) exhibits an increase in editing efficiency of at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, or more. In some embodiments, the system (e.g., the system described herein containing the Cas endonuclease) exhibits an increase in editing efficiency of at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, or more, relative to the editing efficiency of the reference Cas endonuclease shown in SEQ ID NO: 321. In some embodiments, relative to the editing efficiency of the system containing SEQ ID NO: The editing efficiency of the reference Cas endonuclease system shown in 321 (e.g., the system described herein containing the Cas endonuclease described herein) exhibits an increase in editing efficiency of approximately 30%-200%, 40%-200%, 50%-200%, 60%-200%, 70%-200%, 80%-200%, 90%-200%, 100%-200%, 150%-200%, 30%-150%, 40%-150%, 50%-150%, 60%-150%, 70%-150%, 80%-150%, 90%-150%, 100%-150%, 30%-100%, 40%-100%, 50%-100%, 60%-100%, 70%-100%, 80%-100%, or 90%-100% or more. 4.5.4 Methods for evaluating the nucleic acid editing activity of the system

[0275] Standard methods for evaluating the systematic editing of target nucleic acid molecules (e.g., in cells) described herein are known in the art and are described herein. See, for example, Maja Gehre et al., Efficient strategies to detect genome editing and integrity in CRISPR-Cas9 engineered ESCs, bioRxiv 635151; doi:https: / / doi.org / 10.1101 / 635151. Glaser A, McColl B, Vadolas J. GFP to BFP Conversion: A Versatile Assay for the Quantification of CRISPR / Cas9-mediated Genome Editing. [Published corrections can be found in Mol Ther Nucleic Acids. Sep 13, 2016; 5(9):e360]. Mol Ther Nucleic Acids. 2016;5(7):e334. Published on July 12, 2016. doi:10.1038 / mtna.2016.48, the entire contents of which are incorporated herein by reference for all purposes. For example, standard nucleic acid sequencing methods (e.g., next-generation sequencing, Sanger sequencing), assessment of phenotypes associated with specific target editing, mismatch detection assays, or restriction fragment length polymorphism assays.

[0276] For example, to monitor gene editing of target DNA, mammalian cells carrying the target DNA, such as HEK293T or U2OS cells, can be used. In other embodiments for monitoring gene editing of target DNA, mammalian cells carrying a target DNA genome landing pad, such as HEK293T or U2OS cells, can be used. In certain embodiments, the target DNA genome landing pad may contain a gene to be edited to treat a target disease or disorder. In other specific embodiments, the target DNA is a gene sequence expressing a protein exhibiting detectable characteristics that can be monitored to determine whether gene editing has occurred. For example, in some embodiments, a genome landing pad expressing blue fluorescent protein (BFP) or green fluorescent protein (GFP) is used. In some embodiments, mammalian cells (e.g., HEK293T or U2OS cells) containing the target DNA (e.g., a target DNA genome landing pad) are seeded in culture plates at 500x-3000x cells per editing system and transduced at a multiplicity of infection (MOI) of 0.2-0.3 to minimize multiple infection per cell. Puromycin (2.5 ug / mL) can be added 48 hours post-infection to select infected cells. In such an embodiment, cells can be held under puromycin selection for at least 7 days and then scaled up to introduce gRNA (e.g., template RNA) (e.g., electroporation, e.g., template RNA electroporation).

[0277] To determine whether gene editing has occurred, mammalian cells containing the target DNA to be edited can be infected with a candidate endonuclease (or its fusion protein, such as a reverse transcriptase-based fusion protein), followed by transfection with a guide RNA (e.g., template RNA) designed to edit the target DNA. The cells can then be analyzed to determine whether the editing of the target DNA occurred as designed, or whether no editing occurred or imperfect editing occurred, for example, through cell sorting and sequence analysis.

[0278] In a particular embodiment, to determine whether gene editing has occurred, mammalian cells expressing BFP or GFP (e.g., HEK293T or U2OS cells) can be infected with a candidate endonuclease (or its fusion protein, e.g., a reverse transcriptase-based fusion protein) followed by transfection or electroporation with a guide RNA plasmid or RNA (e.g., a template RNA plasmid or RNA), for example, by electroporating approximately 250,000 cells / well with 200 ng of a guide RNA plasmid or RNA (e.g., a template RNA plasmid or RNA) designed to convert BFP to GFP or GFP to BFP, ensuring a cell count with >250x-1000x coverage for each candidate. In such an embodiment, the gene editing capability of various constructs in the assay can be assessed by sorting cells using fluorescence-activated cell sorting (FACS) for the expression of the color-converted fluorescent protein (FP) 4–10 days after electroporation. Cells are sorted and harvested into different populations: unedited cells (exhibiting the original fluorescent protein signal), edited cells (exhibiting the transformed fluorescent protein signal), and imperfectly edited cells (not exhibiting the fluorescent protein signal). Unsorted cell samples can also be harvested as an input population to identify candidate enrichments during analysis. Targeted editing sites can also be analyzed using standard sequencing methods (e.g., next-generation sequencing). 4.5.5 Exemplary System

[0279] The following provides exemplary systems incorporating the components described above. Exemplary systems include exemplary homology-directed repair (HDR) based editing systems; reverse transcriptase based editing systems; and nucleobase editor based editing systems. These systems are exemplary and not intended to be limiting. 4.5.5.1 HDR-based editing system

[0280] This article provides, in particular, HDR-based systems for use in editing target nucleic acid molecules (e.g., in cells, for example, in the body of a subject). In some embodiments, the system comprises (a) (i) a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) as described herein; (ii) a fusion protein comprising a Cas endonuclease (or a functional fragment, functional variant thereof) as described herein (e.g., as described herein); (iii) a conjugate comprising a Cas endonuclease (or a functional fragment, functional variant thereof) as described herein (e.g., as described herein); (iv) a nucleic acid molecule encoding (a)(i), (a)(ii), or (a)(iii) (e.g., a nucleic acid molecule described herein); (v) a vector comprising (a)(iv) (e.g., a vector described herein); (vi) a carrier comprising any one of (a)(i)-(a)(v) (e.g., a carrier described herein); or (vii) a composition comprising any one of (a)(i)-(a)(vi) (e.g., a composition described herein (e.g., a pharmaceutical composition)); (b) (i) gRNA comprising (ia) crRNA and tracrRNA, wherein the crRNA and tracrRNA are on separate nucleic acid molecules, or (ib) sgRNA; (ii) one or more DNA molecules encoding (b)(i); (iii) a vector comprising (b)(i) or (b)(ii) (e.g., the vector described herein); (iv) a carrier comprising any one of (b)(i)-(b)(iii) (e.g., the carrier described herein); or (v) a composition comprising any one of (b)(i)-(b)(iv) (e.g., a pharmaceutical composition) (e.g., the composition described herein); and (c) (i) a donor template nucleic acid (e.g., DNA) molecule (e.g., as defined herein); (ii) a vector comprising (c)(i) (e.g., the vector described herein); (iii) a carrier comprising any one of (c)(i)-(c)(ii) (e.g., the carrier described herein); or (iv) comprising (c)(i)-(c)(iii). Any of the following compositions (e.g., pharmaceutical compositions) (e.g., the compositions described herein (e.g., pharmaceutical compositions)).

[0281] Without being bound by theory, the HDR system can be used, for example, in methods for editing target nucleic acid molecules (e.g., the methods described herein), where cellular (e.g., in a subject, in vitro, or in vitro) molecular mechanisms utilize donor template nucleic acid molecules to repair and / or resolve cleavage sites in the target nucleic acid molecule mediated by a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) (e.g., a systemic Cas endonuclease), wherein the donor sequence is incorporated into the target nucleic acid molecule via, for example, HDR. See, for example, US 8697359, the entire contents of which are incorporated herein by reference for all purposes.

[0282] In some embodiments, an endonuclease (or a functional fragment, functional variant, or domain thereof) has the ability to mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule.

[0283] In some embodiments, the donor template nucleic acid molecule comprises at least about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 300, 400, or 500 or more nucleotides. In some embodiments, the donor template nucleic acid molecule comprises about 10-500, 10-400, 10-300, 10-200, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, or 10-20 nucleotides. In some embodiments, the donor template nucleic acid molecule comprises about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 300, 400, or 500 or more nucleotides. In some embodiments, the donor sequence of the donor template nucleic acid molecule includes substitution, addition, deletion, inversion, or another modification (e.g., relative to the nucleotide sequence of the target nucleic acid molecule).

[0284] In some embodiments, each homologous arm of the donor template nucleic acid molecule comprises at least about 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, or 300 nucleotides. In some embodiments, each homologous arm of the donor template nucleic acid molecule comprises about 10-300, 10-200, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-20, or 10-15 nucleotides. In some embodiments, each homologous arm of the donor template nucleic acid molecule comprises about 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, or 300 nucleotides. In some embodiments, each homologous arm shares at least about 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence homology with its target sequence. In some embodiments, the target sequence of the homologous arm is adjacent to the endonuclease cleavage site. In some embodiments, the target sequence of the homologous arm is within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 30 nucleotides of the endonuclease cleavage site.

[0285] In some embodiments, the donor template nucleic acid molecule is an ssDNA molecule, an ssRNA molecule, a dsDNA molecule, or a dsRNA molecule. In some embodiments, the donor template nucleic acid molecule of the system is a linear nucleic acid molecule. In some embodiments, the donor template nucleic acid molecule is a circular nucleic acid molecule. In some embodiments, the donor template nucleic acid molecule is contained in a vector and / or a carrier. In some embodiments, the donor template nucleic acid molecule contains one or more modified nucleotides. Nucleotide modifications are known in the art and are described herein. For example, one or more nucleotides may be modified to increase stability and reduce degradation (e.g., by endonucleases and / or exonucleases). Exemplary modifications include, but are not limited to, 2'-O-methyl (2'-OMe); 2'-O-methoxyethyl (2'-O-MOE); 2'-deoxy-2'-fluorine (2'-F); 2'-arabinose-fluorine (2'-Ara-F); 2'-O-benzyl; 2'-O-methyl-4-pyridine (2-O-methyl-4-pyridine (2'-O-CH2Py(4)); 2'F-4'-Cα-OMe; or 2',4'-di-Cα-OMe, deoxyribose, thiophosphate (PS (Rp isomer or Sp isomer)) (e.g., 5'-thiophosphate) (e.g., chiral thiophosphate), phosphorus Triesters, aminophosphates (e.g., 3'-aminoaminophosphate and aminoalkylaminophosphate), chiral thiophosphates, dithiophosphates (PS2), aminoalkyl phosphate triesters, methyl and other alkylphosphonates (e.g., methylphosphonate (MP), 3'-alkylenephosphonate), methoxypropylphosphonate (MOP), 5'-(E)vinylphosphonate, 5'-methylphosphonate, (S)-5'-C-methylphosphonate, hypophosphonates, thiocarbonylaminophosphate, thiocarbonylalkylphosphonate, thiocarbonylalkyl phosphate triester, boron phosphate, hypophosphonates and peptide nucleic acids (PNA), and any combination thereof. See also § 4.5.2.2 of this document, which describes modified gRNAs. Any modifications described in § 4.5.2.2 can also be used in the context of the donor template nucleic acid molecule.

[0286] In some embodiments, the donor sequence of the donor template nucleic acid molecule includes, for example, restriction sites, nucleotide polymorphisms, optional markers (e.g., drug resistance genes, fluorescent proteins, enzymes, etc.), which can be used to assess the successful addition of the donor sequence at the cleavage site, or in some cases, for other purposes (e.g., indicating expression at a target nucleic acid sequence (e.g., a gene)). In some cases, if located in a coding region, such nucleotide sequence differences will not alter the amino acid sequence or will result in silent amino acid changes (i.e., no change in protein structure or function). Alternatively, these sequence differences may include flanking recombination sequences, such as FLP, loxP sequences, etc., which can be activated at a later time to remove the marker sequence. 4.5.5.2 RT-based editing system

[0287] This article provides, in particular, RT-based systems for use in editing target nucleic acid molecules (e.g., in cells, or, for example, in a subject). In some embodiments, the system comprises (a) (i) a fusion protein comprising a Cas endonuclease (or a functional fragment, functional variant, or domain thereof) as described herein (e.g., as described herein) and a reverse transcriptase (or a functional fragment, functional variant, or domain thereof) (e.g., as described herein) (see, e.g., § 4.3.1.1); (ii) a nucleic acid molecule encoding (a)(i) (e.g., a nucleic acid molecule as described herein); (iii) a vector comprising (a)(ii) (e.g., a vector as described herein); (iv) a carrier comprising any one of (a)(i)-(a)(iii) (e.g., a carrier as described herein); or (v) a composition comprising any one of (a)(i)-(a)(iv) (e.g., a composition as described herein (e.g., a pharmaceutical composition)); and (b) (i) a template RNA (e.g., as described herein) (see, e.g., § 4.5.2); (ii) a DNA molecule encoding (b)(i); (iii) a fusion protein comprising (b)(i) (b)(ii) a carrier (e.g., the carrier described herein); (iv) a carrier comprising any one of (b)(i)-(b)(iii) (e.g., the carrier described herein); or (v) a composition comprising any one of (b)(i)-(b)(iv) (e.g., the composition described herein (e.g., a pharmaceutical composition)).

[0288] Without being bound by theory, RT-based editing systems can be used, for example, in methods for editing target nucleic acid molecules (e.g., the methods described herein), where a template nucleic acid binds to a target nucleic acid molecule (e.g., a double-stranded nucleic acid molecule (e.g., dsDNA molecule)) and to a fusion protein, thereby localizing the fusion protein to the target nucleic acid molecule. Subsequently, a Cas endonuclease of the fusion protein cleaves a single strand of the target nucleic acid molecule (e.g., a target double-stranded nucleic acid molecule (e.g., dsDNA molecule), allowing the 3' homology domain to bind to a sequence adjacent to the site to be edited on the target nucleic acid molecule (e.g., on the edited strand of the double-stranded nucleic acid molecule (e.g., dsDNA molecule)). It is assumed that the reverse transcriptase domain of the fusion protein utilizes the 3' target homology domain as a primer and the editing template as a template to, for example, aggregate sequences complementary to the editing template. Without being bound by theory, it is assumed that selecting an appropriate editing template can lead to the editing of the nucleotide sequence at the target site (e.g., substitution, deletion, or addition of one or more nucleotides at the target site), where the cell's endogenous DNA repair mechanisms resolve mismatched double-stranded nucleic acid molecules (e.g., dsDNA) to incorporate the desired edit. See, for example, WO 2021178720 and WO 2023039424, the entire contents of which are incorporated herein by reference for all purposes.

[0289] In some embodiments, the Cas endonuclease (a) has the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule; (b) cannot mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule; (c) has the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule and cannot mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule (i.e., nicking enzyme activity); and / or (d) has RNA-directed DNA endonuclease activity; or any combination of the foregoing.

[0290] In some embodiments, the target nucleic acid molecule of the system is a double-stranded nucleic acid (e.g., dsDNA) molecule, wherein one strand of the double-stranded nucleic acid (e.g., dsDNA) molecule is targeted for editing. In some embodiments, the system further comprises a gRNA (e.g., sgRNA) capable of guiding a Cas endonuclease (e.g., as described herein) to form a single-strand break (i.e., nick) in the unedited strand of the target double-stranded nucleic acid (e.g., dsDNA) molecule. It is not desired to be bound by theory, but rather that nicking of the unedited strand of the target double-stranded nucleic acid molecule (e.g., the target dsDNA molecule) induces a preferential substitution of the edited strand. In some embodiments, at least a portion of the nucleotide sequence of the gRNA (e.g., sgRNA) is complementary to a portion of the nucleotide sequence of the edited strand of the target double-stranded nucleic acid (e.g., dsDNA) molecule (as defined herein). In some embodiments, at least a portion of the nucleotide sequence of a second gRNA (e.g., sgRNA) binds to a portion of the nucleotide sequence of the edited strand of the double-stranded nucleic acid (e.g., dsDNA) molecule (as defined herein). In some embodiments, the gRNA is sgRNA. In some embodiments, the gRNA (e.g., sgRNA) is present on the same nucleic acid molecule as the template gRNA (or the nucleic acid (e.g., DNA) molecule encoding the gRNA is present on the same nucleic acid (e.g., DNA) molecule encoding the template gRNA). In some embodiments, the gRNA (e.g., sgRNA) is present on a different nucleic acid molecule than the template gRNA (or the nucleic acid (e.g., DNA) molecule encoding the gRNA is present on a different nucleic acid (e.g., DNA) molecule encoding the template gRNA).

[0291] In some embodiments, the Cas endonuclease (or its functional fragments, functional variants or domains) described herein is used in systems described in WO 2021178720 or WO 2023039424 (the entire contents of which are incorporated herein by reference for all purposes) (e.g., the Gene Writer™ system). 4.5.5.3 Nucleotide Editor Editing System

[0292] This article provides, in particular, a system based on a nucleobase editor (e.g., for use in editing target nucleic acid molecules (e.g., in cells, for example, in a subject's body)). In some embodiments, the system comprises (a) (i) a fusion protein comprising a Cas endonuclease (or a functional fragment or functional variant thereof) as described herein (e.g., as described herein) and a nucleobase editor (or a functional fragment or functional variant thereof) (e.g., as described herein) (see, e.g., § 4.3.1.2); (ii) a nucleic acid molecule encoding (a)(i) (e.g., a nucleic acid molecule as described herein); (iii) a vector comprising (a)(ii) (e.g., a vector as described herein); (iv) a carrier comprising any one of (a)(i)-(a)(iii) (e.g., a carrier as described herein); or (v) a composition comprising any one of (a)(i)-(a)(iv) (e.g., a composition as described herein (e.g., a pharmaceutical composition)); and (b) (i) a first gRNA comprising (ia) crRNA and tracrRNA, wherein the crRNA and tracrRNA are on separate nucleic acid molecules, or (ib) sgRNA; (ii) encoding (b)(i) (iii) one or more DNA molecules; (iv) a vector comprising (b)(i) or (b)(ii) (e.g., the vector described herein); (v) a carrier comprising any one of (b)(i)-(b)(iii) (e.g., the carrier described herein); or (v) a composition comprising any one of (b)(i)-(b)(iv) (e.g., the composition described herein (e.g., a pharmaceutical composition)).

[0293] Without being bound by theory, a nucleobase editor-based editing system can be used, for example, in methods for editing target nucleic acid molecules (e.g., the methods described herein), where a gRNA (e.g., sgRNA) binds to a target nucleic acid molecule (e.g., a double-stranded nucleic acid molecule (e.g., a dsDNA molecule)) and to a fusion protein, thereby localizing the fusion protein to the target nucleic acid molecule. Subsequently, an endonuclease (e.g., a cleavage enzyme) of the fusion protein cleaves a single strand of the target nucleic acid molecule (e.g., the target double-stranded nucleic acid molecule (e.g., the dsDNA molecule), allowing a nucleobase editor (e.g., a deaminase) to edit one or more nucleobases in the nucleotide sequence of the target nucleic acid molecule (e.g., in the single strand of the target double-stranded nucleic acid molecule (e.g., the edited strand)). See, for example, WO 2021050571 A1, WO 2022 / 204268, WO 2019079347 A1, the entire contents of which are incorporated herein by reference for all purposes.

[0294] In some embodiments, the Cas endonuclease (a) has the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule; (b) cannot mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule; (c) has the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule and cannot mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule (i.e., nicking enzyme activity); and / or (d) has RNA-directed DNA endonuclease activity; or any combination of the foregoing.

[0295] In some embodiments, the target nucleic acid molecule of the system is a double-stranded nucleic acid (e.g., dsDNA) molecule, wherein one strand of the double-stranded nucleic acid (e.g., dsDNA) molecule is targeted for editing. In some embodiments, the system further comprises a gRNA (e.g., sgRNA) capable of guiding an endonuclease (e.g., as described herein) of the system to form a single-strand break (i.e., nick) in the unedited strand of the target double-stranded nucleic acid (e.g., dsDNA) molecule. It is not desired to be bound by theory, but rather that nicking of the unedited strand of the target double-stranded nucleic acid molecule (e.g., the target dsDNA molecule) induces a preferential substitution of the edited strand. In some embodiments, at least a portion of the nucleotide sequence of the gRNA (e.g., sgRNA) is complementary to a portion of the nucleotide sequence of the edited strand of the target double-stranded nucleic acid (e.g., dsDNA) molecule (as defined herein). In some embodiments, at least a portion of the nucleotide sequence of a second gRNA (e.g., sgRNA) binds to a portion of the nucleotide sequence of the edited strand of the double-stranded nucleic acid (e.g., dsDNA) molecule (as defined herein). In some embodiments, the gRNA is sgRNA. In some embodiments, the gRNA (e.g., sgRNA) is present on the same nucleic acid molecule as the template gRNA (or the nucleic acid (e.g., DNA) molecule encoding the gRNA is present on the same nucleic acid (e.g., DNA) molecule encoding the template gRNA). In some embodiments, the gRNA (e.g., sgRNA) is present on a different nucleic acid molecule than the template gRNA (or the nucleic acid (e.g., DNA) molecule encoding the gRNA is present on a different nucleic acid (e.g., DNA) molecule encoding the template gRNA). 4.6 Nucleic Acid Molecules

[0296] This document further provides nucleic acid (e.g., DNA, RNA) molecules encoding any of the proteins described herein (e.g., Cas endonucleases (or functional fragments, functional variants, or domains thereof), heterologous proteins (e.g., reverse transcriptases, nucleobase editors), fusion proteins, conjugates), or any RNA molecules described herein (e.g., gRNA (e.g., sgRNA, template RNA)). The nucleic acid molecules described herein can be produced using methods commonly known in the art (e.g., chemical synthesis).

[0297] In some embodiments, the nucleic acid molecule is DNA. In some embodiments, the nucleic acid molecule is RNA (e.g., mRNA or circular RNA). In some embodiments, the nucleic acid (e.g., RNA) molecule is translatable RNA. In some embodiments, the nucleic acid molecule is single-stranded. In some embodiments, the nucleic acid molecule is double-stranded. In some embodiments, the nucleic acid molecule is a single-stranded RNA molecule. In some embodiments, the nucleic acid molecule is a single-stranded DNA molecule. In some embodiments, the nucleic acid molecule is a double-stranded RNA molecule. In some embodiments, the nucleic acid molecule is a double-stranded DNA molecule.

[0298] In some embodiments, the nucleic acid molecule is a linearly encoding nucleic acid construct. In some embodiments, the nucleic acid molecule is contained within a vector (e.g., a plasmid, a viral vector). In some embodiments, the nucleic acid molecule is contained within a non-viral vector. In some embodiments, the nucleic acid molecule is contained within a plasmid. In some embodiments, the nucleic acid molecule is contained within a viral vector. A more detailed description of vectors (e.g., non-viral, such as plasmids, and viruses) for both RNA and DNA nucleic acids is provided in § 4.7.

[0299] In some embodiments, nucleic acid molecules may be modified (compared to the sequence of a reference nucleic acid molecule), for example, to impart one or more of the following: (a) improved resistance to in vivo degradation, (b) improved in vivo stability, (c) reduced secondary structure, and / or (d) improved in vivo translatability compared to a reference nucleic acid sequence. Modifications include, but are not limited to, codon optimization, nucleotide variation (see, for example, described below), etc. Modifications are known in the art and are described herein (see, for example, § 4.5.2.2).

[0300] In some embodiments, the nucleotide sequence of a nucleic acid molecule is codon-optimized, for example, for expression. In some embodiments, this can be used to match codon frequencies in the target and host organisms to ensure correct folding; favor guanosine (G) and / or cytosine (C) content to increase nucleic acid stability; minimize the execution of tandem repeat codons or bases that may impair gene construction or expression; customize transcription and translation control regions; insert or remove protein transport sequences; remove / add post-translational alteration sites (e.g., glycosylation sites) in encoded proteins; add, remove, or reorganize protein domains; insert or delete restriction sites; modify ribosome binding sites and mRNA degradation sites; adjust translation rates to ensure correct folding of individual protein domains; or reduce or eliminate problematic secondary structures within polynucleotides. In some embodiments, the codon-optimized nucleic acid sequence exhibits one or more of the above (compared to a reference nucleic acid sequence). In some embodiments, the codon-optimized nucleic acid sequence exhibits one or more of the following compared to a reference nucleic acid sequence: improved resistance to in vivo degradation, improved in vivo stability, reduced secondary structures, and / or improved in vivo translatability. Codon optimization methods, tools, algorithms, and services are known in the art, and non-limiting examples include services from GeneArt (Life Technologies) and DNA2.0 (Menlo Park, California). In some embodiments, optimization algorithms are used to optimize open reading frame (ORF) sequences. In some embodiments, nucleic acid sequences are modified to optimize the number of G and / or C nucleotides compared to a reference nucleic acid sequence. The increase in the number of G and C nucleotides can be achieved by replacing codons containing adenosine (T) or thymidine (T) (or uracil (U)) nucleotides with codons containing G or C nucleotides. 4.7 Carrier

[0301] In some embodiments, the nucleic acid (DNA, RNA) molecules described herein are contained in a vector (e.g., a non-viral vector (e.g., a plasmid), a viral vector). Therefore, this document provides vectors (e.g., non-viral vectors (e.g., plasmids), viral vectors) that contain one or more nucleic acid molecules described herein (e.g., any protein described herein (e.g., a Cas endonuclease (or a functional fragment, functional variant, or domain thereof), a heterologous protein (e.g., a reverse transcriptase, a nucleobase editor), a fusion protein, a conjugate, etc.) or any RNA molecule described herein (e.g., gRNA (e.g., sgRNA, template RNA)) (e.g., see, e.g., §4.6). Such vectors can be readily manipulated by methods well known to those skilled in the art. The vector used can be any vector suitable for cloning nucleic acid molecules that can be transcribed for the desired nucleic acid molecule.

[0302] In some embodiments, the vector is a plasmid. Those skilled in the art know suitable plasmids for expressing target DNA. For example, plasmid DNA can be generated to allow efficient production of encoded endonucleases in cell lines (e.g., insect cell lines), for example using vectors as described in WO2009150222 A2 and vectors as defined in PCT claims 1 to 33, relating to the disclosure of claims 1 to 33 of WO2009150222 A2, the entire contents of which are incorporated herein by reference for all purposes.

[0303] In some embodiments, the vector is a viral vector. Viral vectors include both RNA-based and DNA-based vectors. Vectors can be designed to meet various specifications. For example, a viral vector can be engineered to be able to or not replicate in prokaryotic and / or eukaryotic cells. In some embodiments, the vector is replication-deficient. In some embodiments, the vector is replication-capable. The vector can be engineered or selected so that it will (or will not) fully or partially integrate into the genome of the host cell, thereby producing (or not producing, e.g., appendage expression) stable host cells containing the desired nucleic acids in their genome.

[0304] Exemplary viral vectors include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, lentiviral vectors, retroviral vectors, poxvirus vectors, parapoxvirus vectors, vaccinia virus vectors, fowlpox virus vectors, herpesvirus vectors, adeno-associated virus vectors, alphavirus vectors, lentiviral vectors, rhabdovirus vectors, measles virus, Newcastle disease virus vectors, piconemavirus vectors, or lymphocytic choriomeningitis virus vectors. In some embodiments, the viral vector is an adenovirus vector, an adeno-associated virus vector, a lentiviral vector, or a ring vector (as described, for example, in U.S. Patent 11,446,344, the entire contents of which are incorporated herein by reference for all purposes).

[0305] In some embodiments, the vector is an adenovirus vector (e.g., a human adenovirus vector, such as HADV or AdHu). In some embodiments, the adenovirus vector lacks the E1 region, making it replication-defective in human cells. Other regions of the adenovirus (such as E3 and E4) may also be missing. Exemplary adenovirus vectors include, but are not limited to, those described in, for example, WO2005071093 or WQ2006048215 (the entire contents of which are incorporated herein by reference for all purposes). Exemplary simian adenovirus vectors include AdCh63 (see, for example, WO2005071093, the entire contents of which are incorporated herein by reference for all purposes) or AdCh68.

[0306] Viral vectors can be produced using standard methods known to those skilled in the art in packaging / production cell lines (e.g., mammalian cell lines). Typically, a nucleic acid construct (e.g., a plasmid) encoding a transgene (e.g., the Cas endonuclease described herein) (along with additional elements, such as a promoter, inverted terminal repeat (ITR) flanking the transgene, and plasmids encoding, for example, viral replication and structural proteins) is transfected into the host cell line (i.e., the packaging / production cell line) along with one or more helper plasmids from the host cell (e.g., the host cell line). In some cases, depending on the viral vector, helper plasmids containing helper genes from another virus may also be required (e.g., in the case of an adeno-associated virus vector). Eukaryotic expression plasmids are commercially available from multiple suppliers, such as plasmid families: pcDNA™, pCR3.1™, pCMV™, pFRT™, pVAX1™, pCI™, Nanoplasmid™, and Pcaggs. Those skilled in the art are familiar with various transfection methods and can employ any suitable transfection method (e.g., using biochemical substances as carriers (e.g., lipid transfection with amines), by mechanical means, or by electroporation). Cells are cultured under conditions suitable for plasmid expression for a sufficient time. Viral particles can be purified from the cell culture medium using standard methods known to those skilled in the art. For example, by centrifugation followed by, for example, chromatography or ultrafiltration. 4.8 Carrier

[0307] In some embodiments, the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (see, for example, § 4.2), the fusion protein described herein (see, for example, § 4.3), the conjugate described herein (see, for example, § 4.3), the system described herein (see, for example, § 4.5) (or any one or more components thereof), the nucleic acid molecule described herein (see, for example, § 4.6), the vector described herein (see, for example, § 4.7), the cell described herein (see, for example, § 4.9), the reaction mixture described herein (see, for example, § 4.10), or the pharmaceutical composition described herein (see, for example, § 4.11) is formulated in one or more carriers.

[0308] Therefore, this disclosure specifically provides carriers comprising any one or more of the following: the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (see, for example, § 4.2); the fusion protein described herein (see, for example, § 4.3); the conjugate described herein (see, for example, § 4.3); the system described herein (see, for example, § 4.5) (or any one or more components thereof); the nucleic acid molecule described herein (see, for example, § 4.6); the vector described herein (see, for example, § 4.7); the cell described herein (see, for example, § 4.9); the reaction mixture described herein (see, for example, § 4.10); or the pharmaceutical composition described herein (see, for example, § 4.11).

[0309] Any of the foregoing elements (e.g., proteins, nucleic acid molecules, carriers, etc.) can be encapsulated within a carrier, chemically conjugated to a carrier, or associated with a carrier. In this context, the term "association" means that any of the foregoing elements (e.g., proteins, nucleic acid molecules, etc.) and one or more molecules of a carrier (e.g., one or more lipids of a lipid-based carrier (e.g., LNPs, liposomes, lipid complexes, and / or nanoliposomes)) substantially stably combine to form a larger complex or assembly without covalent bonding. In this context, the term "encapsulation" means incorporating any of the foregoing elements (e.g., proteins, nucleic acid molecules, etc.) into a carrier (e.g., lipid-based carriers, such as LNPs, liposomes, lipid complexes, and / or nanoliposomes), wherein the molecule (e.g., protein, nucleic acid molecule, etc.) is completely contained within the internal space of the carrier (e.g., lipid-based carriers, such as LNPs, liposomes, lipid complexes, and / or nanoliposomes).

[0310] Exemplary loaders include, but are not limited to, lipid-based loaders (e.g., lipid nanoparticles (LNPs), liposomes, lipid complexes, and nanoliposomes). In some embodiments, the loader is a lipid-based loader. In some embodiments, the loader is an LNP. In some embodiments, the LNP comprises cationic lipids, neutral lipids, cholesterol, and / or PEG lipids. Lipid-based loaders are further described in § 4.8.1 below. 4.8.1 Lipid-based carriers

[0311] In some embodiments, the Cas endonucleases described herein (or functional fragments, functional variants, or domains thereof) (see, for example, § 4.2), the fusion proteins described herein (see, for example, § 4.3), the conjugates described herein (see, for example, § 4.3), the systems described herein (see, for example, § 4.5) (or any one or more components thereof), the nucleic acid molecules described herein (see, for example, § 4.6), the carriers described herein (see, for example, § 4.7), the cells described herein (see, for example, § 4.9), the reaction mixtures described herein (see, for example, § 4.10), or the pharmaceutical compositions described herein (see, for example, § 4.11) are encapsulated or associated with one or more lipids (e.g., cationic lipids and / or neutral lipids) to form lipid-based carriers, such as lipid nanoparticles (LNPs), liposomes, lipid complexes, or nanoliposomes.

[0312] In some embodiments, any of the aforementioned molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) are encapsulated in one or more lipids (e.g., cationic lipids and / or neutral lipids) to form lipid-based carriers such as lipid nanoparticles (LNPs), liposomes, lipid complexes, or nanoliposomes. In some embodiments, molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) associate with one or more lipids (e.g., cationic lipids and / or neutral lipids) to form lipid-based carriers such as lipid nanoparticles (LNPs), liposomes, lipid complexes, or nanoliposomes. In some embodiments, molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) are encapsulated in LNPs (e.g., as described herein). The use of LNPs for mRNA delivery is described in further detail below: Hou X et al. Lipid nanoparticles for mRNA delivery. Nat Rev Mater. [Nature Review: Materials] 2021;6(12):1078-1094. doi: 10.1038 / s41578-021-00358-0. Electronic publication on 10 August 2021. PMID: 34394960; PMCID: PMC8353930, the entire contents of which are incorporated herein by reference for all purposes.

[0313] The molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) may be wholly or partially located within the internal space of LNPs, liposomes, lipid complexes, and / or nanoliposomes, associated within a lipid layer / membrane, or with the outer surface of a lipid layer / membrane. One purpose of incorporating molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) into LNPs, liposomes, lipid complexes, and / or nanoliposomes is to protect said molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) from environments or conditions that may contain enzymes or chemicals that degrade said molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) or from molecules or conditions that lead to rapid excretion of said molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.). Furthermore, incorporating molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) into LNPs, liposomes, lipid complexes, and / or nanoliposomes can promote the uptake of molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) and thus can enhance the therapeutic effects of proteins or nucleic acid molecules (e.g., RNA, e.g., mRNA). Therefore, incorporating molecules (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) into LNPs, liposomes, lipid complexes, and / or nanoliposomes may be particularly suitable for the pharmaceutical compositions described herein, for example, for intramuscular and / or intradermal administration.

[0314] In some embodiments, the molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) are formulated as lipid-based carriers (or lipid nanoformations). In some embodiments, the lipid-based carrier (or lipid nanoformation) is a liposome or lipid nanoparticle (LNP). In one embodiment, the lipid-based carrier is an LNP.

[0315] In some embodiments, the lipid-based carrier (or lipid nanoformation) comprises cationic lipids (e.g., ionizable lipids), non-cationic lipids (e.g., phospholipids), structural lipids (e.g., cholesterol), and PEG-modified lipids. In some embodiments, the lipid-based carrier (or lipid nanoformation) contains one or more molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc.) or pharmaceutically acceptable salts thereof.

[0316] As described herein, suitable compounds for use in lipid-based carriers (or lipid nanoformations) include all isomers and isotopes of the compounds described above, as well as all pharmaceutically acceptable salts, solvates or hydrates thereof, and all crystalline forms, mixtures of crystalline forms and anhydrides or hydrates thereof.

[0317] In addition to one or more molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc.), lipid-based carriers (or lipid nanoformulations) may further comprise a second lipid. In some embodiments, the second lipid is a cationic lipid, a non-cationic (e.g., neutral, anionic, or zwitterionic) lipid, or an ionizable lipid.

[0318] One or more naturally occurring and / or synthetic lipid compounds can be used in the preparation of lipid-based carriers (or lipid nanoformations).

[0319] Lipid-based carriers (or lipid nanoformations) may contain positively charged (cationic) lipids, neutral lipids, negatively charged (anionic) lipids, or combinations thereof. 4.8.1.1 Cationic lipids (positively charged) and ionizable lipids

[0320] In some embodiments, the lipid-based carrier (or lipid nanoformation) comprises one or more cationic lipids, such as cationic lipids that can exist in a positively charged or neutral form depending on pH, or amine-containing lipids that can be readily kinetically converted. In some embodiments, the cationic lipid is, for example, a lipid capable of carrying a positive charge under physiological conditions.

[0321] Exemplary cationic lipids include one or more positively charged amine groups. Examples of positively charged (cationic) lipids include, but are not limited to, N,N'-dimethyl-N,N'-bis(octadecyl)ammonium bromide (DDAB) and ammonium chloride (DDAC), N-(l-(2,3-dioleyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTMA), 3β-[N-(N',N'-dimethylaminoethyl)carbamoyl]cholesterol (DC-chol), 1,2-dioleoyloxy-3-[trimethylammonium]-propane (DOTAP), 1,2-bis(octadecyloxy)-3-[trimethylammonium]-propane (DSTAP), and 1,2-dioleoyloxypropyl-3-dimethyl-hydroxyethylammonium chloride (DORI), N,N-dioleyl-N N-Dimethylammonium chloride (DODAC), N,N-Dimethyl-2,3-diolenyloxypropylammonium (DODMA), 1,2-dioleoyl-3-dimethylammonium-propane (DODAP), 1,2-dioleoylcarbamoyl-3-dimethylammonium-propane (DOCDAP), 1,2-dilinoleoyl-3-dimethylammonium-propane (DLINDAP), 3-Dimethylamino-2-(cholest-5-en-3-β-oxybut-4-oxy)-1-(cis,cis-9,12-octadecadienoxy)propane (CLinDMA), 2-[5'-(cholest-5-en-3-β-oxy)-3'-oxaproxy]-3-dimethyl-1-(cis, Cis-9',12'-octadecadienoxy)propane (CpLin DMA), N,N-dimethyl-3,4-diolenoxybenzylamine (DMOBA), and cationic lipids such as those described in Martin et al., Current Pharmaceutical Design, pp. 1-394 (the entire contents of which are incorporated herein by reference for all purposes). In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises more than one cationic lipid.

[0322] In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises a cationic lipid with an effective pKa greater than 6.0. In some embodiments, the lipid-based carrier (or lipid nanoformulation) further comprises a second cationic lipid having an effective pKa different from that of the first cationic lipid (e.g., greater than the first effective pKa).

[0323] In some embodiments, cationic lipids that may be used in lipid-based carriers (or lipid nanoformations) include, for example, those described in Table 4 of WO 2019 / 217941 (the entire contents of which are incorporated herein by reference for all purposes).

[0324] In some embodiments, the cationic lipids are ionizable lipids (e.g., lipids that are protonated at low pH but remain neutral at physiological pH). In some embodiments, lipid-based carriers (or lipid nanoformations) may comprise one or more additional ionizable lipids different from those described herein. Exemplary ionizable lipids include, but are not limited to: (LP01) (SM-086) (SM-102) (ALC-0315) (Lipid 10) (Lipid A9) and (DLin-MC3-DMA). (See WO 2017004143 A1, the entire contents of which are incorporated herein by reference for all purposes.)

[0325] In some embodiments, the lipid-based carrier (or lipid nanoformulation) further comprises one or more compounds described in WO 2021 / 113777 (the entire contents of which are incorporated herein by reference for all purposes) (e.g., lipids having formula (3), such as lipids in Table 3 of WO 2021 / 113777).

[0326] In one embodiment, the ionizable lipids are those disclosed in Hou, X. et al., Nat Rev Mater [Nature Review Materials] 6, 1078-1094 (2021). https: / / doi.org / 10.1038 / s41578-021-00358-0 (e.g., L319, C12-200, and DLin-MC3-DMA) (the entire contents of which are incorporated herein by reference for all purposes).

[0327] Examples of other ionizable lipids that can be used in lipid-based carriers (or lipid nanoformulations) include, but are not limited to, compounds having one or more of the following formulas: X of US 2016 / 0311759; I of US 20150376115 or US 2016 / 0376224; compound 5 or compound 6 of US 2016 / 0376224; I, IA, or II of US 9,867,888; I, II, or III of US 2016 / 0151284; I, IA, II, or IIA of US 2017 / 0210967; Ic of US 2015 / 0140070; A of US 2013 / 0178541; I of US 2013 / 0303587 or US 2013 / 0123338; ​​US US 2015 / 0141678, Part I; US 2015 / 0239926, Parts II, III, IV, or V; US 2017 / 0119904, Part I; WO 2017 / 117528, Part I or II; US 2012 / 0149894, Part A; US 2015 / 0057373, Part A; WO 2013 / 116126, Part A; US 2013 / 0090372, Part A; US 2013 / 0274523, Part A; US 2013 / 0274504, Part A; US 2013 / 0053572, Part A; WO 2013 / 016058, Part A; WO 2012 / 162210, Part A; US 2008 / 042973, Part I; US US 2012 / 01287670, I, II, III, or IV; US 2014 / 0200257, I or II; US 2015 / 0203446, I, II, or III; US 2015 / 0005363, I or III; US 2014 / 0308304, I, IA, IB, IC, ID, II, IIA, IIB, IIC, IID, or III-XXIV; US 2013 / 0338210; WO 2009 / 132131, I, II, III, or IV; US 2012 / 01011478, A; US 2012 / 0027796, I or XXXV; US 2012 / 0058144, XIV or XVII; US 2013 / 0323269; US US 2011 / 0117125, Part I; US 2011 / 0256175, Part I, II, or III; US 2012 / 0202871, Part I, II, III, IV, V, VI, VII, VIII, IX, X, XI, XII; US 2011 / 0076335, Part I, II, III, IV, V, VI, VII, VIII, X, XII, XIII, XIV, XV, or XVI; US 2006 / 008378, Part I or II;WO2015 / 074085, Part I (e.g., ATX-002); US 2013 / 0123338, Part I; US 2015 / 0064242, Part I or XAYZ; US 2013 / 0022649, Part XVI, XVII, or XVIII; US 2013 / 0116307, Part I, II, or III; US 2013 / 0116307, Part I, II, or III; US 2010 / 0062967, Part I or II; US 2013 / 0189351, Part IX; US 2014 / 0039032, Part I; US 2018 / 0028664, Part V; US 2016 / 0317458, Part I; US 2013 / 0195920, Part I; US 10,221,127 of 5, 6 or 10; WO 2018 / 081480 of III-3; WO 2020 / 081938 of I-5 or I-8; WO 2015 / 199952 of I (e.g., compound 6 or 22) and Table 1 therein; US 9,867,888 of 18 or 25; US 2019 / 0136231 of A; WO 2020 / 219876 of II; US 2012 / 0027803 of 1; US ​​2019 / 0240349 of OF-02; US 10,086,013 of 23; Miao et al. (2020) of cKK-E12 / A6; WO C12-200 of WO 2010 / 053572; 7C1 of Dahlman et al. (2017); 304-O13 or 503-O13 of Whitehead et al.; TS-P4C2 of U.S. 9,708,628; I of WO 2020 / 106946; I of WO 2020 / 106946; (1), (2), (3) or (4) of WO 2021 / 113777; and any one of Tables 1-16 of WO 2021 / 113777, the entire contents of which are incorporated herein by reference for all purposes.

[0328] In some embodiments, the lipid-based carrier (or lipid nanoformulation) further comprises a biodegradable, ionizable lipid, such as octadecano-9,12-dienoic acid (9Z,12Z)-3-((4,4-bis(octyloxy)butyryl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl ester, also known as (9Z,12Z)-octadecano-9,12-dienoic acid 3-((4,4-bis(octyloxy)butyryl)oxy)-2-((((3-(diethylamino)propoxy)carbonyl)oxy)methyl)propyl ester. See, for example, lipids of WO 2019 / 067992, WO 2017 / 173054, WO 2015 / 095340 and WO 2014 / 136086, the entire contents of which are incorporated herein by reference for all purposes. 4.8.1.2 Non-cationic lipids (e.g., phospholipids)

[0329] In some embodiments, the lipid-based carrier (or lipid nanoformulation) further comprises one or more non-cationic lipids. In some embodiments, the non-cationic lipid is a phospholipid. In some embodiments, the non-cationic lipid is a phospholipid substitute or replacement. In some embodiments, the non-cationic lipid is a negatively charged (anionic) lipid.

[0330] Exemplary non-cationic lipids include, but are not limited to, distearyl-sn-glycero-phosphoethanolamine, distearylphosphatidylcholine (DSPC), dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), dioleoylphosphatidylcholine (DOPG), dipalmitoylphosphatidylglycerol (DPPG), dioleoylphosphatidylethanolamine (DOPE), palmitoyloleoylphosphatidylcholine (POPC), palmitoyloleoylphosphatidylethanolamine (POPE), and dioleoylphosphatidyl... Acylethanolamine 4-(N-maleimidemethyl)-cyclohexane-1-carboxylate (DOPE-mal), dipalmitoylphosphatidylethanolamine (DPPE), dimyristoylphosphatidylethanolamine (DMPE), distearate-phosphatidylethanolamine (DSPE), monomethylphosphatidylethanolamine (such as 16-O-monomethylPE), dimethylphosphatidylethanolamine (such as 16-O-dimethylPE), 18-l-transPE, 1-stearoyl-2-oleoylphosphatidylethanolamine (SOPE), hydrogenated soybean phosphatidylcholine (HSPC), lecithin choline (EPC), dioleoylphosphatidylserine (DOPS), sphingomyelin (SM), dimyristoylphosphatidylcholine (DMPC), dimyristoylphosphatidylglycerol (DMPG), distearate phosphatidylglycerol (DSPG), disqualylphosphatidylcholine (DEPC), palmitoylphosphatidylglycerol (POPG), ditransolenoylphosphatidylethanolamine (DEPE), 1,2-di Lauroyl-sn-glycerol-3-phosphate choline (DLPC), sodium 1,2-di-tetradecanoyl-sn-glycerol-3-phosphate (DMPA), phosphatidylcholine (lecithin), phosphatidylethanolamine, lysophosphatidylcholine, lysophosphatidylethanolamine, phosphatidylserine, phosphatidylinositol, sphingomyelin, lecithin (ESM), phosphatidylethanolamine (cephalin), cardiolipin, phosphatidic acid, cerebroside, dihexadecanophosphate, lysophosphatidylcholine, dilinoleylphosphatidylcholine, or mixtures thereof. It should be understood that other diacylphosphatidylcholines and diacylphosphatidylethanolamine phospholipids may also be used. The acyl group in these lipids is preferably derived from a C-terminated compound. 10- C 24 The acyl group of the fatty acid in the carbon chain, for example, lauroyl, myristoyl, palmitoyl, stearoyl, or oleoyl. In some embodiments, other exemplary lipids include, but are not limited to, those described in Kim et al. (2020) dx.doi.org / 10.1021 / acs.nanolett.0c01386 (the entire contents of which are incorporated herein by reference for all purposes). In some embodiments, such lipids include plant lipids (e.g., DGTS) that have been found to improve mRNA liver transfection.

[0331] In some embodiments, the lipid-based carrier (or lipid nanoformation) may comprise a combination of distearylphosphatidylcholine / cholesterol, dipalmitoylphosphatidylcholine / cholesterol, myristoylphosphatidylcholine / cholesterol, 1,2-dioleoyl-sn-glycerol-3-phosphocholine (DOPC) / cholesterol, or lecithin / cholesterol.

[0332] Other examples of suitable noncationic lipids include, but are not limited to, nonphospholipids such as stearylamine, dodecylamine, hexadecylamine, acetyl palmitate, glyceryl ricinoleate, hexadecyl stearate, isopropyl myristate, amphoteric acrylic polymers, triethanolamine-lauryl sulfate, alkyl-aryl sulfates, polyethoxylated fatty acid amides, dioctadecyl dimethyl ammonium bromide, ceramides, sphingomyelin, etc. Other noncationic lipids are described in WO 2017 / 099823 or US 2018 / 0028664, the entire contents of which are incorporated herein by reference for all purposes.

[0333] In one embodiment, the lipid-based carrier (or lipid nanoformulation) further comprises one or more non-cationic lipids, said one or more non-cationic lipids being oleic acid or compounds having formula I, II or IV of US 2018 / 0028664 (the entire contents of which are incorporated herein by reference for all purposes).

[0334] The non-cationic lipid content can be, for example, 0-30% (mol) of the total lipid component present. In some embodiments, the non-cationic lipid content is 5%-20% (mol) or 10%-15% (mol) of the total lipid component present.

[0335] In some embodiments, the lipid-based carrier (or lipid nanoformulation) further comprises neutral lipids, and the molar ratio of ionizable lipids to neutral lipids ranges from about 2:1 to about 8:1 (e.g., about 2:1, 3:1, 4:1, 5:1, 6:1, 7:1 or 8:1).

[0336] In some embodiments, the lipid-based carrier (or lipid nanoformation) does not contain any phospholipids.

[0337] In some embodiments, the lipid-based carrier (or lipid nanoformation) may further comprise one or more phospholipids, and optionally one or more additional molecules having similar molecular shapes and sizes, which have both hydrophobic and hydrophilic portions (e.g., cholesterol). 4.8.1.3 Structural lipids

[0338] The lipid-based carriers (or lipid nanoformulations) described herein may further comprise one or more structured lipids. As used herein, the term "structured lipid" refers to sterols (e.g., cholesterol) and also to lipids containing sterol moieties.

[0339] Incorporating structural lipids into lipid nanoparticles can help reduce the aggregation of other lipids within the particles. Structural lipids can be selected from the group consisting of, but not limited to, cholesterol or cholesterol derivatives, coprosterol, sitosterol, ergosterol, campesterol, stigmasterol, phytosterol, tomatine, tomatine, ursolic acid, α-tocopherol, hornet's salt, phytosterols, steroids, and mixtures thereof. In some embodiments, the structural lipid is a steroid. In some embodiments, the structural lipid is a steroid. In some embodiments, the structural lipid is cholesterol. In some embodiments, the structural lipid is an analogue of cholesterol. In some embodiments, the structural lipid is α-tocopherol.

[0340] In some embodiments, structural lipids may be incorporated into lipid-based carriers at a molar ratio (cholesterol phospholipids) ranging from about 0.1 to 1.0.

[0341] In some embodiments, sterols (when present) may include one or more of cholesterol or cholesterol derivatives, such as those described in WO 2009 / 127060 or US 2010 / 0130588 (the entire contents of which are incorporated herein by reference for all purposes). Other exemplary sterols include phytosterols, including those described in Eygeris et al. (2020), NanoLett. [Nano Communications] 2020;20(6):4543-4549 (the entire contents of which are incorporated herein by reference for all purposes).

[0342] In some embodiments, the structural lipid is a cholesterol derivative. Non-limiting examples of cholesterol derivatives include polar analogs such as 5α-cholesterol, 53-cosanosterol, cholesterol-(2'-hydroxy)-ethyl ether, cholesterol-(4'-hydroxy)-butyl ether, and 6-ketocholesterol; nonpolar analogs such as 5α-cholesterol, cholesterolenone, 5α-cholesterone, 5p-cholesterone, and cholesterol decanoate; and mixtures thereof. In some embodiments, the cholesterol derivative is a polar analog, for example, cholesterol-(4'-hydroxy)-butyl ether. Exemplary cholesterol derivatives are described in WO 2009 / 127060 and US2010 / 0130588, the entire contents of which are incorporated herein by reference for all purposes.

[0343] In some embodiments, the lipid-based carrier (or lipid nanoformulation) further comprises sterols in an amount of 0-50 mol% (e.g., 0-10 mol%, 10-20 mol%, 20-50 mol%, 20-30 mol%, 30-40 mol%, or 40-50 mol%) of the total lipid component. 4.8.1.4 Polymers and polyethylene glycol (PEG)-lipids

[0344] In some embodiments, the lipid-based carrier (or lipid nanoformulation) may comprise one or more polymers or copolymers, such as poly(lactic acid-co-hydroxyacetic acid) (PFAG) nanoparticles.

[0345] In some embodiments, the lipid-based carrier (or lipid nanoformulation) may comprise one or more polyethylene glycol (PEG) lipids. Examples of usable PEG-lipids include, but are not limited to, 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-350] (mPEG 350 PE); 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-550] (mPEG 550 PE); 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-750] (mPEG 750 PE); 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-1000] (mPEG 1000 PE); 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-2000] (mPEG 2000 PE). PE); 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-3000] (mPEG3000 PE); 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-5000] (mPEG 5000 PE); N-acyl-sphingosine-1-[succinyl(methoxypolyethylene glycol) 750] (mPEG 750 ceramide); N-acyl-sphingosine-1-[succinyl(methoxypolyethylene glycol) 2000] (mPEG 2000 ceramide); and N-acyl-sphingosine-1-[succinyl(methoxypolyethylene glycol) 5000] (mPEG 5000 ceramide). In some embodiments, the PEG lipid is a polyethylene glycol-diacylglycerol (i.e., polyethylene glycol diacylglycerol (PEG-DAG), PEG-cholesterol, or PEG-DMB) conjugate.

[0346] In some embodiments, the lipid-based carrier (or nanoformulation) comprises one or more conjugated lipids (such as PEG-conjugated lipids or lipids conjugated to polymers as described in Table 5 of WO 2019 / 217941, the entire contents of which are incorporated herein by reference for all purposes). In some embodiments, one or more conjugated lipids are formulated with one or more ionic lipids (e.g., non-cationic lipids, such as neutral or anionic or zwitterionic lipids); and one or more sterols (e.g., cholesterol).

[0347] PEG conjugates may include PEG-dilaurylglycerol (C12), PEG-dimyristylglycerol (C14), PEG-dipalmitylglycerol (C16), PEG-distearatelglycerol (C18), PEG-dilaurylglycamide (C12), PEG-dimyristylglycamide (C14), PEG-dipalmitylglycamide (C16), and PEG-distearatelglycamide (C18).

[0348] In some embodiments, the conjugated lipids (when present) may include one or more of the following: PEG-diacylglycerol (DAG) (such as 1-(monomethoxy-polyethylene glycol)-2,3-dimyristoylglycerol (PEG-DMG)), PEG-dialkoxypropyl (DAA), PEG-phospholipids, PEG-ceramides (Cer), polyethylene glycol-modified phosphatidylethanolamine (PEG-PE), PEG-succinate diacylglycerol (PEGS-DAG) (such as 4-O-(2',3'-di(tetradecanoyloxy)propyl-1-O-(w-methoxy(polyethoxy)ethyl)succinate (PEG-S-DMG)), PEG-dialkoxypropylcarbamate, N-(carbonyl-methoxy-polyethylene glycol 2000)-1,2-distearate-sn-glycerol-3-phosphate ethanolamine sodium salt, and in WO The ones described in Table 2 of 2019 / 051289 (the entire contents of which are incorporated herein by reference for all purposes) and the combinations thereof.

[0349] Other exemplary PEG-lipid conjugates are described, for example, in US 5,885,613, US 6,287,591, US2003 / 0077829, US 2003 / 0077829, US 2005 / 0175682, US 2008 / 0020058, US 2011 / 0117125, US 2010 / 0130588, US 2016 / 0376224, US 2017 / 0119904, US 2018 / 0028664 and WO 2017 / 099823, the entire contents of which are incorporated herein by reference for all purposes.

[0350] In some embodiments, the PEG-lipid is a compound having formula III, III-aI, III-a-2, III-b-1, III-b-2, or V, as specified in US 2018 / 0028664 (which is incorporated herein by reference in its entirety). In some embodiments, the PEG-lipid has formula II as specified in US 2015 / 0376115 or US 2016 / 0376224 (the entire contents of which are incorporated herein by reference for all purposes). In some embodiments, the PEG-DAA conjugate may be, for example, PEG-dilauryloxypropyl, PEG-dimyristyloxypropyl, PEG-dispalmityloxypropyl, or PEG-distearateloxypropyl. In some embodiments, the PEG-lipid comprises one of the following: .

[0351] In some embodiments, lipids conjugated to molecules other than PEG can also be used instead of PEG-lipids. For example, polyoxazoline (POZ)-lipid conjugates, polyamide-lipid conjugates (such as ATTA-lipid conjugates), and cationic polymer lipid (GPL) conjugates can be used instead of or in combination with PEG-lipids.

[0352] Exemplary conjugated lipids (e.g., PEG-lipids, (POZ)-lipid conjugates, ATTA-lipid conjugates, and cationic polymer-lipids) include those described in Table 2 of WO 2019 / 051289A9 (the entire contents of which are incorporated herein by reference for all purposes).

[0353] In some embodiments, the amount of conjugated lipids (e.g., polyethylene glycol-modified lipids) may be 0-20 mol% of the total lipid component present in the lipid-based carrier (or lipid nanoformulation). In some embodiments, the content of conjugated lipids (e.g., polyethylene glycol-modified lipids) may be 0.5-10 mol% or 2-5 mol% of the total lipid component.

[0354] When needed, the lipid-based carriers (or lipid nanoformations) described herein can be coated with a polymer layer to enhance in vivo stability (e.g., spatially stable LNPs).

[0355] Examples of suitable polymers include, but are not limited to, poly(ethylene glycol) that can form a hydrophilic surface layer that improves the circulating half-life of liposomes and increases the amount of lipid nanoformulations (e.g., liposomes or LNPs) reaching therapeutic targets. See, for example, Working et al., J Pharmacol Exp Ther, 289: 1128-1133 (1999); Gabizon et al., J Controlled Release, 53: 275-279 (1998); Adlakha Hutcheon et al., Nat Biotechnol, 17: 775-779 (1999); and Koning et al., Biochim Biophys Acta, 1420: 153-167 (1999), the entire contents of which are incorporated herein by reference for all purposes. 4.8.1.5 Percentage of lipid nanoform formulation components

[0356] In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises one or more of the molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc.), optionally non-cationic lipids (e.g., phospholipids), sterols, neutral lipids, and optionally conjugated lipids that inhibit particle aggregation (e.g., PEGylated lipids). In some embodiments, the lipid-based carrier (or lipid nanoformulation) further comprises a payload (e.g., the molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc.)). The amounts of these components can be varied independently to obtain desired properties. For example, in some embodiments, ionizable lipids (including lipid compounds described herein) are present in amounts from about 20 mol% to about 100 mol% of the total lipid composition (e.g., 20-90 mol%, 20-80 mol%, 20-70 mol%, 25-100 mol%, 30-70 mol%, 30-60 mol%, 30-40 mol%, 40-50 mol%, or 50-90 mol%); non-cationic lipids (e.g., phospholipids) are present in amounts from about 0 mol% to about 50 mol% of the total lipid composition (e.g., 0-40 mol%, 0-30 mol%, 5-50 mol%, 5-40 mol%, 5-30 mol%, or 5-10 mol%); and conjugated lipids (e.g., polyethylene glycol-modified lipids) are present in amounts from about 0.5 mol% to about 20 mol% of the total lipid composition (e.g., 1-10 mol%). (5-10% mol%), and the amount of sterol present is about 0 mol% to about 60 mol% of the total lipid fraction (e.g., 0-50 mol%, 10-60 mol%, 10-50 mol%, 15-60 mol%, 15-50 mol%, 20-50 mol%, 20-40 mol%), provided that the total mol% of the lipid fraction does not exceed 100%.

[0357] In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises about 25-100 mol% of ionizable lipids (including lipid compounds described herein), about 0-50 mol% of phospholipids, about 0-50 mol% of sterols, and about 0-10 mol% of polyethylene glycol-modified lipids.

[0358] In some embodiments, the lipid-based load comprises a payload (e.g., a molecule described herein (e.g., a protein, nucleic acid molecule, carrier, system, etc.)) formulated in lipid nanoparticles, wherein the lipid nanoparticles comprise about 25-100 mol% of ionizable lipids (including lipid compounds described herein), about 0-50 mol% of phospholipids, about 0-50 mol% of sterols, and about 0-10 mol% of polyethylene glycol-modified lipids. In some embodiments, the encapsulation efficiency of the payload may be at least 70%.

[0359] In one embodiment, the lipid-based carrier (or lipid nanoformulation) comprises about 25-100 mol% of ionizable lipids (including lipid compounds described herein); about 0-40 mol% of phospholipids (e.g., DSPC), about 0-50 mol% of sterols (e.g., cholesterol), and about 0-10 mol% of polyethylene glycol-modified lipids.

[0360] In some embodiments, the lipid-based carrier comprises a payload (e.g., a molecule described herein (e.g., a protein, nucleic acid molecule, carrier, system, etc.)) formulated in lipid nanoparticles, wherein the lipid nanoparticles comprise about 25-100 mol% of ionizable lipids (including lipid compounds described herein), about 0-40 mol% of phospholipids (e.g., DSPC), about 0-50 mol% of sterols (e.g., cholesterol), and about 0-10 mol% of polyethylene glycol-modified lipids. In some embodiments, the encapsulation efficiency of the payload may be at least 70%.

[0361] In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises about 30-60 mol% (e.g., about 35-55 mol% or about 40-50 mol%) of ionizable lipids (including lipid compounds described herein), about 0-30 mol% (e.g., 5-25 mol% or 10-20 mol%) of phospholipids, about 15-50 mol% (e.g., 18.5-48.5 mol% or 30-40 mol%) of sterols and about 0-10 mol% (e.g., 1-5 mol% or 1.5-2.5 mol%) of polyethylene glycol-modified lipids.

[0362] In some embodiments, the lipid-based load comprises a payload (e.g., a molecule described herein (e.g., a protein, nucleic acid molecule, carrier, system, etc.)) formulated in lipid nanoparticles, wherein the lipid nanoparticles comprise about 30-60 mol% (e.g., about 35-55 mol% or about 40-50 mol%) of ionizable lipids (including lipid compounds described herein), about 0-30 mol% (e.g., 5-25 mol% or 10-20 mol%) of phospholipids, about 15-50 mol% (e.g., 18.5-48.5 mol% or 30-40 mol%) of sterols, and about 0-10 mol% (e.g., 1-5 mol% or 1.5-2.5 mol%) of polyethylene glycol-modified lipids. In some embodiments, the encapsulation efficiency of the payload may be at least 70%.

[0363] In some embodiments, the molar ratio of ionizable lipid / sterol / phospholipid (or another structural lipid) / PEG-lipid / another component varies within the following ranges: ionizable lipid (25%-100%); phospholipid (DSPC) (0-40%); sterol (0-50%); and PEG lipid (0-5%).

[0364] In some embodiments, the lipid-based carrier comprises a payload (e.g., a molecule described herein (e.g., a protein, nucleic acid molecule, carrier, system, etc.)) formulated in lipid nanoparticles, wherein the lipid nanoparticles comprise a molar ratio of ionizable lipid / sterol / phospholipid (or another structural lipid) / PEG-lipid / another component in the following range: ionizable lipid (25%-100%); phospholipid (DSPC) (0-40%); sterol (0-50%); and PEG lipid (0-5%). In some embodiments, the encapsulation efficiency of the payload may be at least 70%.

[0365] In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises 50%-75% ionizable lipids (including lipid compounds as described herein) by mol% or wt% of total lipid components, 20%-40% sterols (e.g., cholesterol or derivatives), 0-10% noncationic lipids, and 1%-10% conjugated lipids (e.g., polyethylene glycol-modified lipids).

[0366] In some embodiments, the lipid-based carrier comprises a payload (e.g., a molecule described herein (e.g., a protein, nucleic acid molecule, carrier, system, etc.)) formulated in lipid nanoparticles, wherein the lipid nanoparticles comprise 50%-75% by mol% or wt% of ionizable lipids (including lipid compounds as described herein), 20%-40% by sterols (e.g., cholesterol or its derivatives), 0-10% by noncationic lipids, and 1%-10% by conjugated lipids (e.g., polyethylene glycol-modified lipids). In some embodiments, the encapsulation efficiency of the payload may be at least 70%.

[0367] In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises (i) molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc. described herein); (ii) cationic lipids comprising 50 mol% to 65 mol% of the total lipids present in the lipid-based carrier; (iii) non-cationic lipids comprising a mixture of phospholipids and their cholesterol derivatives, wherein the phospholipids comprise 3 mol% to 15 mol% of the total lipids present in the lipid-based carrier, and cholesterol or its derivatives comprise 30 mol% to 40 mol% of the total lipids present in the lipid-based carrier; and (iv) conjugated lipids comprising 0.5 mol% to 2 mol% of the total lipids present in the particles.

[0368] In some embodiments, the lipid-based carrier (or lipid nanoformulation) comprises (i) molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc. described herein); (ii) cationic lipids comprising 50 mol% to 85 mol% of the total lipids present in the lipid-based carrier; (iii) non-cationic lipids comprising 13 mol% to 49.5 mol% of the total lipids present in the lipid-based carrier; and (d) conjugated lipids comprising 0.5 mol% to 2 mol% of the total lipids present in the lipid-based carrier.

[0369] In some embodiments, the phospholipid component in the mixture may be present at 2 mol% to 20 mol%, 2 mol% to 15 mol%, 2 mol% to 12 mol%, 4 mol% to 15 mol%, 4 mol% to 10 mol%, 5 mol% to 10 mol% (or any portion of these ranges) of the total lipid component. In some embodiments, the lipid-based carrier (or lipid nanoformulation) is phospholipid-free.

[0370] In some embodiments, the sterol component (e.g., cholesterol or its derivatives) in the mixture may account for 25 mol% to 45 mol%, 25 mol% to 40 mol%, 25 mol% to 35 mol%, 25 mol% to 30 mol%, 30 mol% to 45 mol%, 30 mol% to 40 mol%, 30 mol% to 35 mol%, 35 mol% to 40 mol%, 27 mol% to 37 mol%, or 27 mol% to 35 mol% (or any portion of these ranges) of the total lipid component.

[0371] In some embodiments, the non-ionizable lipid component in the lipid-based carrier (or lipid nanoformulation) may be present as 5 mol% to 90 mol%, 10 mol% to 85 mol%, or 20 mol% to 80 mol% (or any portion of these ranges) of the total lipid component.

[0372] The ratio of total lipid components to payload (e.g., encapsulated therapeutic agents, such as molecules described herein (e.g., proteins, nucleic acid molecules, carriers, systems, etc.)) can be varied as needed. For example, the ratio of total lipid components to payload (mass or weight) can be from about 10:1 to about 30:1. In some embodiments, the ratio of total lipid components to payload (mass / mass ratio; w / w ratio) can range from about 1:1 to about 25:1, from about 10:1 to about 14:1, from about 3:1 to about 15:1, from about 4:1 to about 10:1, from about 5:1 to about 9:1, or from about 6:1 to about 9:1. The total lipid component and payload can be adjusted to provide the desired N / P ratio, such as 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or higher. Generally, the total lipid content of lipid-based carriers (or lipid nanoformulations) can range from about 5 mg / mL to about 30 mg / mL. The nitrogen:phosphate ratio (N:P ratio) is evaluated at values ​​between 0.1 and 100.

[0373] Encapsulation efficiency of a payload (such as proteins and / or nucleic acids) describes the amount of protein and / or nucleic acid that is encapsulated or otherwise associated with a lipid nanoparticle formulation (e.g., liposomes or LNPs) after preparation, relative to the initial amount provided. Encapsulation efficiency is ideally high (e.g., at least 70%, 80%, 90%, 95%, close to 100%). Encapsulation efficiency can be measured, for example, by comparing the amount of protein or nucleic acid in a solution containing liposomes or LNPs before and after cleavage with one or more organic solvents or detergents. The amount of free protein or nucleic acid (e.g., RNA) in solution can be measured using anion exchange resins. The amount of free protein and / or nucleic acid (e.g., RNA) in solution can be measured using fluorescence. For the lipid-based carriers (or lipid nanoformations) described herein, the encapsulation efficiency of proteins and / or nucleic acids may be at least 50%, such as 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In some embodiments, the encapsulation efficiency may be at least 70%. In some embodiments, the encapsulation efficiency may be at least 80%. In some embodiments, the encapsulation efficiency may be at least 90%. In some embodiments, the encapsulation efficiency may be at least 95%. 4.9 cells

[0374] This disclosure provides, in particular, cells (e.g., host cells) comprising any one or more of the following: the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (see, e.g., § 4.2); the fusion protein described herein (see, e.g., § 4.3); the conjugate described herein (see, e.g., § 4.3); the system described herein (see, e.g., § 4.5) (or any one or more components thereof); the nucleic acid molecule described herein (see, e.g., § 4.6); the vector described herein (see, e.g., § 4.7); the reaction mixture described herein (see, e.g., § 4.10); the carrier described herein (see, e.g., § 4.8); or the pharmaceutical composition described herein (see, e.g., § 4.11).

[0375] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an animal cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is in vitro. In some embodiments, the cell is in vivo. In some embodiments, the cell is ex vivo.

[0376] Any of the foregoing methods (e.g., endonucleases, fusion proteins, systems, vectors, delivery agents, etc.) can be delivered into cells (e.g., host cells) using standard methods known in the art. Cells (e.g., host cells) can be cultured in vitro or ex vivo using standard methods known in the art. 4.10 Reaction mixture

[0377] This disclosure provides, in particular, reaction mixtures comprising any one or more of the following: the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (see, for example, § 4.2); the fusion protein described herein (see, for example, § 4.3); the conjugate described herein (see, for example, § 4.3); the system described herein (see, for example, § 4.5) (or any one or more components thereof); the nucleic acid molecule described herein (see, for example, § 4.6); the vector described herein (see, for example, § 4.7); the loading agent described herein (see, for example, § 4.8); or the pharmaceutical composition described herein (see, for example, § 4.11).

[0378] In some embodiments, the reaction mixture comprises a target nucleic acid molecule (e.g., as described herein). In some embodiments, the nucleic acid molecule comprises a DNA molecule. In some embodiments, the target nucleic acid molecule comprises a dsDNA molecule. In some embodiments, the target nucleic acid molecule is a gene or genome. In some embodiments, the target nucleic acid molecule (e.g., a target DNA molecule (e.g., a target gene or genome)) is intracellular. In some embodiments, the cell is in vitro, ex vivo, or in vivo. In some embodiments, the cell is a eukaryotic cell (e.g., a mammalian cell, an animal cell, a primate cell, a non-human primate cell, or a human cell). In some embodiments, the cell is a human cell. 4.11 Pharmaceutical Composition

[0379] This disclosure provides, in particular, pharmaceutical compositions comprising any one or more of the following: the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (see, for example, § 4.2); the fusion protein described herein (see, for example, § 4.3); the conjugate described herein (see, for example, § 4.3); the system described herein (see, for example, § 4.5) (or any one or more components thereof); the nucleic acid molecule described herein (see, for example, § 4.6); the carrier described herein (see, for example, § 4.7); the reaction mixture described herein (see, for example, § 4.10); the loading agent described herein (see, for example, § 4.8); and pharmaceutically acceptable excipients (see, for example, Remington's Pharmaceutical Sciences (1990) Mack Publishing Co., Easton, PA, the entire contents of which are incorporated herein by reference for all purposes).

[0380] This disclosure provides, in particular, methods for preparing pharmaceutical compositions comprising any one or more of the following: the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (see, for example, § 4.2); the fusion protein described herein (see, for example, § 4.3); the conjugate described herein (see, for example, § 4.3); the system described herein (see, for example, § 4.5) (or any one or more components thereof); the nucleic acid molecule described herein (see, for example, § 4.6); the carrier described herein (see, for example, § 4.7); the reaction mixture described herein (see, for example, § 4.10); the loading agent described herein (see, for example, § 4.8); and methods for formulating them into pharmaceutically acceptable compositions by adding one or more pharmaceutically acceptable excipients.

[0381] This document also provides pharmaceutical compositions comprising any one or more of the following: the Cas endonuclease described herein (or a functional fragment, functional variant, or domain thereof) (see, for example, § 4.2); the fusion protein described herein (see, for example, § 4.3); the conjugate described herein (see, for example, § 4.3); the system described herein (see, for example, § 4.5) (or any one or more components thereof); the nucleic acid molecule describe...

Claims

1. A Cas endonuclease (or a functional fragment, functional variant, or domain thereof) comprising an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with any Cas endonuclease shown in Table 1 or any of SEQ ID NO: 1-320.

2. The Cas endonuclease of claim 1, wherein the amino acid sequence has at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of any Cas endonuclease shown in Table 1 or any of SEQ ID NO: 1-320.

3. The Cas endonuclease of claim 1 or 2, wherein the amino acid sequence has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of any Cas endonuclease shown in Table 1 or any of SEQ ID NO: 1-320.

4. The Cas endonuclease as described in any of the preceding claims, wherein the amino acid sequence of the Cas endonuclease has less than 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% identity with the amino acid sequence of the reference Cas endonuclease shown in SEQ ID NO:

321.

5. The Cas endonuclease as described in any of the preceding claims, wherein the amino acid sequence of the Cas endonuclease has less than 90% (e.g., 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 60%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%) amino acid sequence of the reference Cas endonuclease shown in SEQ ID NO:

321. The identity of 50% (52%, 51%) and greater than 50% (e.g., 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%).

6. The Cas endonuclease as claimed in any of the preceding claims, wherein the amino acid sequence of the Cas endonuclease has less than 90% (e.g., 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%) and greater than 76% (e.g., 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%) identity with the amino acid sequence of the reference Cas endonuclease shown in SEQ ID NO:

321.

7. The Cas endonuclease as described in any of the preceding claims, having one or more of the following properties (e.g., 1, 2, 3, 4, 5 and / or 6) (or engineered to have one or more of the following properties): (a) The ability to mediate double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules; (b) The ability to mediate single-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules; (c) Cannot mediate double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules; (d) The ability to mediate single-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules, and not to mediate double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules (i.e., nicking enzyme activity). (f) DNA endonuclease activity; and / or (g) RNA-directed DNA endonuclease activity.

8. The Cas endonuclease as described in any of the preceding claims, wherein the amino acid sequence of the Cas endonuclease comprises one or more amino acid variations (e.g., substitution, deletion, addition).

9. The Cas endonuclease of claim 8, wherein one or more amino acid variations (e.g., substitution, deletion, addition) reduce or eliminate the ability of the Cas endonuclease to mediate double-strand breaks in target double-stranded nucleic acid (e.g., DNA) molecules.

10. The Cas endonuclease of claim 8 or 9, wherein the modified Cas endonuclease comprising the one or more amino acid variations (e.g., substitution, deletion, addition) has the ability to mediate single-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule and does not have the ability to mediate double-strand breaks in a target double-stranded nucleic acid (e.g., DNA) molecule (i.e., nicking enzyme activity).

11. The Cas endonuclease of any one of claims 8-10, wherein one or more amino acid variations (e.g., substitution, deletion, addition) alter the PAM nucleotide sequence recognized by the Cas endonuclease.

12. The Cas endonuclease of any one of claims 8-11, wherein the one or more amino acid variations (e.g., substitution, deletion, addition) (a) reduce the Cas endonuclease activity of the endonuclease by at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% relative to an endonuclease lacking the one or more amino acid variations (e.g., substitution, deletion, addition); or (b) increase the Cas endonuclease activity of the endonuclease by at least 1, 2, 5, 10, or 100 times relative to an endonuclease lacking the one or more amino acid variations (e.g., substitution, deletion, addition).

13. The Cas endonuclease as described in any of the preceding claims, further comprising one or more heterologous portions (e.g., heterologous proteins).

14. The Cas endonuclease of claim 13, further comprising 2, 3, 4, or 5 or more heterologous portions.

15. The Cas endonuclease of claim 13 or 14, wherein the heterologous portion is attached to the N-terminus, C-terminus, and / or interior between the N-terminus and C-terminus of the endonuclease.

16. The Cas endonuclease according to any one of claims 13-15, wherein the heterologous portion (e.g., a heterologous protein) is directly attached to the endonuclease.

17. The Cas endonuclease of any one of claims 13-15, wherein the heterologous portion (e.g., a heterologous protein) is indirectly attached to the Cas endonuclease.

18. The Cas endonuclease according to any one of claims 13-15, wherein the heterologous portion (e.g., a heterologous protein) is indirectly attached to the Cas endonuclease via a linker.

19. The Cas endonuclease according to any one of claims 13-18, wherein the heterologous portion is a peptide, protein, carbohydrate, lipid, polymer, or small molecule.

20. The Cas endonuclease according to any one of claims 13-19, wherein the heterologous portion is a nuclear localization signal (NLS), a tag, and / or a reporter gene.

21. A conjugate comprising a Cas endonuclease as claimed in any one of claims 1-20 and one or more heterologous moieties.

22. The conjugate of claim 21, wherein the heterologous portion is a protein, peptide, small molecule, nucleic acid molecule (e.g., DNA, RNA, DNA / RNA hybrid molecule), carbohydrate, lipid, or synthetic polymer.

23. The conjugate of claim 21 or 22, wherein the heterologous portion is operatively attached to the N-terminus, C-terminus, and / or interior between the N-terminus and C-terminus of the Cas endonuclease.

24. The conjugate according to any one of claims 21-23, wherein the heterologous portion is directly operatively linked to the Cas endonuclease.

25. The conjugate according to any one of claims 21-23, wherein the heterologous portion is indirectly operably linked to the Cas endonuclease.

26. The conjugate according to any one of claims 21-23, wherein the heterologous portion is indirectly operably linked to the Cas endonuclease via a linker.

27. A fusion protein comprising a Cas endonuclease as described in any one of claims 1-20 and one or more heterologous proteins.

28. The fusion protein of claim 27, wherein the heterologous protein is fused to the N-terminus, C-terminus, and / or interior between the N-terminus and C-terminus of the Cas endonuclease.

29. The fusion protein of claim 27 or 28, wherein the heterologous protein is directly fused to the Cas endonuclease.

30. The fusion protein of claim 27 or 28, wherein the heterologous protein is indirectly fused to the Cas endonuclease.

31. The fusion protein of claim 27 or 28, wherein the heterologous protein is indirectly fused to the Cas endonuclease via a peptide linker.

32. The fusion protein of any one of claims 27-31, wherein the heterologous protein exhibits polymerase (e.g., reverse transcriptase) activity, nucleobase editing activity (e.g., deaminase activity), methyltransferase activity, demethylase activity, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity or double-stranded DNA cleavage activity and nucleic acid binding activity, or any combination thereof.

33. The fusion protein according to any one of claims 27-32, wherein the heterologous protein is a polymerase.

34. The fusion protein of claim 33, wherein the polymerase has RNA-dependent DNA polymerase activity.

35. The fusion protein of claim 33 or 34, wherein the polymerase is a reverse transcriptase (or a functional fragment, functional variant, or domain thereof).

36. The fusion protein of any one of claims 33-35, wherein the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) is derived from a retrovirus or a retrotransposon.

37. The fusion protein of any one of claims 33-36, wherein the reverse transcriptase (or a functional fragment, functional variant, or domain thereof) comprises an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of the protein shown in Table 2 or any one of SEQ ID NO: 324-476.

38. The fusion protein of any one of claims 27-32, wherein the heteropeptide is a nucleobase editor.

39. The fusion protein of claim 38, wherein the nucleobase editor is a deaminase (or a functional fragment, functional variant, or domain thereof).

40. The fusion protein of claim 38 or 39, wherein the deaminase (or a functional fragment, functional variant, or domain thereof) exhibits adenosine deaminase activity and / or cytidine deaminase activity.

41. The fusion protein according to any one of claims 38-40, wherein the deaminase (or a functional fragment, functional variant, or domain thereof) comprises an amino acid sequence having at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of the protein shown in Table 3 or any one of SEQ ID NO: 477-536.

42. The fusion protein of any one of claims 38-41, wherein the nucleobase editor is fused to a base excision repair inhibitor (or a functional fragment or functional variant thereof) (e.g., uracil glycosylase inhibitor (UGI), nuclease-dead inosine-specific nuclease (dISN)).

43. A nucleic acid molecule encoding a Cas endonuclease as described in any one of claims 1-20, a conjugate as described in any one of claims 21-26, or a fusion protein as described in any one of claims 27-42.

44. The nucleic acid molecule of claim 43, wherein the nucleic acid molecule is a DNA or RNA (e.g., mRNA) molecule.

45. The nucleic acid molecule of claim 43 or 44, wherein the nucleic acid molecule is codon-optimized.

46. ​​The nucleic acid molecule of any one of claims 43-45, further comprising one or more transcriptional or translational regulatory elements (e.g., promoters, enhancers (e.g., cell or tissue-specific transcriptional regulatory elements)).

47. The nucleic acid molecule of any one of claims 43-46, further encoding one or more gRNAs (e.g., crRNA, tracrRNA, sgRNA, template RNA (e.g., as described herein)).

48. A vector comprising a nucleic acid molecule as described in any one of claims 43-47.

49. The vector of claim 48, wherein the vector is a viral vector or a non-viral vector (e.g., plasmid, microcircle).

50. The vector of claim 48 or 49, wherein the vector is a viral vector (e.g., adeno-associated virus (AAV) vector, lentiviral vector, adenovirus vector).

51. A carrier comprising a Cas endonuclease as described in any one of claims 1-20, a conjugate as described in any one of claims 21-26, a fusion protein as described in any one of claims 27-42, a nucleic acid molecule as described in any one of claims 43-47, and / or a vector as described in any one of claims 48-50.

52. The carrier of claim 51, wherein the carrier is a nanoparticle, a polymer, a virus (e.g., a recombinant virus), a virus-like particle, a virion, a fusion, a vesicle, or a lipid-based carrier.

53. The carrier as described in claim 51 or 52, wherein the carrier is a recombinant virus (e.g., adeno-associated virus (AAV), lentivirus, adenovirus).

54. The carrier according to any one of claims 51-53, wherein the carrier is a lipid-based carrier.

55. The carrier of claim 54, wherein the lipid-based carrier is a lipid nanoparticle (LNP), liposome, lipid complex, nanoliposome, exosome, or micelle.

56. The carrier according to any one of claims 51-55, further comprising one or more gRNAs (e.g., crRNA, tracrRNA, sgRNA, template RNA (e.g., as described herein)).

57. A reaction mixture comprising (a) a cell (e.g., a cell containing a target nucleic acid molecule) or a target nucleic acid molecule; and (b) a Cas endonuclease as described in any one of claims 1-20, a conjugate as described in any one of claims 21-26, a fusion protein as described in any one of claims 27-42, a nucleic acid molecule as described in any one of claims 43-47, a carrier as described in any one of claims 48-50, a loading agent as described in any one of claims 51-56, and / or a pharmaceutical composition as described in claim 59.

58. A cell comprising the Cas endonuclease as described in any one of claims 1-20, the conjugate as described in any one of claims 21-26, the fusion protein as described in any one of claims 27-42, the nucleic acid molecule as described in any one of claims 43-47, the vector as described in any one of claims 48-50, the loading agent as described in any one of claims 51-56, the reaction mixture as described in claim 58, and / or the pharmaceutical composition as described in claim 59.

59. A pharmaceutical composition comprising a Cas endonuclease as described in any one of claims 1-20, a conjugate as described in any one of claims 21-26, a fusion protein as described in any one of claims 27-42, a nucleic acid molecule as described in any one of claims 43-47, a carrier as described in any one of claims 48-50, a loading agent as described in any one of claims 51-56, a reaction mixture as described in claim 57, and / or a cell as described in claim 58, and a pharmaceutically acceptable excipient.

60. A kit comprising the Cas endonuclease as described in any one of claims 1-20, the conjugate as described in any one of claims 21-26, the fusion protein as described in any one of claims 27-42, the nucleic acid molecule as described in any one of claims 43-47, the vector as described in any one of claims 48-50, the loading agent as described in any one of claims 51-56, the reaction mixture as described in claim 57, the cells as described in claim 58, and / or the pharmaceutical composition as described in claim 59; and optionally instructions for use with any one or more of the foregoing.

61. A system for modifying a target nucleic acid (e.g., DNA) molecule, said system comprising: (a) The Cas endonuclease of any one of claims 1-20, the conjugate of any one of claims 21-26, the fusion protein of any one of claims 27-42, the nucleic acid molecule of any one of claims 43-47, the vector of any one of claims 48-50, the loading agent of any one of claims 51-56, the reaction mixture of any one of claims 57, the cell of any one of claims 58, and / or the pharmaceutical composition of any one of claims 59; and (b) A first gRNA (e.g., crRNA and tracrRNA; sgRNA; pegRNA, template RNA (e.g., as described herein)) or a nucleic acid (e.g., DNA) molecule encoding the first gRNA (e.g., crRNA and tracrRNA; sgRNA; template RNA (e.g., as described herein)).

62. The system of claim 61, wherein it has one or more of the following features: (a) The Cas endonuclease of the system is capable of binding to the first gRNA; (b) The Cas endonuclease of the system is capable of forming breaks in target nucleic acid (e.g., DNA (e.g., dsDNA)) molecules; (c) The Cas endonuclease of the system is capable of forming single-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules; (d) The Cas endonuclease of the system is capable of forming single-strand breaks in the modified strand (as defined herein) of the target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule. (e) The Cas endonuclease of the system is capable of forming double-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules; (f) The Cas endonuclease of the system cannot form double-strand breaks in the target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; (g) The Cas endonuclease of the system is capable of forming single-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules and is not capable of forming double-strand breaks in target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecules; (h) The Cas endonuclease of the system is capable of forming single-strand breaks in the modified strand (as defined herein) of the target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule, and is unable to form double-strand breaks in the target double-stranded nucleic acid (e.g., DNA (e.g., dsDNA)) molecule; and / or (i) The system is capable of editing target nucleic acid (e.g., DNA) molecules (e.g., target double-stranded DNA molecules) (e.g., mediating the addition of one or more nucleotides to the target nucleic acid, the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid).

63. The system of claim 61 or 62, wherein the system is capable of editing a target nucleic acid (e.g., DNA) molecule (e.g., a target double-stranded DNA molecule) (e.g., mediating the addition of one or more nucleotides to the target nucleic acid, the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid).

64. The system of any one of claims 61-63, wherein, relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., the reference Cas endonuclease shown in SEQ ID NO: 321), the system is capable of editing target nucleic acid (e.g., DNA) molecules (e.g., target double-stranded DNA molecules) with increased efficiency (e.g., mediating the addition of one or more nucleotides to the target nucleic acid, the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid).

65. The system of any one of claims 61-64, wherein, relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., the reference Cas endonuclease shown in SEQ ID NO: 321), the system is capable of editing target nucleic acid (e.g., DNA) molecules (e.g., target double-stranded DNA molecules) with an efficiency increase of at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200% (e.g., mediating the addition of one or more nucleotides to the target nucleic acid, the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid).

66. The system of any one of claims 61-65, wherein, relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., the reference Cas endonuclease shown in SEQ ID NO: 321), the system is capable of editing target nucleic acid (e.g., DNA) molecules (e.g., target double-stranded DNA molecules) with an efficiency increase of at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100% (e.g., mediating the addition of one or more nucleotides to the target nucleic acid, the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid).

67. The system of any one of claims 61-66, wherein, relative to a reference system (e.g., a reference system comprising a reference Cas endonuclease (e.g., the reference Cas endonuclease shown in SEQ ID NO: 321), the system is capable of increasing by about 30%-200%, 40%-200%, 50%-200%, 60%-200%, 70%-200%, 80%-200%, 90%-200%, 100%-200%, 150%-200%, 30%-150%, 40%-150%, 50%-150%, 60%-150%, 70%-150%, 80%-150%, 90%-150%, 80%-150%, 90%-150%, 90%-150%, 100%-200%, 150%-200%, 30%-150%, 40%-150%, 50%-150%, 60%-150%, 70%-150%, 80%-150%, 9 ... Editing target nucleic acid (e.g., DNA) molecules (e.g., target double-stranded DNA molecules) with efficiencies of 0%-150%, 100%-150%, 30%-100%, 40%-100%, 50%-100%, 60%-100%, 70%-100%, 80%-100%, or 90%-100% (e.g., mediating the addition of one or more nucleotides to the target nucleic acid, the deletion of one or more nucleotides from the target nucleic acid, or the substitution of one or more nucleotides in the target nucleic acid).

68. The system of any one of claims 61-67, wherein the target nucleic acid molecule is a DNA molecule.

69. The system of any one of claims 61-68, wherein the target nucleic acid molecule is a double-stranded DNA (dsDNA) molecule.

70. The system of any one of claims 61-69, wherein a portion of the nucleotide sequence of the unmodified strand (as defined herein) of the target dsDNA molecule is complementary to at least a portion of the nucleotide sequence of the first gRNA.

71. The system of any one of claims 61-70, wherein the target nucleic acid molecule is within the genome of a cell (e.g., a eukaryotic cell) (e.g., within a subject (e.g., a human subject), or a plant).

72. The system of any one of claims 61-71, wherein (b) comprises the first gRNA (e.g., crRNA and tracrRNA; or template RNA (e.g., as described herein)).

73. The system of any one of claims 61-72, wherein (b) comprises a nucleic acid (e.g., DNA) molecule encoding the first gRNA.

74. The system of any one of claims 61-73, wherein at least a portion of the nucleotide sequence of the first gRNA is complementary to a portion of the nucleotide sequence of the target nucleic acid molecule (e.g., a gene).

75. The system of any one of claims 61-74, wherein at least a portion of the nucleotide sequence of the first gRNA is complementary to a portion of the nucleotide sequence of the unmodified strand (as defined herein) of a dsDNA target nucleic acid molecule (e.g., a gene).

76. The system of any one of claims 61-75, wherein at least a portion of the nucleotide sequence of the first gRNA binds to a portion of the nucleotide sequence of the unmodified strand (as defined herein) of a dsDNA target nucleic acid molecule (e.g., a gene).

77. The system of any one of claims 61-76, wherein the first gRNA comprises sgRNA (e.g., a single sgRNA, or multiple different sgRNAs).

78. The system of any one of claims 61-77, wherein the first gRNA comprises crRNA (e.g., a single crRNA, multiple different crRNAs) and tracrRNA (e.g., a single tracrRNA, multiple different tracrRNAs), wherein the crRNA and the tracrRNA are on separate RNA nucleic acid molecules (or encoded by separate nucleic acid (e.g., DNA) molecules).

79. The system of any one of claims 61-78, wherein the first gRNA comprises a template RNA (e.g., a single template RNA, multiple different template RNAs), the template RNA comprising (e.g., from 5' to 3') crRNA, tracrRNA, a heterologous object sequence, and a 3' target homologous domain.

80. The system of any one of claims 61-79, wherein the template RNA further comprises a sequence for binding a polymerase (e.g., a reverse transcriptase, such as the reverse transcriptase of the fusion protein of any one of claims 33-37).

81. The system of any one of claims 61-80, wherein the template RNA comprises (e.g., from 5' to 3') crRNA, tracrRNA, a sequence of a binding polymerase (e.g., a reverse transcriptase, such as the reverse transcriptase of the fusion protein of any one of claims 33-37), a heterologous object sequence, and a 3' target homologous domain.

82. The system of any one of claims 61-81, wherein the first gRNA comprises one or more nucleotides, the one or more nucleotides comprising one or more chemical modifications (e.g., base, ribose and / or internucleotide linking chemical modifications) (i.e., modified nucleotides).

83. The system of claim 82, wherein the modified nucleotide comprises 2'-O-methyl (2'-OMe); 2'-O-methoxyethyl (2'-O-MOE); 2'-deoxy-2'-fluoro (2'-F); 2'-arabinose-fluorine (2'-Ara-F); 2'-O-benzyl; 2'-O-methyl-4-pyridine (2'-O-methyl-4-pyridine (2'-O-CH2Py(4)); 2'-F-4'-Cα-OMe; or 2',4'-di-Cα-OMe, 2'-O-methyl-3'-thioPACE and / or S-restricted ethyl (cEt).

84. The system of claim 81 or 82, wherein the modified nucleotide comprises chemically modified internucleotide (or internucleotide) bonds.

85. The system of any one of claims 81-84, wherein the modified internucleotide (or internucleotide) linkage comprises thiophosphate (e.g., chiral thiophosphate), dithiophosphate, phosphate triester, aminoalkyl phosphate triester, alkyl (e.g., methyl)phosphonate (e.g., 3'-alkylene phosphonate, chiral phosphonate), hypophosphonate, aminophosphate (e.g., 3'-aminoaminophosphate, aminoalkylaminophosphate), thiocarbonylaminophosphate, thiocarbonylalkylphosphonate, thiocarbonylalkyl phosphate triester, or boroalkyl phosphate.

86. The system of any one of claims 61-85, wherein the first gRNA (e.g., template RNA, sgRNA) comprises a nucleic acid molecule containing a toe loop, hairpin, stem loop, pseudoknot (e.g., Mpknot1 portion), aptamer, G-quadruplex, tRNA, riboswitch, or ribozyme.

87. The system of any one of claims 61-86, wherein the first gRNA (e.g., template RNA, sgRNA) is a pseudoknot (e.g., Mpknot1 portion).

88. The system of any one of claims 61-87, further comprising the endonuclease guiding the system to form a second gRNA (or a nucleic acid (e.g., DNA) molecule encoding the gRNA) in the unedited strand of the target dsDNA molecule.

89. The system of any one of claims 61-88, wherein at least a portion of the nucleotide sequence of the second gRNA is complementary to a portion of the nucleotide sequence of the edited strand (as defined herein) of the dsDNA target nucleic acid molecule.

90. The system of any one of claims 61-89, wherein at least a portion of the nucleotide sequence of the second gRNA binds to a portion of the nucleotide sequence of the edited strand (as defined herein) of the dsDNA target nucleic acid molecule.

91. The system of any one of claims 61-90, wherein the second gRNA is present on the same nucleic acid molecule as the first gRNA (or the nucleic acid (e.g., DNA) molecule encoding the second gRNA is present on the same nucleic acid (e.g., DNA) molecule encoding the first gRNA).

92. The system of any one of claims 61-91, wherein the second gRNA is present on a nucleic acid molecule different from the first gRNA (or the nucleic acid (e.g., DNA) molecule encoding the second gRNA is present on a different nucleic acid (e.g., DNA) molecule encoding the first gRNA).

93. The system of any one of claims 61-92, further comprising a donor template nucleic acid (e.g., DNA) molecule (e.g., as defined herein).

94. A system for modifying dsDNA molecules, comprising: (a) A fusion protein as described in any one of claims 33-37 or a nucleic acid molecule (e.g., DNA, RNA molecule) encoding a fusion protein as described in any one of claims 33-37; and (b) A template RNA (e.g., a single template RNA, multiple different template RNAs) comprising (e.g., from 5' to 3') crRNA, tracrRNA, a heterologous object sequence and a 3' target homologous domain; or a nucleic acid molecule (e.g., a DNA molecule) encoding the template RNA.

95. A nucleic acid molecule encoding the system described in any one of claims 61-94.

96. The nucleic acid molecule of claim 95, wherein the nucleic acid molecule is a DNA or RNA (e.g., mRNA) molecule.

97. The nucleic acid molecule of claim 95 or 96, wherein the nucleic acid molecule is codon-optimized.

98. The nucleic acid molecule of any one of claims 95-97, further comprising one or more transcriptional or translational regulatory elements (e.g., promoters, enhancers (e.g., cell or tissue-specific transcriptional regulatory elements)).

99. A vector comprising a nucleic acid molecule as described in any one of claims 95-98.

100. The vector of claim 99, wherein the vector is a viral vector or a non-viral vector (e.g., plasmid, microcircle).

101. The vector as claimed in claim 99 or 100, wherein the vector is a viral vector (e.g., adeno-associated virus (AAV) vector, lentiviral vector, adenovirus vector).

102. A carrier comprising the system as described in any one of claims 61-94, the nucleic acid molecule as described in any one of claims 95-98, and / or the vector as described in any one of claims 99-101.

103. The carrier of claim 102, wherein the carrier is a nanoparticle, a polymer, a virus (e.g., a recombinant virus), a virus-like particle, a virion, a fusion, a vesicle, or a lipid-based carrier.

104. The carrier of any one of claims 102 or 103, wherein the carrier is a recombinant virus (e.g., adeno-associated virus (AAV), lentivirus, adenovirus).

105. The carrier as claimed in claim 102 or 104, wherein the carrier is nanoparticles.

106. The carrier of any one of claims 102, 104 or 105, wherein the carrier is a lipid-based carrier.

107. The carrier of claim 106, wherein the lipid-based carrier is a lipid nanoparticle (LNP), liposome, lipid complex, nanoliposome, exosome, or micelle.

108. The carrier according to any one of claims 102-107, further comprising one or more gRNAs (e.g., crRNA, tracrRNA, sgRNA, template RNA (e.g., as described herein)).

109. A reaction mixture comprising (a) a cell (e.g., a cell containing a target nucleic acid molecule) or a target nucleic acid molecule; and (b) a system as described in any one of claims 61-94, a nucleic acid molecule as described in any one of claims 95-98, a carrier as described in any one of claims 99-101, and / or a loading agent as described in any one of claims 102-108.

110. A cell comprising the system as described in any one of claims 61-94, the nucleic acid molecule as described in any one of claims 95-98, the vector as described in any one of claims 99-101, the carrier as described in any one of claims 102-108, and / or the reaction mixture as described in claim 109.

111. A pharmaceutical composition comprising the system as described in any one of claims 61-94, a nucleic acid molecule as described in any one of claims 95-98, a carrier as described in any one of claims 99-101, a loading agent as described in any one of claims 102-108, a reaction mixture as described in claim 109, and / or cells as described in claim 110, and a pharmaceutically acceptable excipient.

112. A kit comprising the system of any one of claims 61-94, the nucleic acid molecule of any one of claims 85-88, the vector of any one of claims 99-101, the loading agent of any one of claims 102-108, the reaction mixture of any one of claims 109, the cells of any one of claims 110, and / or the pharmaceutical composition of any one of claims 111; and instructions for use with any one or more of the foregoing.

113. A method for delivering a Cas endonuclease, fusion protein, conjugate, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition to a cell, the method comprising introducing into the cell the Cas endonuclease as claimed in any one of claims 1-20, the conjugate as claimed in any one of claims 21-26, the fusion protein as claimed in any one of claims 27-42, the system as claimed in any one of claims 61-94, the nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, the vector as claimed in any one of claims 48-50 or 99-101, the loading agent as claimed in any one of claims 51-56 or 102-108, the reaction mixture as claimed in any one of claims 57 or 109, the cell as claimed in any one of claims 58 or 110, or the pharmaceutical composition as claimed in any one of claims 59 or 111, thereby delivering the Cas endonuclease, fusion protein, conjugate, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition to the cell.

114. The method of claim 113, wherein the cells are in vitro, ex vivo, or in vivo.

115. The method of claim 113 or 114, wherein the cell is euploid, not immortalized, part of a tissue, part of an organism, a primary cell, non-dividing, haploid (e.g., germline cell), non-cancerous polyploid cell, or derived from a subject with a genetic disease.

116. The method of any one of claims 113-115, wherein the cells are in a subject (e.g., a human subject).

117. The method of any one of claims 113-116, wherein the cells are in a human subject.

118. A method for delivering a Cas endonuclease, fusion protein, conjugate, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition to a cell, the method comprising the Cas endonuclease as claimed in any one of claims 1-20, the conjugate as claimed in any one of claims 21-26, the fusion protein as claimed in any one of claims 27-42, the system as claimed in any one of claims 61-94, the nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, the vector as claimed in any one of claims 48-50 or 99-101, the loading agent as claimed in any one of claims 51-56 or 102-108, the reaction mixture as claimed in any one of claims 57 or 109, the cell as claimed in any one of claims 58 or 110, or the pharmaceutical composition as claimed in any one of claims 59 or 111, thereby delivering the Cas endonuclease, fusion protein, conjugate, system, nucleic acid molecule, vector, loading agent, reaction mixture, cell, or pharmaceutical composition to the subject (e.g., a human subject).

119. A method of cleaving a target site in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))), the method comprising contacting the cell with a Cas endonuclease as claimed in any one of claims 1-20, a conjugate as claimed in any one of claims 21-26, a fusion protein as claimed in any one of claims 27-42, a system as claimed in any one of claims 61-94, a nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, a vector as claimed in any one of claims 48-50 or 99-101, a carrier as claimed in any one of claims 51-56 or 102-108, a reaction mixture as claimed in any one of claims 57 or 109, a cell as claimed in any one of claims 58 or 110, or a drug combination as claimed in any one of claims 59 or 111, thereby cleaving the target site in the target nucleic acid (e.g., DNA) molecule.

120. A method of editing a target site in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))), the method comprising contacting the cell with a Cas endonuclease as claimed in any one of claims 1-20, a conjugate as claimed in any one of claims 21-26, a fusion protein as claimed in any one of claims 27-42, a system as claimed in any one of claims 61-94, a nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, a vector as claimed in any one of claims 48-50 or 99-101, a carrier as claimed in any one of claims 51-56 or 102-108, a reaction mixture as claimed in any one of claims 57 or 109, a cell as claimed in any one of claims 58 or 110, or a pharmaceutical composition as claimed in any one of claims 59 or 111, thereby editing the target site in the target nucleic acid (e.g., DNA) molecule.

121. A method for editing a target site in genomic dsDNA of a cell, the method comprising contacting a Cas endonuclease as described in any one of claims 1-20, a conjugate as described in any one of claims 21-26, a fusion protein as described in any one of claims 27-42, a system as described in any one of claims 61-94, a nucleic acid molecule as described in any one of claims 43-47 or 95-98, a vector as described in any one of claims 48-50 or 99-101, a carrier as described in any one of claims 51-56 or 102-108, a reaction mixture as described in any one of claims 57 or 109, a cell as described in any one of claims 58 or 110, or a pharmaceutical composition as described in any one of claims 59 or 111, thereby editing the target site in the genomic DNA of the cell.

122. A method for editing a target site in a dsDNA molecule (e.g., genomic dsDNA (e.g., in a cell)), the method comprising: Contact the dsDNA molecule with the following (a) The fusion protein of any one of claims 33-37 (or the nucleic acid molecule (e.g., DNA, RNA nucleic acid molecule) encoding the fusion protein of any one of claims 33-37), and (b) A template RNA (e.g., a single template RNA, multiple different template RNAs) comprising (e.g., from 5' to 3') crRNA, tracrRNA, a heterologous object sequence and a 3' target homologous domain, thereby modifying the target site in the dsDNA molecule (or a nucleic acid molecule encoding the template RNA (e.g., a DNA nucleic acid molecule)) to edit the target site in the dsDNA molecule (e.g., genomic dsDNA (e.g., in a cell)).

123. The method of any one of claims 119-122, wherein the nucleic acid molecule is in a cell (e.g., a eukaryotic cell).

124. The method of claim 123, wherein the cells are in vitro, ex vivo, or in vivo.

125. The method of claim 123 or 124, wherein the cells are in a subject (e.g., a human subject).

126. The method of any one of claims 123-125, wherein the cells are in a human subject.

127. The method of any one of claims 120-126, wherein the editing comprises adding one or more nucleotides to the target site of the genomic dsDNA in the cell, deleting one or more nucleotides from the target site, or replacing one or more nucleotides in the target site.

128. The method of any one of claims 120-127, wherein the editing comprises adding one or more nucleotides to the target site of the target nucleic acid molecule, deleting one or more nucleotides from the target site, or replacing one or more nucleotides in the target site.

129. The method of any one of claims 127-128, wherein the addition comprises adding about 1-500, 1-3200, 1-300, 1-200, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 nucleotides to the target site.

130. The method of any one of claims 127-129, wherein the deletion comprises the deletion of about 1-500, 1-3200, 1-300, 1-200, 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-320, 1-30, 1-20, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 nucleotides at the target site.

131. A method of treating, improving, or preventing a disease in a subject (e.g., a human subject), the method comprising administering a Cas endonuclease as claimed in any one of claims 1-20, a conjugate as claimed in any one of claims 21-26, a fusion protein as claimed in any one of claims 27-42, a system as claimed in any one of claims 61-94, a nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, a vector as claimed in any one of claims 48-50 or 99-101, a loading agent as claimed in any one of claims 51-56 or 102-108, a reaction mixture as claimed in any one of claims 57 or 109, a cell as claimed in any one of claims 58 or 110, or a pharmaceutical composition as claimed in any one of claims 59 or 111, thereby treating, improving, or preventing the disease in the subject.

132. The method of claim 131, wherein the disease is related to a genetic defect.

133. The method of claim 132, wherein the gRNA of the system is capable of targeting the endonuclease to the site of the genetic defect.

134. The method of claim 132 or 133, wherein the genetic defect includes gene duplication, gene deletion, or gene mutation.

135. The method of any one of claims 132-134, wherein the application results in the correction of the genetic defect.

136. The method of any one of claims 131-135, wherein the subject is a human subject.

137. The Cas endonuclease of any one of claims 1-20, the conjugate of any one of claims 21-26, the fusion protein of any one of claims 27-42, the system of any one of claims 61-94, the nucleic acid molecule of any one of claims 43-47 or 95-98, the vector of any one of claims 48-50 or 99-101, the carrier of any one of claims 51-56 or 102-108, the reaction mixture of any one of claims 57 or 109, the cell of any one of claims 58 or 110, or the pharmaceutical composition of any one of claims 59 or 111, for use at a target site in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))) in a subject in need.

138. Use of any Cas endonuclease as claimed in any one of claims 1-20, any conjugate as claimed in any one of claims 21-26, any fusion protein as claimed in any one of claims 27-42, any system as claimed in any one of claims 61-94, any nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, any vector as claimed in any one of claims 48-50 or 99-101, any carrier as claimed in any one of claims 51-56 or 102-108, any reaction mixture as claimed in any one of claims 57 or 109, any cell as claimed in any one of claims 58 or 110, or any pharmaceutical composition as claimed in any one of claims 59 or 111, for the manufacture of a medicament for cleaving a target site in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))) in a subject in need.

139. The Cas endonuclease of any one of claims 1-20, the conjugate of any one of claims 21-26, the fusion protein of any one of claims 27-42, the system of any one of claims 61-94, the nucleic acid molecule of any one of claims 43-47 or 95-98, the vector of any one of claims 48-50 or 99-101, the carrier of any one of claims 51-56 or 102-108, the reaction mixture of any one of claims 57 or 109, the cell of any one of claims 58 or 110, or the pharmaceutical composition of any one of claims 59 or 111, for use at a target site in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))) in a subject in need of editing.

140. Use of any Cas endonuclease as claimed in any one of claims 1-20, any conjugate as claimed in any one of claims 21-26, any fusion protein as claimed in any one of claims 27-42, any system as claimed in any one of claims 61-94, any nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, any vector as claimed in any one of claims 48-50 or 99-101, any carrier as claimed in any one of claims 51-56 or 102-108, any reaction mixture as claimed in any one of claims 57 or 109, any cell as claimed in any one of claims 58 or 110, or any pharmaceutical composition as claimed in any one of claims 59 or 111, for the manufacture of a medicament for editing a target site in a target nucleic acid (e.g., DNA) molecule (e.g., a double-stranded target nucleic acid sequence (e.g., dsDNA (e.g., genomic dsDNA))) in a subject in need.

141. The Cas endonuclease of any one of claims 1-20, the conjugate of any one of claims 21-26, the fusion protein of any one of claims 27-42, the system of any one of claims 61-94, the nucleic acid molecule of any one of claims 43-47 or 95-98, the vector of any one of claims 48-50 or 99-101, the carrier of any one of claims 51-56 or 102-108, the reaction mixture of any one of claims 57 or 109, the cell of any one of claims 58 or 110, or the pharmaceutical composition of any one of claims 59 or 111, used as a drug.

142. The Cas endonuclease of any one of claims 1-20, the conjugate of any one of claims 21-26, the fusion protein of any one of claims 27-42, the system of any one of claims 61-94, the nucleic acid molecule of any one of claims 43-47 or 95-98, the vector of any one of claims 48-50 or 99-101, the loading agent of any one of claims 51-56 or 102-108, the reaction mixture of any one of claims 57 or 109, the cell of any one of claims 58 or 110, or the pharmaceutical composition of any one of claims 59 or 111, for use in treating a disease (e.g., a disease related to a genetic defect) in a subject in need.

143. Use of any Cas endonuclease as claimed in any one of claims 1-20, any conjugate as claimed in any one of claims 21-26, any fusion protein as claimed in any one of claims 27-42, any system as claimed in any one of claims 61-94, any nucleic acid molecule as claimed in any one of claims 43-47 or 95-98, any vector as claimed in any one of claims 48-50 or 99-101, any carrier as claimed in any one of claims 51-56 or 102-108, any reaction mixture as claimed in any one of claims 57 or 109, any cell as claimed in any one of claims 58 or 110, or any pharmaceutical composition as claimed in any one of claims 59 or 111, for the manufacture of a medicament for treating a disease (e.g., a disease related to a genetic defect) in a subject in need.

Citation Information

Patent Citations

  • Amino acid-, peptide- and polypeptide-lipids, isomers, compositions, and uses thereof

    US10086013B2

  • Lipids and lipid nanoparticle formulations for delivery of nucleic acids

    US10221127B2

  • Anellovirus compositions and methods of use

    US11446344B1

  • Lipid-based formulations

    US20030077829A1

  • Polyethyleneglycol-modified lipid compounds and uses thereof

    US20050175682A1