Methods and compositions for regulating the genome
Patent Information
- Application Number
- JP2024515072
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-08
- Filing Date
- 2022-09-07
- Publication Date
- 2025-09-17
AI Technical Summary
Existing methods for integrating nucleic acids into genomes lack site specificity and efficiency, particularly for long sequences, and require multiple steps or rely on host repair pathways.
Development of recombinant polypeptides with specific domains, such as polymerase and Cas nickase, linked by a linker, and template nucleic acids to facilitate targeted insertion, deletion, or modification of genomic sequences using a genetically engineered system.
Enables precise and efficient insertion, deletion, or modification of genomic sequences, including long sequences, with improved site specificity and reduced reliance on host machinery.
Smart Images

Figure 00000297_0000 
Figure 00000297_0001 
Figure 00000297_0002
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 241,931, filed September 8, 2021, the entire contents of each of which are incorporated herein by reference.
[0002] Sequence Listing This application contains a Sequence Listing, which has been submitted electronically in XML format and is incorporated herein by reference in its entirety. November 16 The XML copy created in is named V2065-7020WO_SL.xml, 11,405,833 The size in bytes. [Background technology]
[0003] Integration of a nucleic acid of interest into a genome occurs at low frequency in the absence of specialized proteins to facilitate the insertion event and has little site specificity. Some existing methods, such as CRISPR / Cas9, are more suitable for small edits that rely on host repair pathways and are less effective at integrating long sequences. Other existing methods, such as Cre / loxP, require a first step of inserting a loxP site into the genome, followed by a second step of inserting a sequence of interest into the loxP site. There is a need in the art for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, modifying, or deleting a sequence of interest in a genome. Summary of the Invention [Means for solving the problem]
[0004] The present disclosure relates to novel compositions, systems, and methods for modifying the genome of one or more locations in a host cell, tissue, or subject in vivo or in vitro. In particular, the present disclosure features compositions, systems, and methods for inserting, modifying, or deleting a sequence of interest into a host genome. For example, the disclosure provides systems capable of modulating gene activity (e.g., inserting, altering, or deleting a sequence of interest), and methods for treating disease by administering one or more such systems to modify genomic sequences at nucleotides to correct disease-causing pathogenic mutations.
[0005] The composition or method configuration may include one or more of the embodiments listed below.
[0006] 1. A recombinant polypeptide comprising: A DNA binding domain (DBD) that binds to a target nucleic acid sequence and a polymerase (Pol) domain of Table 1 or 23, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto; wherein the DBD is heterologous to the Pol domain; and Linker between the Pol domain and DBD A recombinant polypeptide comprising:
[0007] 2. A recombinant polypeptide comprising: a Cas domain (e.g., a Cas nickase domain, e.g., a Cas9 nickase domain); a polymerase (Pol) domain of Table 1 or 23, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto, wherein the Pol domain is C-terminal to the Cas domain; and Linker between the Pol and Cas domains A recombinant polypeptide comprising:
[0008] 3. The recombinant polypeptide of embodiment 1 or 2, wherein the linker has a sequence from Table 6 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto.
[0009] 4. The recombinant polypeptide of the previous embodiment, wherein the Pol domain has a sequence with at least 90% identity to a Pol domain of Table 1 or 23.
[0010] 5. The recombinant polypeptide of any of the previous embodiments, wherein the Pol domain has a sequence with at least 95% identity to a Pol domain of Table 1 or 23.
[0011] 6. The recombinant polypeptide of any of the previous embodiments, wherein the Pol domain has a sequence with at least 98% identity to a Pol domain of Table 1 or 23.
[0012] 7. The recombinant polypeptide of any of the previous embodiments, wherein the Pol domain has a sequence with at least 99% identity to a Pol domain of Table 1 or 23.
[0013] 8. The recombinant polypeptide of any of the previous embodiments, wherein the Pol domain has a sequence with 100% identity to a Pol domain of Table 1 or 23.
[0014] 9. The recombinant polypeptide of any of the previous embodiments, wherein the linker has a sequence with at least 90% identity to a linker sequence from Table 6.
[0015] 10. The recombinant polypeptide of any of the previous embodiments, wherein the linker has a sequence with at least 95% identity to a linker sequence from Table 6.
[0016] 11. The recombinant polypeptide of any of the previous embodiments, wherein the linker has a sequence with at least 97% identity to a linker sequence from Table 6.
[0017] 12. The recombinant polypeptide of any of the previous embodiments, wherein the linker has a sequence with 100% identity to a linker sequence from Table 6.
[0018] 13. The recombinant polypeptide of any of the previous embodiments, wherein the Cas domain comprises a sequence in Table 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto.
[0019] 14. The recombinant polypeptide of any of the preceding embodiments, wherein the Cas domain is a Cas nickase domain.
[0020] 15. The recombinant polypeptide of any of the preceding embodiments, wherein the Cas domain is a Cas9 nickase domain.
[0021] 16. The recombinant polypeptide of any of the preceding embodiments, wherein the Cas domain comprises an N863A mutation.
[0022] 17. The recombinant polypeptide of any of the previous embodiments, which comprises an NLS, such as two NLSs.
[0023] 18. The recombinant polypeptide of any of the previous embodiments, comprising an NLS N-terminal to the Cas9 domain.
[0024] 19. The recombinant polypeptide of any of the previous embodiments, comprising an NLS C-terminal to the Pol domain.
[0025] 20. The recombinant polypeptide of any of the previous embodiments, comprising a first NLS that is N-terminal to the Cas9 domain and a second NLS that is C-terminal to the Pol domain.
[0026] 21. A nucleic acid (e.g., DNA or RNA, e.g., mRNA) encoding the recombinant polypeptide of any of the preceding embodiments.
[0027] 22. A cell comprising a recombinant polypeptide of any one of embodiments 1 to 20 or a nucleic acid of embodiment 21.
[0028] 23. A system comprising: i) a recombinant polypeptide according to any one of embodiments 1 to 20, and ii) a template nucleic acid (e.g., a template RNA), a) a gRNA spacer complementary to a portion of the target nucleic acid sequence; b) gRNA scaffold binding to the Cas domain of the recombinant polypeptide; c) a heterologous sequence of interest; and d) Primer binding site sequence (PBS sequence) A template nucleic acid (e.g., template RNA) comprising: A system including:
[0029] 24. The system of embodiment 23, wherein the template nucleic acid comprises RNA.
[0030] 25. The system of embodiment 23 or 24, wherein the template nucleic acid comprises DNA.
[0031] 26. The system of any of embodiments 23 to 25, wherein the gRNA spacer and the gRNA scaffold comprise RNA.
[0032] 27. The system of any of embodiments 23-26, wherein the heterologous target sequence comprises DNA and the PBS sequence comprises RNA.
[0033] 28. The system of any of embodiments 23-26, wherein the heterologous sequence of interest and the PBS sequence comprise DNA.
[0034] 29. A method for modifying a target nucleic acid in a cell (e.g., a human cell), comprising contacting the cell with the system of any of embodiments 23 to 28 or a nucleic acid encoding same, thereby modifying the target nucleic acid.
[0035] 30. A method of treating a subject having a disease or condition associated with a genetic defect, comprising: administering to a subject the system, polypeptide, template RNA or DNA encoding same of any of the preceding embodiments, thereby treating the subject with a disease or condition associated with a genetic defect. A method comprising:
[0036] 31. The method of embodiment 30, wherein the disease or condition associated with a genetic defect is an indication listed in any of Tables 12-15, and / or the genetic defect is a defect in a gene listed in any of Tables 12-15.
[0037] 32. The method of embodiment 30 or 31, wherein the subject is a human patient.
[0038] In one aspect, the disclosure relates to a system for recombining genes, the system comprising: (a) a nucleic acid encoding a recombination polypeptide capable of target-primed reverse transcription, the polypeptide comprising (i) a polymerase (Pol) domain and (ii) a Cas9 nickase that binds to DNA and has endonuclease activity; and (b) a template RNA, DNA, or hybrid having both ribonucleotide and deoxyribonucleotide residues in the same strand, the template RNA, DNA, or hybrid comprising: (i) a gRNA spacer complementary to a first portion of a first portion of a target gene; (ii) a gRNA scaffold that binds to the polypeptide; (iii) a heterologous sequence of interest comprising a mutation region for recombining the gene; and (iv) a primer binding site (PBS) sequence at the 3' end of the template RNA that comprises at least 3, 4, 5, 6, 7, or 8 bases of 100% homology to the target DNA strand.
[0039] The gRNA spacer may comprise at least 15 bases at the 5' end of the template RNA that are 100% homologous to the target DNA. The template RNA may further comprise a PBS sequence that comprises at least 5 bases that are at least 80% homologous to the target DNA strand. The template RNA may comprise one or more chemical modifications.
[0040] The domains of the recombinant polypeptide may be linked by a peptide linker. The polypeptide may contain one or more peptide linkers. The recombinant polypeptide may further contain a nuclear localization signal. The polypeptide may contain two or more nuclear localization signals, for example, multiple adjacent nuclear localization signals or one or more nuclear localization signals in different regions of the polypeptide, for example, one or more nuclear localization signals at the N-terminus of the polypeptide and one or more nuclear localization signals at the C-terminus of the polypeptide. The nucleic acid encoding the recombinant polypeptide may encode one or more intein domains.
[0041] Introduction of the system into a target cell can result in the insertion of at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 base pairs of exogenous DNA. Introduction of the system into a target cell can result in a deletion, where the deletion is less than 2, 3, 4, 5, 10, 50, or 100 base pairs of genomic DNA upstream or downstream of the insertion. Introduction of the system into a target cell can result in a substitution, for example, of 1, 2, or 3 nucleotides, for example, consecutive nucleotides.
[0042] The heterologous sequence of interest can be at least 5, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600, or 700 base pairs.
[0043] In one aspect, the present disclosure relates to a pharmaceutical composition comprising the above-described system and a pharmaceutically acceptable excipient or carrier, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle. In one aspect, the present disclosure relates to a pharmaceutical composition comprising the above-described system and a plurality of pharmaceutically acceptable excipients or carriers, wherein the pharmaceutically acceptable excipients or carriers are selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle, for example, wherein the above-described system is delivered by two different excipients or carriers, for example, two lipid nanoparticles, two viral vectors, or one lipid nanoparticle and one viral vector. The viral vector may be an adeno-associated virus (AAV).
[0044] In one aspect, the present disclosure relates to a host cell (e.g., a mammalian cell, e.g., a human cell) comprising the above-described system.
[0045] In one aspect, the present disclosure relates to a method for correcting a mutation in a human gene in a cell, tissue, or subject, comprising administering the above-described system to the cell, tissue, or subject. The system can be introduced in vivo, in vitro, ex vivo, or in situ. The nucleic acid (a) can be integrated into the genome of the host cell. In some embodiments, the nucleic acid (a) is not integrated into the genome of the host cell. In some embodiments, the heterologous sequence of interest is inserted at only one target site within the host cell genome. The heterologous sequence of interest can be inserted at two or more target sites within the host cell genome, for example, at the same corresponding sites on two homologous chromosomes or at two different sites on the same or different chromosomes. The heterologous sequence of interest can encode a mammalian polypeptide, or a fragment or variant thereof. The components of the system can be delivered on one, two, three, four, or more different nucleic acid molecules. The system can be introduced into the host cell by electroporation or by using at least one vehicle selected from a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle. [Brief explanation of the drawings]
[0046] [Figure 1] 1 shows the genetic engineering system described herein. The diagram on the left shows a genetic engineering polypeptide including a Cas nickase domain (e.g., spCas9 N863A) and a Pol domain connected by a linker. The diagram on the right shows a template nucleic acid including, from 5' to 3', a gRNA spacer, a gRNA scaffold, a heterologous sequence of interest, and a primer binding site sequence (PBS sequence). The heterologous sequence of interest may include a mutation region containing one or more sequence differences compared to the target site. The heterologous sequence of interest may also include pre-editing and post-editing homology regions flanking the mutation region. Without intending to be bound by any particular theory, it is believed that the gRNA spacer of the template nucleic acid binds to the second strand of the target site in the genome, and the gRNA scaffold of the template nucleic acid binds to the genetic engineering polypeptide, for example, to localize the genetic engineering polypeptide to the target site in the genome. It is believed that the Cas domain of the recombinant polypeptide nicks the target site (e.g., the first strand of the target site) and, for example, binds a PBS sequence to a sequence adjacent to the site to be modified on the first strand of the target site. It is believed that the Pol domain of the recombinant polypeptide polymerizes, for example, a sequence complementary to the heterologous target sequence, using the PBS sequence of the template nucleic acid as a primer and the first strand of the target site bound to a complementary sequence containing the heterologous target sequence of the template nucleic acid as a template. Without intending to be bound by any particular theory, it is believed that DNA polymerization can then proceed through the pre-edited homology region, then the mutation region, and then the post-edited homology region, to generate a DNA strand containing the mutation specified by the heterologous target sequence. [Figure 2] 1 is a diagram showing exemplary cleavages of human DNA polymerase theta. [Figure 3]
[0023] Figure 3A is a series of graphs showing editing activity by Cas-Pol recombinant polypeptides on the indicated template nucleic acid molecules in HEK293 cells (Figure 3A) and U2OS cells (Figure 3B). DETAILED DESCRIPTION OF THE INVENTION
[0047] definition The term "expression cassette," as used herein, refers to a nucleic acid construct that contains sufficient nucleic acid elements for expression of a nucleic acid molecule of the invention.
[0048] "gRNA spacer," as used herein, refers to a portion of a nucleic acid that has complementarity to a target nucleic acid and, together with the gRNA scaffold, can target a Cas protein to the target nucleic acid.
[0049] "gRNA scaffold," as used herein, refers to a portion of a nucleic acid that can bind to a Cas protein and, together with a gRNA spacer, target the Cas protein to a target nucleic acid. In some embodiments, the gRNA scaffold comprises a crRNA sequence, a tetraloop, and a tracrRNA sequence.
[0050] A "genetically modified polypeptide," as used herein, refers to a polypeptide comprising a polymerase or retroviral reverse transcriptase, or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to a polymerase or retroviral reverse transcriptase that is capable of integrating a nucleic acid sequence (e.g., a sequence provided on a template nucleic acid) into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule within the host cell). In some embodiments, the genetically modified polypeptide is capable of integrating a sequence substantially independent of the host machinery. In some embodiments, the genetically modified polypeptide integrates a sequence at a random location within the genome; in some embodiments, the genetically modified polypeptide integrates a sequence at a specific target site. In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively 1) bind to a template nucleic acid, 2) facilitate binding to a target DNA molecule, and 3) facilitate integration of at least a portion of the template nucleic acid into the target DNA. Transgenic polypeptides include both naturally occurring polypeptides and engineered variants thereof, e.g., having one or more amino acid substitutions relative to the native sequence. Transgenic polypeptides also include heterologous constructs, e.g., wherein one or more of the domains listed above are heterologous to one another, whether by heterologous fusion (or other conjugation) of otherwise wild-type domains as well as fusion of modified domains, e.g., by substitution or fusion of heterologous subdomains or other replacement domains. Exemplary transgenic polypeptides that can be used in the methods provided herein, as well as systems comprising and methods of using them, are described in PCT / US2021 / 020948, incorporated herein by reference, for example, with respect to transgenic polypeptides comprising retroviral reverse transcriptase domains. In some embodiments, transgenic polypeptides incorporate sequences into genes. In some embodiments, transgenic polypeptides incorporate sequences into sequences external to a gene. "Transgenic system," as used herein, refers to a system comprising a transgenic polypeptide and a template nucleic acid.
[0051] The term "domain," as used herein, refers to a structure of a biomolecule that contributes to a specific function of the biomolecule. A domain can include a continuous region (e.g., a contiguous sequence) or discrete, non-contiguous regions (e.g., non-contiguous sequences) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, polymerase (Pol) domains, recruitment domains, and reverse transcription domains; examples of nucleic acid domains are regulatory domains, such as transcription factor binding domains. In some embodiments, a domain (e.g., a Cas domain) can include two or more smaller domains (e.g., a DNA-binding domain and an endonuclease domain).
[0052] As used herein, the term "exogenous," when used in reference to a biomolecule (e.g., a nucleic acid sequence or a polypeptide), means that the biomolecule has been introduced into a host genome, cell, or organism by human intervention. For example, a nucleic acid that is added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods is exogenous to the existing nucleic acid sequence, cell, tissue, or subject.
[0053] As used herein, the terms "first strand" and "second strand" used to describe individual DNA strands of a target DNA distinguish between the two DNA strands upon which a Pol domain initiates polymerization, e.g., upon which target-primed synthesis is initiated. The term "first strand" refers to the strand of target DNA upon which a Pol domain initiates polymerization, e.g., upon which target-primed synthesis is initiated. The term "second strand" refers to the other strand of target DNA. The terms "first strand" and "second strand" do not otherwise describe target site DNA strands; for example, in some embodiments, the first strand and second strand are nicked by the polypeptides described herein, but the terms "first" and "second" strand are independent of the order in which such nicks appear.
[0054] A "genomic safe harbor site" (GSH site) is a site within a host genome that can accommodate the integration of new genetic material, such that the inserted genetic element does not cause significant alterations to the host genome that pose a risk to the host cell or organism. GSH sites generally meet one, two, three, four, five, six, seven, eight, or nine of the following criteria: (i) located >300 kb from a cancer-associated gene; (ii) located >300 kb from an miRNA / other functional small RNA; (iii) located >50 kb from the 5' gene end; (iv) located >50 kb from a replication origin; (v) located >50 kb away from an ultraconserved element; (vi) having low transcriptional activity (i.e., no mRNA + / - 25 kb); (vii) not within a variable copy number region; (viii) located within open chromatin; and / or (ix) having one copy and being unique within the human genome. Examples of GSH sites within the human genome that meet some or all of these criteria include: (i) adenovirus site 1 (AAVS1), the naturally occurring integration site of the AAV virus on chromosome 19; (ii) the chemokine (CC motif) receptor 5 (CCR5) gene, a chemokine receptor gene known as an HIV-1 co-receptor; (iii) the human orthologue of the mouse Rosa26 locus; and (iv) the ribosomal DNA ("rDNA") locus. Additional GSH sites are known and are described, for example, in Pellenz et al., epub August 20, 2018 (https: / / doi.org / 10.1101 / 396390).
[0055] The term "heterologous," when used to refer to a first element with respect to a second element, means that the first and second elements do not naturally exist in the arrangement described. For example, a heterologous polypeptide, nucleic acid molecule, construct, or sequence refers to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed; (b) a polypeptide or nucleic acid molecule or portion of a polypeptide or nucleic acid molecule that has been modified or mutated relative to its natural state; or (c) a polypeptide or nucleic acid molecule that has altered expression compared to native expression levels under similar conditions. For example, heterologous regulatory sequences (e.g., promoters, enhancers) can be used to regulate expression of a gene or nucleic acid molecule in a manner different from that in which the gene or nucleic acid molecule is normally expressed in nature. In another example, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., a DNA-binding domain of a polypeptide, or a nucleic acid encoding a DNA-binding domain of a polypeptide) can be positioned relative to other domains or can be of a different sequence or derived from a different source compared to other domains or portions of a polypeptide or its encoding nucleic acid. In certain embodiments, a heterologous nucleic acid molecule may be present in the native host cell genome, but may have an altered expression level or a different sequence, or both. In other embodiments, a heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but instead may be introduced into the host cell by transformation (e.g., transfection, electroporation), where the added molecule may be integrated into the host genome or may exist as extrachromosomal genetic material, either transiently (e.g., mRNA) or semi-stable for more than one generation (e.g., episomal viral vectors, plasmids, or other self-replicating vectors).
[0056] As used herein, "insertion" of a sequence into a target site refers to the net addition of a DNA sequence at the target site, e.g., where there is a new nucleotide in the heterologous sequence of interest that does not have a cognate position in the unedited target site. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous sequence of interest to the target nucleic acid sequence will result in an alignment gap in the target nucleic acid sequence.
[0057] As used herein, a "deletion" generated by a heterologous sequence of interest at a target site refers to the net deletion of DNA sequence at the target site, e.g., where there is a nucleotide in the unedited target site that does not have a cognate position in the heterologous sequence of interest. In some embodiments, nucleotide alignment of the PBS sequence and heterologous sequence of interest to the target nucleic acid sequence will result in an alignment gap in the molecule comprising the PBS sequence and the heterologous sequence of interest.
[0058] As used herein, the term "inverted terminal repeat" or "ITR" refers to an AAV viral cis element, so named because of its symmetry, which facilitates efficient propagation of the AAV genome. The minimal element for ITR function is a Rep binding site (RBS; for AAV2, 5'-GCGCGCTCGCTCGCTC-3' (SEQ ID NO: 4424)and terminal separation site (TRS; 5'-AGTTGG-3' for AAV2), as well as a variable palindromic sequence that allows hairpin formation. According to the present invention, an ITR contains at least these three elements (RBS, TRS, and a sequence that allows hairpin formation). In addition, in the present invention, the term "ITR" refers to the ITRs of known natural AAV serotypes (e.g., ITRs of serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 AAV), chimeric ITRs formed by fusing ITR elements from different serotypes, and functional variants thereof. A "functional variant" refers to a sequence that shows at least 80%, 85%, 90%, and preferably at least 95% sequence identity with a known ITR and allows for the growth of sequences containing the ITR in the presence of Rep proteins.
[0059] The term "mutation region," as used herein, refers to a region in a template nucleic acid that has one or more sequence differences compared to the corresponding sequence in a target nucleic acid. Sequence differences can include, for example, substitutions, insertions, frameshifts, or deletions.
[0060] The term "mutant," when applied to a nucleic acid sequence, means that nucleotides within a nucleic acid sequence have been inserted, deleted, or changed relative to a reference (e.g., naturally occurring) nucleic acid sequence. A single alteration may be made at a single locus (point mutation), or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.
[0061] "Nucleic acid molecule" refers to both RNA and DNA molecules, including, but not limited to, complementary DNA ("cDNA"), genomic DNA ("gDNA"), and messenger RNA ("mRNA"), and also includes synthetic nucleic acid molecules, such as those chemically synthesized or recombinantly produced, such as RNA templates, as described herein. Nucleic acid molecules can be double-stranded or single-stranded, circular or linear. If single-stranded, the nucleic acid molecule can be the sense or antisense strand. Unless otherwise specified, and as an example of all sequences described herein in the general format "SEQ ID NO:1," a nucleic acid containing "SEQ ID NO:1" refers to a nucleic acid having, at least a portion thereof, either (i) the sequence of SEQ ID NO:1, or (ii) a sequence complementary to SEQ ID NO:1. The choice between the two is determined by the context in which SEQ ID NO:1 is used. For example, if the nucleic acid is used as a probe, the choice between the two is determined by the requirement that the probe be complementary to the desired target. The nucleic acid sequences of the present disclosure may be chemically or biochemically modified or may contain non-natural or derivatized nucleotide bases, as will be readily understood by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalating agents (e.g., acridines, psoralens, etc.), chelators, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Chemically modified bases (e.g., see Table 9 below), backbones (e.g., see Table 10 below), and modified caps (e.g., see Table 11 below) are also included. Synthetic molecules that mimic polynucleotides in their ability to bind to designated sequences through hydrogen bonding and other chemical interactions are also included. Such molecules are known in the art and include, for example, those that substitute peptide linkages for phosphate linkages in the backbone of the molecule, eg, peptide nucleic acids (PNAs).Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structures, such as modifications found in "locked" nucleic acids (LNA). In various embodiments, the nucleic acid is operatively associated with additional genetic elements, such as tissue-specific expression-controlling sequences (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences), as well as additional elements, such as inverted repeats (e.g., inverted terminal repeats, e.g., elements derived from viruses (e.g., AAV ITRs)) and tandem repeats, inverted / direct repeats, homologous regions (segments with varying degrees of homology to the target DNA), untranslated regions (UTRs) (5', 3', or both 5' and 3' UTRs), and various combinations of the foregoing. Nucleic acid elements of the systems provided by the present invention can be provided in various topologies, including single-stranded, double-stranded, circular, linear, open-ended linear, closed-ended linear, and specific versions thereof, e.g., doggybone DNA (dbDNA), closed-ended DNA (ceDNA).
[0062] As used herein, a "gene expression unit" is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences can be contiguous or non-contiguous. Where necessary to link two protein coding regions, operably linked sequences can be in the same reading frame.
[0063] The term "host genome" or "host cell," as used herein, refers to a cell and / or its genome into which proteins and / or genetic material have been introduced. These terms refer not only to the particular subject cell and / or genome, but also to the progeny of such a cell and / or the genomes of the progeny of such a cell. Because certain modifications may occur in subsequent generations due to mutations or environmental influences, it is understood that such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term "host cell" as used herein. A host genome or host cell may be an isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line, or may be a host cell or host genome comprising a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, e.g., as described herein. In certain examples, the host cell may be a mammalian cell, a human cell, an avian cell, a reptilian cell, a bovine cell, an equine cell, a porcine cell, a caprine cell, a ovine cell, a chicken cell, or a turkey cell. In particular examples, the host cell can be a corn cell, a soybean cell, a wheat cell, or a rice cell.
[0064] As used herein, "operably associated" describes the functional relationship between two nucleic acid sequences, e.g., 1) a promoter and 2) a heterologous sequence of interest, and in such instances means that the promoter and heterologous sequence of interest (e.g., a gene of interest) are oriented such that, under appropriate conditions, the promoter drives expression of the heterologous sequence of interest. For example, a template nucleic acid bearing a promoter and a heterologous sequence of interest can be, for example, single-stranded in either a (+) or (-) orientation. The "operably associated" relationship between the promoter and heterologous sequence of interest in this template means that the template nucleic acid will be correctly transcribed when under appropriate conditions (e.g., in the (+) orientation, in the presence of required catalytic factors, NTPs, etc.), regardless of whether the template nucleic acid will be transcribed under a particular condition. Operative association applies similarly to other pairs of nucleic acids, including sequences encoding other tissue-specific expression control sequences (e.g., enhancers, repressors, and microRNA recognition sequences), IR / DR, ITR, UTR, or homologous regions, and a heterologous sequence of interest, or a retroviral RT domain.
[0065] The term "primer binding site sequence" or "PBS sequence," as used herein, refers to a portion of a template nucleic acid that can bind to a region contained in a target nucleic acid sequence. In some cases, a PBS sequence is a nucleic acid sequence that includes at least 3, 4, 5, 6, 7, or 8 bases that are 100% identical to a region contained in a target nucleic acid sequence. In some embodiments, a primer region includes at least 5, 6, 7, or 8 bases that are 100% identical to a region contained in a target nucleic acid sequence. Without intending to be bound by theory, in some embodiments, when a template nucleic acid includes a PBS sequence and a heterologous sequence of interest, the PBS sequence binds to a region contained in the target nucleic acid sequence, enabling the Pol domain to use that region as a primer for DNA polymerization and the heterologous sequence of interest as a template.
[0066] As used herein, a "stem-loop sequence" refers to a nucleic acid sequence (e.g., an RNA sequence) having a stem containing sufficient self-complementarity to form a stem-loop, e.g., at least 2 (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs, and a loop having at least 3 (e.g., 4) base pairs. The stem may contain mismatches or bulges.
[0067] As used herein, "tissue-specific expression-control sequence" refers to a nucleic acid element that increases or decreases the level of a transcript containing a heterologous sequence of interest in a target tissue in a tissue-specific manner, e.g., preferentially in on-target tissue compared to off-target tissue. In some embodiments, the tissue-specific expression-control sequence preferentially drives or represses the transcription, activity, or half-life of a transcript containing a heterologous sequence of interest in a target tissue in a tissue-specific manner, e.g., preferentially in on-target tissue compared to off-target tissue. Exemplary tissue-specific expression-control sequences include tissue-specific promoters, repressors, enhancers, or combinations thereof, and tissue-specific microRNA recognition sequences. Tissue specificity refers to on-target (tissues in which expression or activity of the template nucleic acid is desired or acceptable) and off-target (tissues in which expression or activity of the template nucleic acid is undesirable or unacceptable). For example, a tissue-specific promoter preferentially drives expression in on-target tissue compared to off-target tissue. In contrast, microRNAs that bind to tissue-specific microRNA recognition sequences are preferentially expressed in off-target tissues compared to on-target tissues, thereby reducing the expression of the template nucleic acid in the off-target tissue. Thus, promoters and microRNA recognition sequences specific to the same tissue, such as target tissues, have contrasting functions with respect to the transcription, activity, or half-life of the associated sequence in the tissue (i.e., matching expression levels, i.e., promoting and suppressing high levels of the microRNA in off-target tissues and low levels in on-target tissues, respectively, while the promoter drives high expression in on-target tissues and low expression in off-target tissues).
[0068] List of Headlines 1) Introduction 2) Genetic recombination system a) Polypeptide components of the genetic engineering system i) Lighting Domain ii) Endonuclease domain and DNA binding domain (1) A recombinant polypeptide containing a Cas domain (2) TAL effectors and zinc finger nucleases iii) Linker iv) Localization sequences for recombinant DNA systems v) Evolved variants of genetically engineered polypeptides and systems vi) Intein vii) Further domains b) Template nucleic acid i) gRNA spacer and gRNA scaffold ii) Heterologous sequence of interest iii) PBS sequence iv) Exemplary template sequences c) gRNA with inducible activity d) Circular RNA and ribozymes in recombinant DNA systems e) Target nucleic acid site f) Second Strand Nicking 3) Preparation of compositions and systems 4)Application a) Therapeutic use b) Application to plants 5) Administration and Delivery a) Tissue-specific activity / administration i) Promoter ii) microRNA b) Viral vectors and their components c) AAV administration d) Lipid nanoparticles 6) Kits, Products, and Pharmaceutical Compositions 7) Chemistry, Manufacturing, and Controls (CMC)
[0069] introduction The present disclosure relates to compositions, systems, and methods for targeting, editing, modifying, or manipulating DNA sequences (e.g., inserting a heterologous sequence of interest at a target site in a mammalian genome), e.g., at one or more locations within a DNA sequence in a cell, tissue, or subject, in vivo or in vitro. The heterologous DNA sequence of interest may include, for example, substitutions, deletions, insertions, e.g., of a coding sequence, a regulatory sequence, or a gene expression unit.
[0070] More specifically, the present disclosure provides a DNA polymerase (Pol)-based system for modifying a genomic DNA sequence of interest, for example, by inserting, deleting, or substituting one or more nucleotides into / from the sequence of interest.
[0071] Fusion of a Cas9-related functional group to a polymerase functional group can be used to drive modifications to genomic DNA. The polymerase functional group can be, for example, a DNA polymerase that synthesizes DNA from a nucleic acid template. The nucleic acid template can be, for example, DNA or RNA. In the case of a DNA polymerase that can use an RNA template, such as an RNA-dependent DNA polymerase, e.g., a reverse transcriptase, the recombinant polypeptide component can be provided with template RNA, such as those described above. One such example is DNA polymerase θ (the polypeptide product encoded by POLQ, which may be referred to herein as "POLQ" or "Polθ"), a eukaryotic DNA polymerase that has been shown to use either DNA or RNA as a template. Chandramouly et al. 2021 (DOI: 10.1126 / sciadv.abf1771). Thus, a Cas9 functionality fused (optionally via a linker) to a POLQ (or a component of a POLQ) can be used as a driver for genome modification when administered to an organism or cell along with a template RNA that targets the desired genomic site for modification (via a gRNA spacer), recruiting the Cas9 functionality (via the gRNA scaffold) to prime and template DNA synthesis (via the template RNA).
[0072] The present disclosure provides, in part, a genetic engineering system including a genetically engineered polypeptide component and a template nucleic acid (e.g., template RNA) component. In some embodiments, the genetic engineering system can be used to introduce modifications into a target site in a genome. In some embodiments, the genetically engineered polypeptide component includes a writing domain (e.g., a reverse transcriptase domain), a DNA-binding domain, and an endonuclease domain (e.g., a nickase domain). In some embodiments, the template nucleic acid (e.g., template RNA) includes a sequence (e.g., a gRNA spacer) that binds to the target site in the genome (e.g., binds to the second strand of the target site), a sequence that binds to the genetically engineered polypeptide component (e.g., a gRNA scaffold), a heterologous sequence of interest, and a PBS sequence. Without wishing to be bound by theory, it is believed that the template nucleic acid (e.g., template RNA) binds to the second strand of the target site in the genome and binds to the genetically engineered polypeptide component (e.g., localizes the polypeptide component to the target site in the genome). The endonuclease (e.g., nickase) of the recombinant polypeptide component may cleave the target site (e.g., the first strand of the target site) and, for example, a PBS sequence may bind to the sequence adjacent to the site to be modified on the first strand of the target site. The writing domain (e.g., reverse transcriptase domain) of the polypeptide component may polymerize, for example, a sequence complementary to the heterologous target sequence, using the first strand of the target site bound to a complementary sequence comprising the PBS sequence of the template nucleic acid as a primer and the heterologous target sequence of the template nucleic acid as a template. Without wishing to be bound by theory, it is believed that selection of an appropriate heterologous target sequence may result in the substitution, deletion, and / or insertion of one or more nucleotides at the target site.
[0073] Genetic engineering system In some embodiments, the genetic engineering systems described herein include (A) a genetically engineered polypeptide or a nucleic acid encoding a genetically engineered polypeptide, where the genetically engineered polypeptide includes (i) a reverse transcriptase domain and either (x) an endonuclease domain comprising DNA-binding functionality or (y) an endonuclease domain and a separate DNA-binding domain; and (B) a template RNA. In some embodiments, the genetically engineered polypeptide acts as a substantially autonomous protein machinery capable of integrating a template nucleic acid sequence into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule within the host cell) substantially independent of the host machinery. For example, the genetically engineered protein may include a DNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. In some embodiments, the DNA-binding functionality may include an RNA component, e.g., a gRNA spacer, that guides the protein to the DNA sequence. In other embodiments, the genetically engineered polypeptide may include a reverse transcriptase domain and an endonuclease domain. The RNA template element of the genetic engineering system is typically heterologous to the genetically engineered polypeptide element and provides the sequence of interest to be inserted (reverse transcribed) into the host genome. In some embodiments, the engineered polypeptide is capable of target-primed reverse transcription. In some embodiments, the engineered polypeptide is capable of second strand synthesis.
[0074] In some embodiments, the genetic recombination system is combined with a second polypeptide. In some embodiments, the second polypeptide may include an endonuclease domain. In some embodiments, the second polypeptide may include a polymerase domain, such as a reverse transcriptase domain. In some embodiments, the second polypeptide may include a DNA-dependent DNA polymerase domain. In some embodiments, the second polypeptide assists in completing genome editing, for example, by contributing to second strand synthesis or DNA repair recovery.
[0075] A functional recombinant polypeptide can be composed of unrelated DNA-binding, reverse transcription, and endonuclease domains. This modular structure allows for the combination of functional domains, such as dCas9 (DNA binding), MMLV reverse transcriptase (reverse transcription), and FokI (endonuclease). In some embodiments, multiple functional domains can occur from a single protein, such as Cas9 or Cas9 nickase (DNA binding, endonuclease).
[0076] In some embodiments, the engineered polypeptide comprises one or more domains that collectively: 1) bind to a template nucleic acid; 2) facilitate binding to a target DNA molecule; and 3) facilitate integration of at least a portion of the template nucleic acid into the target DNA. In some embodiments, the engineered polypeptide is a modified polypeptide comprising one or more amino acid substitutions relative to the corresponding native sequence. In some embodiments, the engineered polypeptide comprises two or more domains that are heterologous to each other, e.g., by heterologous fusion (or other conjugate) of otherwise wild-type domains as well as fusion of modified domains, e.g., by substitution or fusion of heterologous subdomains or other replacement domains. For example, in some embodiments, one or more of the RT domain is heterologous to the DBD; the DBD is heterologous to the endonuclease domain; or the RT domain is heterologous to the endonuclease domain.
[0077] In some embodiments, a template RNA molecule for use in the system comprises, from 5' to 3', (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous sequence of interest; and (4) a primer binding site (PBS) sequence. (1) a gRNA spacer of about 18 to 22 nt, for example, 20 nt; (2) A gRNA scaffold comprising one or more hairpin loops, e.g., one, two, or three loops for associating a template with a Cas domain, e.g., a nickase Cas9 domain. In some embodiments, the gRNA scaffold comprises, from 5' to 3', the sequence GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC (SEQ ID NO: 4008). (3) In some embodiments, the heterologous sequence of interest is, for example, 7 to 74, e.g., 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, or 70 to 80 or 80 to 90 nt in length. In some embodiments, the first (usually 5') base of the sequence is not C. (4) In some embodiments, the PBS sequence that binds to the target priming sequence after nicking is, for example, 3 to 20 nt, for example, 7 to 15 nt, for example, 12 to 14 nt, and has a GC content of 40 to 60%.
[0078] In some embodiments, a second gRNA associated with the system can help drive complete integration. In some embodiments, the second gRNA can target a position 0-200 nt away from the first strand nick, e.g., 0-50, 50-100, or 100-200 nt away from the first strand nick. In some embodiments, the second gRNA can only bind to its target sequence after editing has occurred, e.g., the gRNA binds to a sequence present in the heterologous sequence of interest but not in the initial target sequence.
[0079] In some embodiments, the genetic engineering systems described herein are used to perform editing in HEK293, K562, U2OS, or HeLa cells. In some embodiments, the genetic engineering systems are used to perform editing in primary cells, such as primary cortical neurons from E18.5 mice.
[0080] In some embodiments, the recombinant polypeptides described herein comprise a reverse transcriptase or RT domain (e.g., as described herein) comprising a MoMLV RT sequence or a variant thereof. In embodiments, the MoMLV RT sequence comprises one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L. In some embodiments, the MoMLV RT sequence comprises a combination of mutations, such as D200N, L603W, and T330P, optionally further comprising T306K and / or W313F.
[0081] In some embodiments, the endonuclease domain (e.g., as described herein) comprises nCAS9, e.g., comprising an H840A mutation.
[0082] In some embodiments, a heterologous sequence of interest (e.g., in a system described herein) is about 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, or more nucleotides in length.
[0083] In some embodiments, the RT and endonuclease domains are linked by a flexible linker, for example, comprising the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 4006).
[0084] In some embodiments, the endonuclease domain is N-terminal to the RT domain. In some embodiments, the endonuclease domain is C-terminal to the RT domain.
[0085] In some embodiments, the system incorporates a heterologous sequence of interest into a target site by TPRT, for example, as described herein.
[0086] In some embodiments, the genetically engineered polypeptide comprises a DNA-binding domain. In some embodiments, the genetically engineered polypeptide comprises an RNA-binding domain. In some embodiments, the RNA-binding domain comprises an RNA-binding domain of a B-box protein, an MS2 coat protein, dCas, or an element of a sequence in a table herein. In some embodiments, the RNA-binding domain can bind to a template RNA with higher affinity than a standard RNA-binding domain.
[0087] In some embodiments, the genetic recombination system is capable of generating insertions at a target site of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally up to 500, 400, 300, 200, or 100 nucleotides). In some embodiments, the genetic recombination system is capable of generating insertions at a target site of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally up to 500, 400, 300, 200, or 100 nucleotides). In some embodiments, the genetic recombination system is capable of generating insertions at target sites of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally up to 1, 5, 10, or 20 kilobases). In some embodiments, the genetic recombination system is capable of generating deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally up to 500, 400, 300, or 200 nucleotides). In some embodiments, the genetic engineering system is capable of generating deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally up to 500, 400, 300, or 200 nucleotides). In some embodiments, the genetic engineering system is capable of generating deletions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally up to 500, 400, 300, or 200 nucleotides).In some embodiments, the genetic engineering system is capable of generating deletions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally up to 1, 5, 10, or 20 kilobases). In some embodiments, the genetic engineering system is capable of generating substitutions at the target site of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides. In some embodiments, the genetic recombination system can generate substitutions at 1-2, 2-3, 3-4, 4-5, 5-10, 10-15, 15-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 nucleotides in the target site.
[0088] In some embodiments, the substitution is a transition mutation. In some embodiments, the substitution is a transversion mutation. In some embodiments, the substitution converts adenine to thymine, adenine to guanine, adenine to cytosine, guanine to thymine, guanine to cytosine, guanine to adenine, thymine to cytosine, thymine to adenine, thymine to guanine, cytosine to adenine, cytosine to guanine, or cytosine to thymine.
[0089] In some embodiments, the insertion, deletion, substitution, or a combination thereof increases or decreases expression (e.g., transcription or translation) of a gene. In some embodiments, the insertion, deletion, substitution, or a combination thereof increases or decreases expression (e.g., transcription or translation) of a gene by modifying, adding, or deleting sequences in a promoter or enhancer, such as sequences that bind transcription factors. In some embodiments, the insertion, deletion, substitution, or a combination thereof alters the translation of a gene (e.g., alters the amino acid sequence), inserts or deletes start or stop codons, alters or restores the translation frame of a gene. In some embodiments, the insertion, deletion, substitution, or a combination thereof alters the splicing of a gene, for example, by inserting, deleting, or modifying a splice acceptor or donor site. In some embodiments, the insertion, deletion, substitution, or a combination thereof alters the half-life of a transcript or protein. In some embodiments, the insertion, deletion, substitution, or combination thereof alters protein localization in a cell (e.g., from the cytoplasm to mitochondria, from the cytoplasm to the extracellular space (e.g., adding a secretion tag)). In some embodiments, the insertion, deletion, substitution, or combination thereof alters (e.g., improves) protein folding (e.g., to prevent the accumulation of misfolded proteins). In some embodiments, the insertion, deletion, substitution, or combination thereof alters, increases, decreases the activity of a gene, e.g., a protein encoded by the gene.
[0090] Exemplary recombinant polypeptides, and systems comprising and methods of using them, are described in PCT / US2021 / 020948, which is incorporated herein by reference, for example, with respect to retroviral RT domains, including amino acid and nucleic acid sequences therein.
[0091] Exemplary recombinant polypeptide and retroviral RT domain sequences are also described, for example, in International Patent Application No. PCT / US21 / 20948, filed March 4, 2021, e.g., Tables 30, 31, and 44 therein; this application is incorporated herein by reference in its entirety, e.g., with respect to the retroviral RT sequences and tables. Thus, the recombinant polypeptides described herein can comprise an amino acid sequence according to any of the tables described in this paragraph, or a domain thereof (e.g., a retroviral RT domain), or a functional fragment or variant thereof of any of the above, or an amino acid sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0092] In some embodiments, a polypeptide for use in any of the systems described herein can be a molecular or ancestral reconstructor based on aligned polypeptide sequences of multiple homologous proteins. In some embodiments, a reverse transcriptase domain for use in any of the systems described herein can be a molecular or ancestral reconstructor, or can be modified at specific residues based on alignment of reverse transcriptase domains from the same or different sources. Those skilled in the art can align polypeptide or nucleic acid sequences based on the accession numbers provided herein, for example, by using routine sequence analysis tools such as the Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis. Molecular reconstructors can be generated based on sequence consensus, for example, using techniques described in Ivics et al., Cell 1997, 501-510; Wagstaff et al., Molecular Biology and Evolution 2013, 88-99.
[0093] Polypeptide components of recombinant systems In some embodiments, the genetically engineered polypeptide has the functions of DNA target site binding, template nucleic acid (e.g., RNA) binding, DNA target site cleavage, and template nucleic acid (e.g., RNA) writing, e.g., reverse transcription. In some embodiments, each function is contained within a different domain. In some embodiments, a function can be attributed to two or more domains (e.g., two or more domains together exhibit functionality). In some embodiments, two or more domains can have the same or similar function (e.g., two or more domains each independently possess DNA binding functionality, e.g., in two different DNA sequences). In other embodiments, one or more domains can perform one or more functions; for example, a Cas9 domain can perform both DNA binding and target site cleavage. In some embodiments, the domains are all located within a single polypeptide. In some embodiments, the first domain is present in a first polypeptide and the second domain is present in a second polypeptide. For example, in some embodiments, the sequence may be split between a first polypeptide and a second polypeptide, e.g., the first polypeptide comprises a reverse transcriptase (RT) domain and the second polypeptide comprises a DNA-binding domain and an endonuclease domain, e.g., a nickase domain. By way of further example, in some embodiments, the first polypeptide and the second polypeptide each comprise a DNA-binding domain (e.g., a first DNA-binding domain and a second DNA-binding domain). In some embodiments, the first and second polypeptides may be post-translationally joined via a split intein to form a single recombinant polypeptide.
[0094] In some embodiments, an engineered polypeptide described herein comprises (e.g., a system described herein comprises an engineered polypeptide that comprises: 1) a Cas domain (e.g., a Cas nickase domain, e.g., a Cas9 nickase domain); 2) a reverse transcriptase (RT) domain of Table D, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto, wherein the RT domain is C-terminal to the Cas domain; and a linker disposed between the RT domain and the Cas domain, the linker having a sequence from the same row as the RT domain of Table D, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0095] In some embodiments, the RT domain has a sequence having 100% identity to an RT domain of Table D, and the linker has a sequence having 100% identity to a linker sequence from the same row as the RT domain of Table D. In some embodiments, the Cas domain comprises a sequence of Table 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto. In some embodiments, the recombinant polypeptide comprises an amino acid sequence according to any of SEQ ID NOs: 1-3332 in the Sequence Listing, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0096] In some embodiments, the engineered polypeptide comprises a GG amino acid sequence between the Cas domain and the linker, an AG amino acid sequence between the RT domain and the second NLS, and / or a GG amino acid sequence between the linker and the RT domain. In some embodiments, the engineered polypeptide comprises the sequence of SEQ ID NO: 4000 comprising the first NLS and the Cas domain, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto. In some embodiments, the engineered polypeptide comprises the sequence of SEQ ID NO: 4001 comprising the second NLS, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto. Exemplary N-terminal NLS-Cas9 domains [ka] Exemplary C-terminal sequences AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4001)
[0097] Lighting Domain In certain aspects of the present invention, the writing domain of the recombination system uses a polymerase functional group to drive modifications to genomic DNA. The polymerase functional group can be, for example, a DNA polymerase that synthesizes DNA from a nucleic acid template. The nucleic acid template can be, for example, DNA or RNA. In the case of a DNA polymerase that can use an RNA template, such as an RNA-dependent DNA polymerase, e.g., a reverse transcriptase, the recombination polypeptide component can be provided with template RNA, such as those described above. One such example is DNA polymerase θ (the polypeptide product encoded by POLQ, which may be referred to herein as "POLQ" or "Polθ"), a eukaryotic DNA polymerase that has been shown to use either DNA or RNA as a template. Chandramouly et al. 2021 (DOI: 10.1126 / sciadv.abf1771). Thus, a Cas9 functionality fused (optionally via a linker) to a POLQ (or a component of a POLQ) can be used as a driver for genome modification when administered to an organism or cell along with a template RNA that targets the desired genomic site for modification (via a gRNA spacer), recruiting the Cas9 functionality (via the gRNA scaffold) to prime and template DNA synthesis (via the template RNA).
[0098] DNA polymerases that use DNA as a template can also be incorporated into fusions with Cas9 functionality to perform genome modifications. In this case, the template nucleic acid is a fusion of an sgRNA and a DNA template, linked end-to-end by a covalent bond or a linker. Polθ (or its components) can also function in this manner, as it can synthesize DNA from a DNA template. As described herein, it is understood that embodiments that refer to a template RNA can include a template nucleic acid comprising ribonucleotides, or a template nucleic acid comprising ribonucleotides and deoxyribonucleotides (e.g., a template RNA comprising one or more RNA regions linked to one or more DNA regions). In some embodiments, the recombinant polypeptides described herein comprise a polymerase domain having an amino acid sequence according to Table 1, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. In some embodiments, the nucleic acids described herein encode a polymerase domain having an amino acid sequence according to Table 1, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0099] As described herein, it is understood that embodiments that refer to a reverse transcriptase or reverse transcriptase domain can include the polymerases listed in Table 1.
[0100] [Table 1-1]
[0101] [Table 1-2]
[0102] [Table 1-3]
[0103] [Table 1-4]
[0104] In certain aspects of the present invention, the writing domain of the genetic recombination system has reverse transcriptase activity and is also referred to as the reverse transcriptase domain (RT domain). In some embodiments, the RT domain includes an RT catalytic portion and an RNA-binding region (e.g., a region that binds to a template RNA).
[0105] In some embodiments, the nucleic acid encoding the reverse transcriptase is modified from its native sequence to have altered codon usage, e.g., improved for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase from a retrovirus. In some embodiments, the RT domain comprising the recombinant polypeptide has been mutated from its original amino acid sequence, e.g., has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 substitutions. In some embodiments, the RT domain is derived from a retroviral RT, e.g., HIV-1 RT, Moloney Murine Leukemia Virus (MMLV) RT, avian myeloblastosis virus (AMV) RT, or Rous Sarcoma Virus (RSV) RT.
[0106] In some embodiments, the retroviral reverse transcriptase (RT) domain exhibits increased stringency for target-primed reverse transcription (TPRT) initiation, e.g., compared to the endogenous RT domain. In some embodiments, the RT domain initiates TPRT when 3 nt within the target site immediately upstream of the first-strand nick, e.g., the genomic DNA priming the RNA template, are at least 66% or 100% complementary to the 3 nt of homology in the RNA template. In some embodiments, the RT domain initiates TPRT when there is less than a 5 nt mismatch (e.g., less than a 1, 2, 3, 4, or 5 nt mismatch) between the template RNA and the target DNA primed reverse transcription. In some embodiments, the RT domain is modified to increase the stringency of mismatches in priming the TPRT reaction, e.g., the RT domain tolerates no mismatches or tolerates fewer mismatches within the priming region compared to a wild-type (e.g., unmodified) RT domain. In some embodiments, the RT domain comprises an HIV-1 RT domain. In embodiments, the HIV-1 RT domain initiates synthesis at a lower level, even with three nucleotide mismatches, compared to alternative RT domains (e.g., as described by Jamburuthugoda and Eickbush J Mol Biol 407(5):661-672 (2011), which is incorporated herein by reference in its entirety).
[0107] In some embodiments, the RT domain forms a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain is a monomer. In some embodiments, the RT domain naturally functions as a monomer or a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain naturally functions as a monomer, e.g., is derived from a virus that functions as a monomer. In some embodiments, the RT domain is selected from the group consisting of murine leukemia virus (MLV; sometimes referred to as MoMLV) (e.g., P03355), porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2), mouse mammary tumor virus (MMTV) (e.g., UniProt P03365), Mason-Pfizer monkey virus (MPMV) (e.g., UniProt P07572), bovine leukemia virus (BLV) (e.g., UniProt P03361), human T-cell leukemia virus-1 (HTLV-1) (e.g., UniProt P03362), human foamy virus (HFV) (e.g., UniProt P03363), and / or porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2). In some embodiments, the RT domain is selected from the RT domains of simian foamy virus (SFV) (e.g., UniProt P23074), bovine foamy / syncytial virus (BFV / BSV) (e.g., UniProt O41894), or functional fragments or variants thereof (e.g., amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity thereto). In some embodiments, the RT domain is dimeric in its native functionality. In some embodiments, the RT domain is derived from a virus that functions as a dimer.In embodiments, the RT domain is selected from the group consisting of avian sarcoma / leukemia virus (ASLV) (e.g., UniProt A0A142BKH1), Rous sarcoma virus (RSV) (e.g., UniProt P03354), avian myeloblastosis virus (AMV) (e.g., UniProt Q83133), human immunodeficiency virus type I (HIV-1) (e.g., UniProt P03369), human immunodeficiency virus type II (HIV-2) (e.g., UniProt P15833), simian immunodeficiency virus (SIV) (e.g., UniProt P05896), bovine immunodeficiency virus (BIV) (e.g., UniProt P19560), equine infectious anemia virus (EIAV) (e.g., UniProt P03371), or feline immunodeficiency virus (FIV) (e.g., UniProt P16088) (Herschhorn and Hizi Cell Mol Life Sci 67(16):2717-2747 (2010)), or a functional fragment or variant thereof (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto). Naturally, heterodimeric RT domains may, in some embodiments, also be functional as homodimers. In some embodiments, the dimeric RT domain is expressed as a fusion protein, e.g., as a homodimeric fusion protein or a heterodimeric fusion protein. In some embodiments, the RT function of the system is fulfilled by multiple RT domains (e.g., as described herein).In further embodiments, the multiple RT domains can be fused or separate, for example, on the same polypeptide or on different polypeptides.
[0108] In some embodiments, the genetic recombination systems described herein include an integrase domain, e.g., the integrase domain can be part of an RT domain. In some embodiments, the RT domain (e.g., as described herein) includes an integrase domain. In some embodiments, the RT domain (e.g., as described herein) lacks an integrase domain or includes an integrase domain that has been inactivated by mutation or deletion. In some embodiments, the genetic recombination systems described herein include an RNase H domain, e.g., the RNase H domain can be part of the RT domain. In some embodiments, the RNase H domain is not part of the RT domain but is covalently linked via a flexible linker. In some embodiments, the RT domain (e.g., as described herein) includes an RNase H domain, e.g., an endogenous RNase H domain or a heterologous RNase H domain. In some embodiments, the RT domain (e.g., as described herein) lacks an RNase H domain. In some embodiments, an RT domain (e.g., as described herein) comprises an RNase H domain that has been added, deleted, mutated, or exchanged for a heterologous RNase H domain. In some embodiments, the polypeptide comprises an inactivated endogenous RNase H domain. In some embodiments, an endogenous RNase H domain from one of the polypeptide's other domains is genetically removed such that it is not included in the polypeptide, e.g., the endogenous RNase H domain is partially or completely truncated from the polypeptide that comprises the domain. In some embodiments, mutation of the RNase H domain produces a polypeptide that exhibits reduced RNase activity, e.g., by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% less, compared to an otherwise similar domain without the mutation, e.g., as measured by the method in Kotewicz et al. Nucleic Acids Res 16(1):265-277 (1988), incorporated herein by reference in its entirety.In some embodiments, RNase H activity is abolished.
[0109] In some embodiments, the RT domain is mutated to increase fidelity relative to other similar domains that do not have the mutation. For example, in some embodiments, YADD in the RT domain (e.g., in a reverse transcriptase) (SEQ ID NO: 4435) or YMDD (SEQ ID NO: 4436) The motif is YVDD (SEQ ID NO: 4437) In some embodiments, YADD (SEQ ID NO: 4435) , or YMDD (SEQ ID NO: 4436) , or YVDD (SEQ ID NO: 4437) Substitution of results in greater fidelity in retroviral reverse transcriptase activity (e.g., as described in Jamburuthugoda and Eickbush J Mol Biol 2011; incorporated herein by reference in its entirety).
[0110] In some embodiments, an engineered polypeptide described herein comprises an RT domain having an amino acid sequence according to Table 2, or a sequence with at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. In some embodiments, a nucleic acid described herein encodes an RT domain having an amino acid sequence according to Table 2, or a sequence with at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0111] [Table 2-1]
[0112] [Table 2-2]
[0113] [Table 2-3]
[0114]
Table 2-4
[0115]
Table 2-5
[0116]
Table 2-6
[0117]
Table 2-7
[0118]
Table 2-8
[0119]
Table 2-9
[0120]
Table 2-10
[0121]
Table 2-11
[0122]
Table 2-12
[0123]
Table 2-13
[0124]
Table 2-14
[0125] [Table 2-15]
[0126] In some embodiments, the reverse transcriptase domain is modified, for example, by site-directed mutagenesis. In some embodiments, the reverse transcriptase domain is engineered to have improved properties, such as the SuperScript IV (SSIV) reverse transcriptase from MMLV RT. In some embodiments, the reverse transcriptase domain may be engineered to have a lower error rate, for example, as described in International Publication No. WO2001068895, incorporated herein by reference. In some embodiments, the reverse transcriptase domain may be engineered to have increased thermostability. In some embodiments, the reverse transcriptase domain may be engineered to have increased processivity. In some embodiments, the reverse transcriptase domain may be engineered to be resistant to inhibitors. In some embodiments, the reverse transcriptase domain may be engineered to be faster. In some embodiments, the reverse transcriptase domain may be engineered to have increased tolerance to modified nucleotides in the RNA template. In some embodiments, the reverse transcriptase domain may be engineered to insert modified DNA nucleotides. In some embodiments, the reverse transcriptase domain is engineered to bind to the template RNA. In some embodiments, the one or more mutations are selected from D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N in the RT domain of murine leukemia virus reverse transcriptase, or a corresponding mutation at a corresponding position in another RT domain.
[0127] In some embodiments, the recombinant polypeptide comprises an RT domain from a retroviral reverse transcriptase, such as, for example, wild-type M-MLV RT, comprising the sequence: M-MLV(WT): [ka]
[0128] In some embodiments, the recombinant polypeptide comprises an RT domain from a retroviral reverse transcriptase, such as, for example, M-MLV RT, which comprises the following sequence: [ka]
[0129] In some embodiments, the recombinant polypeptide comprises an RT domain from a retroviral reverse transcriptase comprising the sequence of amino acids 659 to 1329 of NP_057933. In embodiments, the recombinant polypeptide further comprises one additional amino acid at the N-terminus of the sequence of amino acids 659 to 1329 of NP_057933, for example, as shown below. [ka] Core RT (bold), annotations above RNAseH (underlined), annotation as above
[0130] In embodiments, the recombinant polypeptide further comprises one additional amino acid at the C-terminus of the sequence of amino acids 659 to 1329 of NP_057933. In embodiments, the recombinant polypeptide comprises an RNase H1 domain (e.g., amino acids 1178 to 1318 of NP_057933).
[0131] In some embodiments, a retroviral reverse transcriptase domain, e.g., M-MLV RT, can contain one or more mutations from the wild-type sequence that can improve characteristics of the RT, such as thermostability, processivity, and / or template binding. In some embodiments, the M-MLV RT domain comprises one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, K103L relative to the M-MLV(WT) sequence above, e.g., a combination of mutations such as D200N, L603W, and T330P, optionally further comprising T306K and W313F. In some embodiments, an M-MLV RT as used herein comprises the mutations D200N, L603W, T330P, T306K, and W313F. In embodiments, the mutant M-MLV RT comprises the following amino acid sequence: M-MLV(PE2): [ka]
[0132] In some embodiments, the writing domain (e.g., the RT domain) comprises an RNA-binding domain that specifically binds to, for example, an RNA sequence. In some embodiments, the template RNA comprises an RNA sequence that is specifically bound by the RNA-binding domain of the writing domain.
[0133] In some embodiments, the reverse transcription domain simply recognizes and reverse transcribes a specific template of the system, e.g., a template RNA. In some embodiments, the template comprises a sequence or structure that allows recognition and reverse transcription by the reverse transcription domain. In some embodiments, the template comprises a sequence or structure that allows association with an RNA-binding domain of a polypeptide component of a genome modification system described herein. In some embodiments, the genome modification system preferentially reverse transcribes a template that includes an association sequence over a template that lacks the association sequence.
[0134] The writing domain may also comprise DNA-dependent DNA polymerase activity, e.g., an enzymatic activity capable of writing DNA into a genome from a template DNA sequence. In some embodiments, DNA-dependent DNA polymerization is used to complete second strand synthesis of target site editing. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in the polypeptide. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that is also capable of DNA-dependent DNA polymerization, e.g., second strand synthesis. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a second polypeptide of the system. In some embodiments, the DNA-dependent DNA polymerase activity is optionally provided by an endogenous host cell polymerase recruited to the target site by a component of the genome modification system.
[0135] In some embodiments, the reverse transcriptase domain exhibits a lower probability of poor termination (P) in vitro compared to a reference reverse transcriptase domain. off In some embodiments, the reference reverse transcriptase domain is a viral reverse transcriptase domain, for example, the RT domain from M-MLV.
[0136] In some embodiments, the reverse transcriptase domain has a nucleotide sequence of about 5×10 in vitro, e.g., as measured on 1094 nt RNA. -3 / nt, 5 × 10 -4 / nt, or 5 x 10 -6 A lower probability of insufficient termination (P off In some embodiments, insufficient termination rates in vitro are determined as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845, which is incorporated herein by reference in its entirety.
[0137] In some embodiments, the reverse transcriptase domain can complete at least about 30% or 50% of integrations in the cell. The percentage of complete integrations can be measured by dividing the number of substantially full-length integration events (e.g., genomic sites containing at least 98% of the expected integration sequence) by the total number of integration events (including substantially full-length and partial) in the cell population. In some embodiments, integration in the cell is determined (e.g., through the integration site) using long-read amplicon sequencing, as described, for example, in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated herein by reference in its entirety).
[0138] In embodiments, quantifying integration in a cell comprises counting the percentage of integrations that comprise at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the DNA sequence corresponding to the template RNA (e.g., a template RNA having a length of at least 0.05, 0.1, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 3, 4, or 5 kb, e.g., 0.5-0.6, 0.6-0.7, 0.7-0.8, 0.8-0.9, 1.0-1.2, 1.2-1.4, 1.4-1.6, 1.6-1.8, 1.8-2.0, 2-3, 3-4, or 4-5 kb).
[0139] In some embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro. In embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro at a rate of 0.1 to 50 nt / sec (e.g., 0.1 to 1, 1 to 10, or 10 to 50 nt / sec). In embodiments, polymerization of dNTPs by the reverse transcriptase domain is measured by a single-molecule assay, e.g., as described in Schwartz and Quake (2009) PNAS 106(48):20294-20299, which is incorporated by reference in its entirety.
[0140] In some embodiments, the reverse transcriptase domain is at least 1×10 nucleotides in length, e.g., as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2):147-153, which is incorporated herein by reference in its entirety. -3 ~1×10 -4 or 1 x 10 -4 ~1×10 -5 In some embodiments, the reverse transcriptase domain has an in vitro error rate (e.g., nucleotide misincorporation) of 1×10 substitutions / nt. In some embodiments, the reverse transcriptase domain is sequenced at 1×10 nucleotides per nucleotide in cells (e.g., HEK293T cells), e.g., by long-read amplicon sequencing, as described, e.g., in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated herein by reference in its entirety). -3 ~1×10 -4 or 1 x 10 -4 ~1×10 -5 It has an error rate (e.g., nucleotide misincorporation) of substitutions / nt.
[0141] In some embodiments, the reverse transcriptase domain is capable of performing reverse transcription of the target RNA in vitro. In some embodiments, the reverse transcriptase requires a primer of at least 3 nucleotides to initiate reverse transcription of the template. In some embodiments, reverse transcription of the target RNA is determined by detecting cDNA from the target RNA (e.g., when an ssDNA primer is provided that anneals to the target with at least 3, 4, 5, 6, 7, 8, 9, or 10 nt at the 3' end), e.g., as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (incorporated herein by reference in its entirety).
[0142] In some embodiments, the reverse transcriptase domain performs reverse transcription (e.g., by generating cDNA) at least 5-fold or 10-fold more efficiently, e.g., when converting its RNA template to cDNA, compared to, e.g., an RNA template lacking a protein-binding motif (e.g., a 3'UTR). In embodiments, the efficiency of reverse transcription is measured as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2):147-153, which is incorporated herein by reference in its entirety.
[0143] In some embodiments, the reverse transcriptase domain specifically binds to a particular RNA template at a higher frequency (e.g., about 5-fold or 10-fold higher frequency) than any endogenous cellular RNA, e.g., when expressed in a cell (e.g., HEK293T cell). In embodiments, the frequency of specific binding between the reverse transcriptase domain and the template RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501, which is incorporated herein by reference in its entirety.
[0144] template nucleic acid binding domain The recombinant polypeptide typically has a region capable of associating with a template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid binding domain is an RNA binding domain. In some embodiments, the RNA binding domain is a modular domain that can associate with RNA molecules containing a particular signature, e.g., a structural motif. In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within a reverse transcription domain, e.g., a component derived from a reverse transcriptase enzyme, which has a known signature for RNA preference.
[0145] In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within the target DNA binding domain. For example, in some embodiments, the DNA binding domain is a CRISPR-associated protein that recognizes the structure of a template nucleic acid (e.g., template RNA) that includes a gRNA. In some embodiments, the genetically engineered polypeptide comprises a DNA binding domain that includes a CRISPR-associated protein that associates with a gRNA scaffold, allowing the DNA binding domain to bind to a target genomic DNA sequence. In some embodiments, the gRNA scaffold and gRNA spacer are contained within the template nucleic acid (e.g., template RNA), such that the DNA binding domain is also the template nucleic acid binding domain. In some embodiments, the polypeptide has RNA binding function in multiple domains, e.g., it may bind to a gRNA structure within the CRISPR-associated DNA binding domain and an additional sequence or structure within the reverse transcriptase domain.
[0146] In some embodiments, the RNA-binding domain can bind to the template RNA with higher affinity than a standard RNA-binding domain. In some embodiments, the standard RNA-binding domain is an RNA-binding domain from S. pyogenes Cas9. In some embodiments, the RNA-binding domain can bind to the template RNA with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM). In some embodiments, the affinity of the RNA-binding domain for its template RNA is measured in vitro, e.g., by thermophoresis, as described, e.g., in Asmari et al. Methods 146:107-119 (2018), which is incorporated herein by reference in its entirety. In some embodiments, the affinity of the RNA-binding domain for its template RNA is measured in a cell (e.g., by FRET or CLIP-Seq).
[0147] In some embodiments, the RNA-binding domain associates with the template RNA in vitro at least about 5-fold or 10-fold more frequently than scrambled RNA. In some embodiments, the frequency of association between the RNA-binding domain and the template RNA or scrambled RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501, incorporated herein by reference in its entirety. In some embodiments, the RNA-binding domain associates with the template RNA in cells (e.g., HEK293T cells) at least about 5-fold or 10-fold more frequently than scrambled RNA. In some embodiments, the frequency of association between the RNA-binding domain and the template RNA or scrambled RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) supra.
[0148] Endonuclease domain and DNA binding domain In some embodiments, the genetically engineered polypeptide functions to cleave a DNA target site via an endonuclease domain. In some embodiments, the genetically engineered polypeptide comprises, for example, a DNA-binding domain for binding to a target nucleic acid. In some embodiments, a domain of the genetically engineered polypeptide (e.g., a Cas domain) comprises two or more smaller domains, for example, a DNA-binding domain and an endonuclease domain. When a DNA-binding domain (e.g., a Cas domain) is described as binding to a target nucleic acid sequence, it is understood that in some embodiments, binding is mediated by a gRNA.
[0149] In some embodiments, the domain has two functions. For example, in some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA)-binding domain. For example, in some embodiments, the polypeptide comprises a CRISPR-associated endonuclease domain that binds to a template RNA, including a gRNA, binds to a target DNA sequence (e.g., having complementarity to a portion of the gRNA), and cleaves the target DNA sequence. In some embodiments, an endonuclease domain or endonuclease / DNA-binding domain derived from a heterologous source can be used or modified (e.g., by inserting, deleting, or substituting one or more residues) in the genetic engineering systems described herein.
[0150] In some embodiments, the nucleic acid encoding the endonuclease domain or endonuclease / DNA-binding domain is modified from its native sequence to have modified codon usage, e.g., improved for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, such as a Cas endonuclease (e.g., Cas9), a Type II restriction endonuclease (e.g., Fok1), a meganuclease (e.g., I-SceI), or other endonuclease domain.
[0151] In certain aspects, the DNA-binding domain of a genetically engineered polypeptide described herein is selected, designed, or engineered for binding to a desired host DNA target sequence. In certain embodiments, the DNA-binding domain of the polypeptide is a heterologous DNA-binding factor. In some embodiments, the heterologous DNA-binding factor is a zinc finger factor or TAL effector factor, e.g., a zinc finger or TAL polypeptide or a functional fragment thereof. In some embodiments, the heterologous DNA-binding factor is a sequence-guided DNA-binding factor, such as Cas9, Cpfl, or other CRISPR-associated protein, that has been modified to lack endonuclease activity. In some embodiments, the heterologous DNA-binding factor retains endonuclease activity. In some embodiments, the heterologous DNA-binding factor retains partial endonuclease activity, such as cleaving ssDNA, e.g., has nickase activity. In certain embodiments, the heterologous DNA-binding domain can be any one or more of Cas9, a TAL domain, a ZF domain, a Myb domain, a combination thereof, or a complex thereof.
[0152] In some embodiments, the DNA-binding domain is modified, e.g., by site-directed mutagenesis, to increase or decrease DNA-binding factors (e.g., the number and / or specificity of zinc fingers), etc., to alter DNA-binding specificity and affinity. In some embodiments, the nucleic acid sequence encoding the DNA-binding domain is modified from its native sequence to have altered codon usage, e.g., improved for human cells. In several embodiments, the DNA-binding domain includes one or more modifications relative to the wild-type DNA-binding domain, e.g., modifications by directed evolution, e.g., phage-assisted continuous evolution (PACE).
[0153] In some embodiments, the DNA-binding domain comprises a meganuclease domain (e.g., an endonuclease domain portion, e.g., as described herein), or a functional fragment thereof. In some embodiments, the meganuclease domain has endonuclease activity, e.g., double-strand cleavage and / or nickase activity. In other embodiments, the meganuclease domain has reduced activity, e.g., lacks endonuclease activity, e.g., the meganuclease is catalytically inactive. In some embodiments, a catalytically inactive meganuclease is used as the DNA-binding domain, e.g., as described in Fonfara et al. Nucleic Acids Res 40(2):847-860 (2012), which is incorporated herein by reference in its entirety.
[0154] In some embodiments, the recombinant polypeptide comprises modifications to the DNA-binding domain, e.g., compared to the wild-type polypeptide. In some embodiments, the DNA-binding domain comprises additions, deletions, substitutions, or modifications to the amino acid sequence of the original DNA-binding domain. In some embodiments, the DNA-binding domain is modified to comprise a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the functional domain replaces at least a portion (e.g., the entirety) of a previous DNA-binding domain of the polypeptide. In some embodiments, the functional domain comprises a zinc finger (e.g., a zinc finger that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the functional domain comprises a Cas domain (e.g., a Cas domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the Cas domain comprises Cas9 or a mutant or variant thereof (e.g., as described herein). In embodiments, the Cas domain is associated with a guide RNA (gRNA), e.g., as described herein. In embodiments, the Cas domain is guided to the target nucleic acid (e.g., DNA) sequence of interest by the gRNA. In some embodiments, the Cas domain is encoded in the same nucleic acid (e.g., RNA) molecule as the gRNA. In some embodiments, the Cas domain is encoded in a different nucleic acid (e.g., RNA) molecule than the gRNA.
[0155] In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with higher affinity than a standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain from S. pyogenes Cas9. In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM).
[0156] In some embodiments, the affinity of a DNA binding domain for its target sequence (e.g., a dsDNA target sequence) is measured in vitro, e.g., by thermophoresis, as described, e.g., in Asmari et al. Methods 146:107-119 (2018), which is incorporated herein by reference in its entirety.
[0157] In embodiments, the DNA-binding domain can bind to its target sequence (e.g., a dsDNA target sequence) with an affinity of, for example, 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM) in the presence of a molar excess, e.g., about a 100-fold molar excess, of scrambled sequence competitor dsDNA.
[0158] In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a higher frequency than any other sequence in the genome of the target cell, e.g., a human target cell, as measured by, e.g., ChIP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21, which is incorporated herein by reference in its entirety. In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a frequency at least about 5-fold or 10-fold higher than any other sequence in the genome of the target cell, as measured by, e.g., ChIP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010) supra.
[0159] In some embodiments, the endonuclease domain has nickase activity and cleaves one strand of the target DNA. In some embodiments, the nickase activity reduces the formation of double-strand breaks at the target site. In some embodiments, the endonuclease domain generates staggered nicks in the first and second strands of the target DNA. In some embodiments, the staggered nicks generate free 3' overhangs at the target site. In some embodiments, the free 3' overhangs at the target site improve editing efficiency, for example, by enhancing access and annealing of the 3' homologous region of the template nucleic acid. In some embodiments, the staggered nicks reduce the formation of double-strand breaks at the target site.
[0160] In some embodiments, the endonuclease domain cleaves both strands of the target DNA, e.g., resulting in a blunt-end cleavage of the target with no ssDNA overhangs on either side of the cleavage site. The amino acid sequence of the endonuclease domain of the genetic recombination systems described herein can be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the amino acid sequence of the endonuclease domain described herein.
[0161] In certain embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof. In certain embodiments, the heterologous endonuclease is a Holliday junction resolvase or a homolog thereof, such as the Holliday junction cleavage enzyme (Ssol Hje) from Sulfolobus solfataricus (Govindaraju et al., Nucleic Acids Research 44:7, 2016). In certain embodiments, the heterologous endonuclease is a large fragment endonuclease of a spliceosomal protein, such as Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). In certain embodiments, the heterologous endonuclease is derived from a CRISPR-associated protein, such as Cas9. In certain embodiments, the heterologous endonuclease is modified to have only ssDNA cleavage activity, e.g., only nickase activity, e.g., a Cas9 nickase, e.g., SpCas9 with a D10A, H840A, or N863A mutation. Table 3 lists exemplary Cas proteins and mutations associated with nickase activity. In yet other embodiments, the homologous endonuclease domain is modified, e.g., by site-directed mutagenesis, to alter DNA endonuclease activity. In yet other embodiments, the endonuclease domain is modified to reduce DNA sequence specificity, e.g., by truncation to remove a domain that confers DNA sequence specificity or mutations to inactivate the region that confers DNA sequence specificity.
[0162] In some embodiments, the endonuclease domain has nickase activity and does not form double-stranded breaks. In some embodiments, the endonuclease domain forms single-stranded breaks more frequently than double-stranded breaks, for example, at least 90%, 95%, 96%, 97%, 98%, or 99% of the cuts are single-stranded breaks, or less than 10%, 5%, 4%, 3%, 2%, or 1% of the cuts are double-stranded breaks. In some embodiments, the endonuclease does not substantially form double-stranded breaks. In some embodiments, the endonuclease does not form detectable levels of double-stranded breaks.
[0163] In some embodiments, the endonuclease domain has a nickase activity that nicks the target site DNA of the first strand; for example, in some embodiments, the endonuclease domain cleaves the genomic DNA of the target site near the modification site on the strand that will be extended by the writing domain. In some embodiments, the endonuclease domain has a nickase activity that nicks the target site DNA of the first strand but does not nick the target site DNA of the second strand. For example, when a polypeptide comprises a CRISPR-associated endonuclease domain with nickase activity, in some embodiments, the CRISPR-associated endonuclease domain nicks the target site DNA strand that contains the PAM site (e.g., does not nick the target site DNA strand that does not contain the PAM site). By way of further example, when a polypeptide comprises a CRISPR-associated endonuclease domain with nickase activity, in some embodiments, the CRISPR-associated endonuclease domain nicks the target site DNA strand that does not contain a PAM site (e.g., and does not nick the target site DNA strand that contains a PAM site).
[0164] In some other embodiments, the endonuclease domain has nickase activity, which creates nicks in the first and second strands of target site DNA. Without intending to be bound by theory, after the writing domain (e.g., RT domain) of a polypeptide described herein polymerizes (e.g., reverse transcribes) from a heterologous target sequence of a template nucleic acid (e.g., template RNA), the cellular DNA repair machinery must repair the nick on the first DNA strand. The target site DNA here contains two distinct sequences relative to the first DNA strand: one corresponding to the original genomic DNA (e.g., with a free 5' end) and the second corresponding to that polymerized from the heterologous target sequence (e.g., with a free 3' end). It is believed that the two distinct sequences equilibrate with each other, with first one hybridizing to the second strand, followed by the other, and the order of incorporation of the cellular DNA repair machinery into its repair target site is a stochastic process. Without intending to be bound by any particular theory, it is believed that the introduction of an additional nick into the second strand can bias cellular DNA repair mechanisms to use sequences based on the heterologous target sequence more frequently than the original genomic sequence (Anzalone et al. Nature 576:149-157 (2019)). In some embodiments, the additional nick is positioned at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides 5' or 3' of the target site modification (e.g., insertion, deletion, or substitution) or relative to the nick on the first strand.
[0165] Alternatively or additionally, without intending to be bound by any particular theory, it is believed that an additional nick in the second strand may facilitate second strand synthesis. In some embodiments, when the genetic engineering system inserts or replaces a portion of the first strand, synthesis of a new sequence corresponding to the insertion / substitution in the second strand is required.
[0166] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain), which nicks both the first strand and the second strand. For example, in such embodiments, the endonuclease domain can be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA spacer that directs nicking of the first strand and an additional gRNA spacer that directs nicking of the second strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, wherein a first endonuclease domain nicks the first strand and a second endonuclease domain nicks the second strand (optionally, the first endonuclease domain does not (e.g., is unable to) nick the second strand and the second endonuclease domain does not (e.g., is unable to) nick the first strand).
[0167] In some embodiments, the endonuclease domain can nick the first and second strands. In some embodiments, the first and second strand nicks occur at the same position in the target site, but not on opposite strands. In some embodiments, the second strand nick occurs at a staggered position, e.g., upstream or downstream from the first nick. In some embodiments, the endonuclease domain generates a deletion of the target site when the second strand nick is upstream of the first strand nick. In some embodiments, the endonuclease domain generates a duplication of the target site when the second strand nick is downstream of the first strand nick. In some embodiments, the endonuclease domain does not generate a duplication and / or deletion when the first and second strand nicks occur at the same position in the target site. In some embodiments, the endonuclease domain has altered activity depending on the protein conformation or RNA binding state, for example, to promote first strand or second strand nicking (e.g., as described in Christensen et al. PNAS 2006; incorporated herein by reference in its entirety).
[0168] In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof. In some embodiments, the endonuclease domain comprises a homing endonuclease, or a functional fragment thereof. In some embodiments, the endonuclease domain comprises a homing endonuclease, e.g., ... (SEQ ID NO: 4555) In some embodiments, the endonuclease domain comprises a meganuclease from the I-SmaMI (Uniprot F7WD42), I-SceI (Uniprot P03882), I-AniI (Uniprot P03880), I-DmoI (Uniprot P21505), I-CreI (Uniprot P05725), I-TevI (Uniprot P13299), I-OnuI (Uniprot Q4VWW5), or I-BmoI (Uniprot Q9ANR6), or a fragment thereof. In some embodiments, the meganuclease is naturally a monomer, e.g., I-SceI, I-TevI, or a dimer, e.g., I-CreI, in its functional form. (SEQ ID NO: 4555) LAGLIDADG with a single copy of the motif (SEQ ID NO: 4555) Meganucleases generally form homodimers, while LAGLIDADG (SEQ ID NO: 4555)Members with two copies of the motif are generally found as monomers. In some embodiments, meganucleases that normally form dimers are expressed as fusions, for example, two subunits are expressed as a single ORF, optionally linked by a linker, for example, as an I-CreI dimer fusion (Rodriguez-Fornes et al. Gene Therapy 2020; the entire contents of which are incorporated herein by reference). In some embodiments, meganucleases, or functional fragments thereof, are engineered to preferentially exhibit nickase activity in one strand of a double-stranded DNA molecule, e.g., I-SceI (K122I and / or K223I) (Niu et al. J Mol Biol 2008), I-AniI (K227M) (McConnell Smith et al. PNAS 2009), I-DmoI (Q42A and / or K120M) (Molina et al. J Biol Chem 2015). In some embodiments, meganucleases or functional fragments thereof with this preference for single-strand cleavage are used, for example, as endonuclease domains with nickase activity. In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof, that naturally targets or has been engineered to target a safe harbor site, e.g., an SH6 site that targets I-CreI (Rodriguez-Fornes et al., supra). In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof, with a sequence-tolerant catalytic domain, e.g., I-TevI, which recognizes the minimal motif CNNNG (Kleinstiver et al. PNAS 2012).In some embodiments, the target sequence-resistant catalytic domain is fused to a DNA-binding domain, e.g., fusion of I-TevI to (i) a Zn finger to create Tev-ZFE (Kleinstiver et al. PNAS 2012), (ii) another meganuclease to create MegaTev (Wolfs et al. Nucleic Acids Res 2014), and / or (iii) Cas9 to create TevCas9 (Wolfs et al. PNAS 2016) induces activity.
[0169] In some embodiments, the endonuclease domain comprises a restriction enzyme, e.g., a Type IIS or Type IIP restriction enzyme. In some embodiments, the endonuclease domain comprises a Type IIS restriction enzyme, e.g., FokI, or a fragment or variant thereof. In some embodiments, the endonuclease domain comprises a Type IIP restriction enzyme, e.g., PvuII, or a fragment or variant thereof. In some embodiments, the dimeric restriction enzyme is expressed as a fusion, e.g., a FokI dimer fusion, such that it functions as a single strand (Minczuk et al. Nucleic Acids Res 36(12):3926-3938 (2008)).
[0170] The use of additional endonuclease domains is described, for example, in Guha and Edgell Int J Mol Sci 18(22):2565 (2017), which is incorporated herein by reference in its entirety.
[0171] In some embodiments, the recombinant polypeptide comprises a modification to the endonuclease domain, e.g., compared to a wild-type Cas protein. In some embodiments, the endonuclease domain comprises an addition, deletion, substitution, or modification to the amino acid sequence of a wild-type Cas protein. In some embodiments, the endonuclease domain is modified to comprise a heterologous functional domain that specifically binds to and / or directs endonucleolytic cleavage of a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the endonuclease domain comprises a zinc finger. In several embodiments, the endonuclease domain, including a Cas domain, associates with a guide RNA (gRNA), e.g., as described herein. In some embodiments, the endonuclease domain is modified to comprise a functional domain that does not target a specific target nucleic acid (e.g., DNA) sequence. In several embodiments, the endonuclease domain comprises a Fok1 domain.
[0172] In some embodiments, the endonuclease domain associates with the target dsDNA at least about 5-fold or 10-fold more frequently than scrambled dsDNA in vitro. In some embodiments, the endonuclease domain associates with the target dsDNA at least about 5-fold or 10-fold more frequently than scrambled dsDNA in vitro, e.g., in a cell (e.g., HEK293T cell). In some embodiments, the frequency of association between the endonuclease domain and the target DNA or scrambled DNA is measured by ChIP-seq, e.g., as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21, incorporated herein by reference in its entirety.
[0173] In some embodiments, the endonuclease domain may catalyze the formation of nicks at the target sequence, e.g., by at least about a 5-fold or 10-fold increase, relative to a non-target sequence (e.g., relative to any other genomic sequence in the genome of the target cell). In some embodiments, the level of nicking is measured using Nick-Seq, e.g., as described in Elacqua et al. (2019) bioRxiv doi.org / 10.1101 / 867937, which is incorporated by reference in its entirety.
[0174] In some embodiments, the endonuclease domain is capable of nicking DNA in vitro. In embodiments, the nick results in an exposed base. In embodiments, the exposed base can be detected using a nuclease sensitivity assay, e.g., as described in Chaudhry and Weinfeld (1995) Nucleic Acids Res 23(19):3805-3809, incorporated herein by reference in its entirety. In embodiments, the level of exposed bases (e.g., as detected by the nuclease sensitivity assay) is increased by at least 10%, 50%, or more compared to the reference endonuclease domain. In some embodiments, the reference endonuclease domain is an endonuclease domain from Cas9 of Streptococcus pyogenes (S. pyogenes).
[0175] In some embodiments, the endonuclease domain is capable of nicking DNA in a cell. In embodiments, the endonuclease domain is capable of nicking DNA in a HEK293T cell. In embodiments, unrepaired nicks that undergo replication in the absence of Rad51 result in an increased rate of NHEJ at the site of the nick, detectable, for example, by using a Rad51 inhibition assay, e.g., as described in Bothmer et al. (2017) Nat Commun 8:13905 (incorporated herein by reference in its entirety). In embodiments, the NHEJ rate is increased by more than 0-5%. In embodiments, the NHEJ rate is increased, for example, by 20-70% (e.g., 30%-60% or 40-50%) upon Rad51 inhibition.
[0176] In some embodiments, the endonuclease domain releases the target after cleavage. In some embodiments, target release is indicated indirectly by assessing multiple enzymatic turnover, e.g., as described in Yourik at al. RNA 25(1):35-44 (2019) (incorporated herein by reference in its entirety) and as shown in Figure 2. In some embodiments, the k of the endonuclease domain exp is measured by this method and is 1×10 -3 ~1×10 -5 It is min-1.
[0177] In some embodiments, the endonuclease domain is capable of binding to about 1 x 10 8 s -1 M -1 Catalytic efficiency (k cat / K m In some embodiments, the endonuclease domain has a nucleotide sequence of about 1 x 10 in vitro. 5 , 1×10 6 , 1×10 7 , or 1×10 8 s -1 M -1In some embodiments, the catalytic efficiency is determined as described in Chen et al. (2018) Science 360(6387):436-439, which is incorporated herein by reference in its entirety. In some embodiments, the endonuclease domain has a catalytic efficiency of greater than about 1×10 in a cell. 8 s -1 M -1 Catalytic efficiency (k cat / K m In some embodiments, the endonuclease domain has a denaturing activity of about 1×10 5 , 1×10 6 , 1×10 7 , or 1×10 8 s -1 M -1 It has a catalytic efficiency of more than
[0178] Engineered polypeptide containing a Cas domain In some embodiments, the transgenic polypeptides described herein comprise a Cas domain. In some embodiments, the Cas domain can guide the transgenic polypeptide to a target site specified by a gRNA spacer, thereby modifying a target nucleic acid sequence in "cis." In some embodiments, the transgenic polypeptide is fused to a Cas domain. In some embodiments, the transgenic polypeptide comprises a CRISPR / Cas domain (also referred to herein as a CRISPR-associated protein). In some embodiments, the CRISPR / Cas domain comprises a protein involved in the clustered regularly interspaced short palindromic repeats (CRISPR) system, e.g., a Cas protein, and optionally binds to a guide RNA, e.g., a single guide RNA (sgRNA).
[0179] The CRISPR system is an adaptive defense system first discovered in bacteria and archaea. CRISPR systems use RNA-guided nucleases called CRISPR-associated or "Cas" endonucleases (e.g., Cas9 or Cpf1) to cleave foreign DNA. For example, in a typical CRISPR-Cas system, the endonuclease is guided to a target nucleotide sequence (e.g., a site in the genome to be sequence-edited) by a sequence-specific, non-coding "guide RNA" that targets single- or double-stranded DNA sequences. Three classes of CRISPR systems (I-III) have been identified. Class II CRISPR systems use a single Cas endonuclease (rather than multiple Cas proteins). One Class II CRISPR system includes a type II Cas endonuclease, such as Cas9, a CRISPR RNA ("crRNA"), and a trans-activating crRNA ("tracrRNA"). The crRNA typically contains a "spacer" sequence (protospacer), an approximately 20-nucleotide RNA sequence that corresponds to the target DNA sequence. In wild-type systems, and in some engineered systems, the crRNA binds to the tracrRNA, forming a partially double-stranded structure that is cleaved by RNase III. cThe crRNA / tracrRNA hybrid molecule also contains a region that results in a rRNA / tracrRNA hybrid molecule. The crRNA / tracrRNA hybrid then guides the Cas endonuclease to recognize and cleave the target DNA sequence. The target DNA sequence is generally flanked by a "protospacer adjacent motif" ("PAM") that is specific for a given Cas endonuclease and required for cleavage activity at the target site that matches the spacer of the crRNA. CRISPR endonucleases identified from various prokaryotic species have unique PAM sequence requirements, for example, as listed for exemplary Cas enzymes in Table 3; example PAM sequences include 5'-NGG (Streptococcus pyogenes), 5'-NNAGAA (Streptococcus thermophilus CRISPR1), 5'-NGGNG (Streptococcus thermophilus CRISPR3), and 5'-NNNGATT (Neisseria meningiditis). Some endonucleases, such as the Cas9 endonuclease, associate with G-rich PAM sites, e.g., 5'-NGG, and perform blunt-end cleavage of target DNA three nucleotides upstream (5') from the PAM site. Another class II CRISPR system includes a type V endonuclease, Cpf1, which is smaller than Cas9; examples include AsCpf1 (from Acidaminococcus sp.) and LbCpf1 (from Lachnospiraceae sp.). Cpf1-associated CRISPR arrays do not require tracrRNA and are processed into mature crRNA; in other words, the Cpf1 system, in some embodiments, exclusively contains Cpf1 nuclease and crRNA to cleave the target DNA sequence. Cpf1 endonuclease typically associates with T-rich PAM sites, such as 5'-TTN. Cpf1 can also recognize the 5'-CTA PAM motif.Cpf1 typically cleaves target DNA by introducing an offset or staggered double-stranded break into a 4- or 5-nucleotide 5' overhang, e.g., cleaving the target DNA with a 5-nucleotide offset or staggered break located 18 nucleotides downstream (3') from the PAM site on the coding strand and 23 nucleotides downstream from the PAM site on the complementary strand; the 5-nucleotide overhang resulting from such an offset break allows for more precise genome editing by DNA insertion via homologous recombination rather than insertion with blunt-ended cut DNA. See, e.g., Zetsche et al. (2015) Cell, 163:759-771.
[0180] Various CRISPR-associated (Cas) genes or proteins can be used in the techniques provided by the present disclosure, and the choice of Cas protein will depend on the specific requirements of the method. Specific examples of Cas proteins include Class II systems, including Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, the Cas protein, e.g., the Cas9 protein, can be derived from any of a variety of prokaryotic species. In some embodiments, a particular Cas protein, e.g., a particular Cas9 protein, is selected to recognize a particular protospacer adjacent motif (PAM) sequence. In some embodiments, the DNA-binding domain or endonuclease domain comprises a sequence-targeting polypeptide, such as a Cas protein, e.g., Cas9. In certain embodiments, the Cas protein, e.g., the Cas9 protein, can be obtained from bacteria or archaea, or can be synthesized using known methods. In certain embodiments, the Cas protein can be derived from Gram-positive or Gram-negative bacteria. In certain embodiments, the Cas protein is selected from the group consisting of Streptococcus (e.g., S. pyogenes or S. thermophilus), Francisella (e.g., F. novicida), Staphylococcus (e.g., S. aureus), Acidaminococcus (e.g., Acidaminococcus sp. BV3L6), and the like. sp. BV3L6), Neisseria (e.g., N. meningitidis), Cryptococcus, Corynebacterium, Haemophilus, Eubacterium, Pasteurella, Prevotella, Veillonella, or Marinobacter.
[0181] In some embodiments, the recombinant polypeptide may comprise a Cas domain listed in Table 3 or 4, or a functional fragment thereof, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto.
[0182] [Table 3]
[0183] [Table 4-1]
[0184] [Table 4-2]
[0185] [Table 4-3]
[0186] [Table 4-4]
[0187] [Table 4-5]
[0188] [Table 4-6]
[0189] [Table 4-7]
[0190] [Table 4-8]
[0191]
Table 4-9
[0192]
Table 4-10
[0193]
Table 4-11
[0194]
Table 4-12
[0195]
Table 4-13
[0196]
Table 4-14
[0197]
Table 4-15
[0198]
Table 4-16
[0199] In some embodiments, Cas proteins require a protospacer adjacent motif (PAM) to be present within or adjacent to the target DNA sequence to which the Cas protein binds and / or functions. In some embodiments, the PAM is or includes, from 5' to 3', NGG, YG, NNGRRT, NNNRRT, NGA, TYCV, TATV, NTTN, or NNNGATT, where N represents any nucleotide, Y represents C or T, R represents A or G, and V represents A, C, or G. In some embodiments, the Cas protein is a protein listed in Table 3 or 4. In some embodiments, the Cas protein comprises one or more mutations that alter its PAM. In some embodiments, the Cas protein comprises the following mutations: E1369R, E1449H, and R1556A, or analogous substitutions for the amino acids corresponding to these positions. In some embodiments, the Cas protein comprises E782K, N968K, and R1015H mutations or analogous substitutions for the amino acids corresponding to said positions. In some embodiments, the Cas protein comprises D1135V, R1335Q, and T1337R mutations or analogous substitutions for the amino acids corresponding to said positions. In some embodiments, the Cas protein comprises S542R and K607R mutations or analogous substitutions for the amino acids corresponding to said positions. In some embodiments, the Cas protein comprises S542R, K548V, and N552R mutations or analogous substitutions for the amino acids corresponding to said positions. Exemplary advances in engineering Cas enzymes to recognize engineered PAM sequences are reviewed in Collias et al. Nature Communications 12:555 (2021), incorporated herein by reference in its entirety.
[0200] In some embodiments, the Cas protein is catalytically active and cleaves one or both strands of the target DNA site, and in some embodiments, following cleavage of the target DNA site, a modification, e.g., an insertion or deletion, is formed, e.g., by cellular repair mechanisms.
[0201] In some embodiments, the Cas protein is modified to inactivate or partially inactivate the nuclease, e.g., nuclease-deficient Cas9. While wild-type Cas9 generates double-strand breaks (DSBs) at specific DNA sequences targeted by gRNAs, several CRISPR endonucleases with modified functionality are available, e.g., partially inactivated "nickase" versions of Cas9 generate only single-strand breaks; catalytically inactive Cas9 ("dCas9") does not cleave target DNA. In some embodiments, binding of dCas9 to a DNA sequence can interfere with transcription at that site due to steric hindrance. In some embodiments, binding of dCas9 to an anchor sequence can interfere with (e.g., reduce or prevent) the formation and / or maintenance of a genome complex (e.g., ASMC). In some embodiments, the DNA-binding domain comprises a catalytically inactive Cas9, e.g., dCas9. Numerous catalytically inactive Cas9 proteins are known in the art. In some embodiments, dCas9 comprises mutations, e.g., D10A and H840A or N863A mutations, within each endonuclease domain of the Cas protein. In some embodiments, a catalytically inactive or partially inactive CRISPR / Cas domain comprises a Cas protein comprising one or more mutations, e.g., one or more of the mutations listed in Table 3. In some embodiments, a Cas protein listed in a given row of Table 3 comprises one, two, three, or all of the mutations listed in the same row of Table 3. In some embodiments, for example, a Cas protein not listed in Table 3 comprises one, two, three, or all of the mutations listed in a row of Table 3, or corresponding mutations at corresponding sites in the Cas protein.
[0202] In some embodiments, catalytically inactive, e.g., dCas9, or partially inactivated Cas9 proteins comprise a D11 mutation (e.g., a D11A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins comprise a H969 mutation (e.g., a H969A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins comprise a N995 mutation (e.g., a N995A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, comprise mutations at one, two, or three of positions D11, H969, and N995 (e.g., a D11A, H969A, and N995A mutations) or an analogous substitution for the amino acid corresponding to said positions.
[0203] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises a D10 mutation (e.g., a D10A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises a H557 mutation (e.g., a H557A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a D10 mutation (e.g., a D10A mutation) and a H557 mutation (e.g., a H557A mutation) or an analogous substitution for the amino acid corresponding to said position.
[0204] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises a D839 mutation (e.g., a D839A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises a H840 mutation (e.g., a H840A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises a N863 mutation (e.g., a N863A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, comprises a D10 mutation (e.g., D10A), a D839 mutation (e.g., D839A), an H840 mutation (e.g., H840A), and an N863 mutation (e.g., N863A) or an analogous substitution for the amino acids corresponding to the positions.
[0205] In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises an E993 mutation (e.g., an E993A mutation) or an analogous substitution for the amino acid corresponding to said position.
[0206] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises a D917 mutation (e.g., a D917A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises an E1006 mutation (e.g., an E1006A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises a D1255 mutation (e.g., a D1255A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a D917 mutation (e.g., D917A), an E1006 mutation (e.g., E1006A), and a D1255 mutation (e.g., D1255A) or an analogous substitution for the amino acid corresponding to said position.
[0207] In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D16 mutation (e.g., a D16A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D587 mutation (e.g., a D587A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the partially inactivated Cas domain has nickase activity. In some embodiments, the partially inactivated Cas9 domain is a Cas9 nickase domain. In some embodiments, the catalytically inactive Cas domain or inactive Cas domain does not form a detectable double-stranded break. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a H588 mutation (e.g., a H588A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises an N611 mutation (e.g., an N611A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, comprises a D16 mutation (e.g., D16A), a D587 mutation (e.g., D587A), an H588 mutation (e.g., H588A), and an N611 mutation (e.g., N611A) or an analogous substitution for the amino acid corresponding to said position.
[0208] In some embodiments, the DNA binding domain or endonuclease domain can comprise a Cas molecule that includes or is linked (e.g., covalently) to a gRNA (e.g., a template nucleic acid that includes a gRNA, e.g., a template RNA).
[0209] In some embodiments, the endonuclease domain or DNA-binding domain comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises a modified SpCas9. In some embodiments, the modified SpCas9 comprises a modification that alters its protospacer adjacent motif (PAM) specificity. In some embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In some embodiments, the modified SpCas9 comprises one or more amino acid substitutions, e.g., at one or more of L1111, D1135, G1218, E1219, A1322, or R1335, e.g., selected from L1111R, D1135V, G1218R, E1219F, A1322R, and R1335V. In some embodiments, the modified SpCas9 comprises the amino acid substitution T1337R and one or more additional amino acid substitutions selected from L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof. In some embodiments, the modified SpCas9 comprises (i) one or more amino acid substitutions selected from D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337; and (ii) one or more additional amino acid substitutions selected from L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof.
[0210] In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas domain, such as a Cas9 domain. In several embodiments, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas). In several embodiments, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas). In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas9 domain of Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises S. pyogenes or S. thermophilus Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas9 sequence, e.g., as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737 (incorporated herein by reference). In some embodiments, the endonuclease domain or DNA binding domain comprises the HNH nuclease subdomain and / or RuvC1 subdomain of a Cas, e.g., Cas9, or a variant thereof, as described herein.In some embodiments, the endonuclease domain or DNA-binding domain comprises Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas polypeptide (e.g., an enzyme), or a functional fragment thereof. In several embodiments, the Cas polypeptide (e.g., an enzyme) is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3 , Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, C sm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Cs a5, a type II Cas effector protein, a type V Cas effector protein, a type VI Cas effector protein, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, hyper accurate Cas9 mutant (HypaCas9), homologs thereof, modified or engineered versions thereof, and / or functional fragments thereof.In some embodiments, the Cas9 comprises one or more substitutions selected from, for example, H840A, D10A, P475A, W476A, N477A, D1125A, W1126A, and D1127A. In some embodiments, the Cas9 comprises one or more mutations at a position selected from D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, for example, one or more substitutions selected from D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. In some embodiments, the endonuclease domain or DNA binding domain is selected from the group consisting of Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Belliella baltica, Psychroflexus torquis, Staphylococcus thermophilus, Listeria innocua, Campylobacter jejuni, Neisseria meningitidis, and the like. meningitidis, Streptococcus pyogenes, or Staphylococcus aureus, or functional fragments or variants thereof.
[0211] In some embodiments, the endonuclease domain or DNA binding domain comprises a Cpf1 domain comprising one or more substitutions selected from, e.g., D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, and D917A / E1006A / D1255A, e.g., at positions D917, E1006A, D1255, or any combination thereof.
[0212] In some embodiments, the endonuclease domain or DNA binding domain is spCas9, spCas9-VRQ R、 spCas9-VRE R、 xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSE R、 spCas9-LRKIQ K、 or spCas9-LRVSQ L include.
[0213] In some embodiments, the engineered polypeptide has an endonuclease domain that includes a Cas9 nickase, e.g., Cas9 H840A. In embodiments, Cas9 H840A has the following amino acid sequence: Cas9 Nickase (H840A): [ka]
[0214] In some embodiments, the recombinant polypeptide comprises a dCas9 sequence comprising a D10A and / or H840A mutation, for example, the sequence: [ka]
[0215] TAL effectors and zinc finger nucleases In some embodiments, the endonuclease domain or DNA-binding domain comprises a TAL effector molecule. A TAL effector molecule, for example, a TAL effector molecule that specifically binds to a DNA sequence, typically comprises multiple TAL effector domains or fragments thereof, and optionally one or more additional portions of a naturally occurring TAL effector (for example, the N-terminus and / or C-terminus of multiple TAL effector domains). Many TAL effectors are known to those skilled in the art and are commercially available, for example, from Thermo Fisher Scientific.
[0216] Naturally occurring TALEs are natural effector proteins secreted by numerous species of bacterial pathogens, including the plant pathogen Xanthomonas, that regulate gene expression in host plants and promote bacterial colonization and survival. The specific binding of TAL effectors is typically based on a central repeat domain (repeated variable dinucleotide, RVD domain) of tandemly arranged, nearly identical repeats of 33 or 34 amino acids.
[0217] Members of the TAL effector family differ primarily in the number and order of their repeats. The number of repeats typically ranges from 1.5 to 33.5 repeats, with the C-terminal repeats usually being shorter in length (e.g., approximately 20 amino acids) and commonly referred to as "half-repeats." Each repeat in a TAL effector is generally characterized by a one-repeat-to-one base-pair correlation (one repeat recognizes one base pair in the target gene sequence), with different repeat types exhibiting different base-pair specificities. Generally, a decrease in the number of repeats weakens the protein-DNA interaction. It has been shown that several 6.5 repeats are sufficient to activate transcription of a reporter gene (Scholze et al., 2010).
[0218] The variation between repeats occurs primarily at amino acid positions 12 and 13, which are therefore termed "hypervariable" and are responsible for the specificity of the interaction with the target DNA promoter sequence, as shown in Table 5, which lists exemplary repeat variable dinucleotides (RVDs) and their correspondence to nucleobase targets.
[0219] [Table 5]
[0220] Therefore, it is possible to modify the repeats of TAL effectors to target specific DNA sequences. Furthermore, studies have shown that RVD NK can target G. Furthermore, the target sites of TAL effectors tend to contain a T adjacent to the 5' base targeted by the first repeat, although the exact mechanism of this recognition is unknown. Over 113 TAL effector sequences are known to date. Non-limiting examples of TAL effectors from Xanthomonas include Hax2, Hax3, Hax4, AvrXa7, AvrXa10, and AvrBs3.
[0221] Thus, the TAL effector domain of the TAL effector molecules described herein can be derived from a TAL effector from any bacterial species (e.g., Xanthomonas species, such as African strains of Xanthomonas oryzae pv. oryzae (Yu et al. 2011), Xanthomonas campestris pv. raphani strain 756C, and Xanthomonas oryzae pv. oryzicola BLS256 (Bogdanove et al. 2011)). In some embodiments, the TAL effector domain also comprises an RVD domain and flanking sequences (sequences N- and / or C-terminal to the RVD domain) from a naturally occurring TAL effector. It may contain more or fewer RVD repeats than the naturally occurring TAL effector. TAL effector molecules can be designed to target a given DNA sequence based on the above codes or others known in the art. The number of TAL effector domains (e.g., repeats (monomers or modules)) and their specific sequences can be selected based on the desired DNA target sequence. For example, TAL effector domains, e.g., repeats, can be removed or added as appropriate for a particular target sequence. In some embodiments, a TAL effector molecule of the invention comprises between 6.5 and 33.5 TAL effector domains, e.g., repeats. In some embodiments, a TAL effector molecule of the invention comprises between 8 and 33.5 TAL effector domains, e.g., repeats, for example, between 10 and 25 TAL effector domains, e.g., repeats, for example, between 10 and 14 TAL effector domains, e.g., repeats.
[0222] In some embodiments, a TAL effector molecule comprises a TAL effector domain that corresponds to a perfect match with the DNA target sequence. In some embodiments, mismatches between repeats and target base pairs on the DNA target sequence are tolerated as long as they allow the polypeptide comprising the TAL effector molecule to function. Generally, TALE binding is inversely correlated with the number of mismatches. In some embodiments, a TAL effector molecule of a polypeptide of the present invention comprises at most seven mismatches, six mismatches, five mismatches, four mismatches, three mismatches, two mismatches, or one mismatch with the target DNA sequence, and optionally no mismatches. While not intending to be bound by a particular theory, generally, as the number of TAL effector domains in a TAL effector molecule decreases, a reduced number of mismatches is not only tolerated but also allows the polypeptide comprising the TAL effector molecule to function. Binding affinity is thought to depend on the sum of matching repeat-DNA combinations. For example, a TAL effector molecule with 25 or more TAL effector domains may be able to tolerate up to seven mismatches.
[0223] In addition to the TAL effector domain, the TAL effector molecules of the present invention may contain additional sequences derived from naturally occurring TAL effectors. The length of the C-terminal and / or N-terminal sequences included on either side of the TAL effector domain portion of the TAL effector molecule can vary and can be selected by those skilled in the art based on, for example, the study of Zhang et al. (2011). Zhang et al. characterized several C-terminal and N-terminal truncation mutants in proteins based on Hax3-derived TAL effectors and identified key elements that contribute to optimal binding to target sequences and, therefore, transcriptional activation. Generally, transcriptional activity was found to be inversely correlated with the length of the N-terminus. Regarding the C-terminus, key elements in the DNA-binding residues within the first 68 amino acids of the Hax3 sequence were identified. Thus, in some embodiments, the first 68 amino acids on the C-terminal side of the TAL effector domain of a naturally occurring TAL effector are included in the TAL effector molecule. Thus, in one embodiment, a TAL effector molecule comprises: 1) one or more TAL effector domains derived from a naturally occurring TAL effector; 2) at least 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260, 270, 280 or more amino acids from a naturally occurring TAL effector N-terminal to the TAL effector domain; and / or 3) at least 68, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260 or more amino acids from a naturally occurring TAL effector C-terminal to the TAL effector domain.
[0224] In some embodiments, the endonuclease domain or DNA-binding domain is or comprises a zinc finger molecule. The zinc finger molecule comprises a zinc finger protein, such as a naturally occurring zinc finger protein or a modified zinc finger protein, or a fragment thereof. Many zinc finger proteins are known to those skilled in the art and are commercially available, for example, from Sigma-Aldrich.
[0225] In some embodiments, the zinc finger molecule comprises a non-naturally occurring zinc finger protein engineered to bind to a selected target DNA sequence (see, e.g., Beerli, et al. (2002) Nature Biotechnol. 20:135-141; Pabo, et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan, et al. (2001) Nature Biotechnol. 19:656-660; Segal, et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo, et al. al. (2000) Curr. Opin. Struct. Biol. 10:411-416; U.S. Patent Nos. 6,453,242; 6,534,261; 6,599,692; 6,503,717; 6,689,558; 7,030,215; 6,794,136; 7,067,317 Nos. 7,262,054; 7,070,934; 7,361,635; 7,253,273; and U.S. Patent Application Publication Nos. 2005 / 0064474; 2007 / 0218528; and 2005 / 0267061 (all of which are incorporated by reference in their entirety).
[0226] The engineered zinc finger proteins may have novel binding specificities compared to naturally occurring zinc finger proteins. Engineering methods include, but are not limited to, rational design and various types of selection. Rational design, for example, involves the use of a database containing triplet (or quadruplet) nucleotide sequences and individual zinc finger amino acid sequences, where each triplet or quadruplet nucleotide sequence is associated with one or more amino acid sequences of zinc fingers that bind to a particular triplet or quadruplet sequence. See, for example, U.S. Patent Nos. 6,453,242 and 6,534,261 (incorporated herein by reference in their entireties).
[0227] Exemplary selection methods, including phage display and two-hybrid systems, are disclosed in U.S. Patent Nos. 5,789,538; 5,925,523; 6,007,988; 6,013,453; 6,410,248; 6,140,466; 6,200,759; and 6,242,568; as well as International Patent Publications WO 98 / 37186; WO 98 / 53057; WO 00 / 27878; and WO 01 / 88197 and GB 2,338,237. Furthermore, increased binding specificity in zinc finger proteins is described, for example, in International Patent Publication WO 02 / 077227.
[0228] Furthermore, as disclosed in these and other references, zinc finger domains and / or multi-fingered zinc finger proteins can be linked together using any suitable linker sequence, including, for example, linkers of five or more amino acids in length. For exemplary linker sequences of six or more amino acids in length, see also U.S. Pat. Nos. 6,479,626; 6,903,185; and 7,153,949. The proteins described herein can include any combination of suitable linkers between the individual zinc fingers of the protein. Furthermore, increased binding specificity in zinc finger binding domains is described, for example, in co-owned International Patent Publication WO 02 / 077227.
[0229] Zinc finger proteins and methods for the design and construction of fusion proteins (and polynucleotides encoding same) are known to those of skill in the art and include those disclosed in U.S. Patent Nos. 6,140,0815; 789,538; 6,453,242; 6,534,261; 5,925,523; 6,007,988; 6,013,453; and 6,200,759; International Patent Publication Nos. WO 95 / 19431; WO 96 / 19432; and WO 03 / 016496.
[0230] Furthermore, as disclosed in these and other references, zinc finger proteins and / or multi-fingered zinc finger proteins can be linked together, e.g., as a fusion protein, using any suitable linker sequence, including, for example, linkers of 5 or more amino acids in length. See also U.S. Pat. Nos. 6,479,626; 6,903,185; and 7,153,949 for exemplary linker sequences of 6 or more amino acids in length. The zinc finger molecules described herein can include any combination of suitable linkers between the individual zinc finger proteins and / or multi-fingered zinc finger proteins of the zinc finger molecule.
[0231] In certain embodiments, the DNA-binding domain or endonuclease domain comprises a zinc finger molecule comprising an engineered zinc finger protein that binds (in a sequence-specific manner) to a target DNA sequence. In some embodiments, the zinc finger molecule comprises one zinc finger protein or a fragment thereof. In other embodiments, the zinc finger molecule comprises multiple zinc finger proteins (or fragments thereof), for example, 2, 3, 4, 5, 6, or more zinc finger proteins (and optionally at most 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 zinc finger proteins). In some embodiments, the zinc finger molecule comprises at least three zinc finger proteins. In some embodiments, the zinc finger molecule comprises four, five, or six fingers. In some embodiments, the zinc finger molecule comprises eight, nine, ten, eleven, or twelve fingers. In some embodiments, a zinc finger molecule comprising three zinc finger proteins recognizes a target DNA sequence comprising 9 or 10 nucleotides. In some embodiments, a zinc finger molecule comprising four zinc finger proteins recognizes a target DNA sequence comprising 12 to 14 nucleotides, and in some embodiments, a zinc finger molecule comprising six zinc finger proteins recognizes a target DNA sequence comprising 18 to 21 nucleotides.
[0232] In some embodiments, the zinc finger molecule comprises a bimanual zinc finger protein. A bimanual zinc finger protein is a protein in which two clusters of zinc finger proteins are separated by an intervening amino acid, such that the two zinc finger domains bind to two discontinuous target DNA sequences. An example of a bimanual zinc finger binding protein is SIP1, in which a cluster of four zinc finger proteins is located at the amino terminus of the protein and a cluster of three zinc finger proteins is located at the carboxyl terminus (see Remade, et al. (1999) EMBO Journal 18(18):5073-5084). Each cluster of zinc fingers in these proteins can bind to a unique target sequence, and the spacing between the two target sequences can include multiple nucleotides.
[0233] Linker In some embodiments, the engineered polypeptide can include a linker, e.g., a peptide linker, e.g., a linker described in Table 6. In some embodiments, the engineered polypeptide includes, from N-terminal to C-terminal, a Cas domain (e.g., a Cas domain of Table 3), a linker of Table 6 (or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto), and an RT domain (e.g., an RT domain of Table 2). In some embodiments, the engineered polypeptide can include a flexible linker between the endonuclease and the RT domain, e.g., the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 4006) In some embodiments, the RT domain of the recombinant polypeptide may be located C-terminal to the endonuclease domain. In some embodiments, the RT domain of the recombinant polypeptide may be located N-terminal to the endonuclease domain.
[0234] [Table 6-1]
[0235] [Table 6-2]
[0236] [Table 6-3]
[0237] [Table 6-4]
[0238] In some embodiments, the linker of the genetically engineered polypeptide is (SGGS) n (SEQ ID NO: 4025), (GGGS) n (SEQ ID NO: 4026), (GGGGS) n (SEQ ID NO: 4027), (G) n , (EAAAK) n (SEQ ID NO: 4028), (GGS) n , or (XP) n The motif comprises a motif selected from:
[0239] Selection of recombinant polypeptides by pooled screening Candidate recombinant polypeptide can be screened to evaluate the gene editing ability of candidate.For example, can use the RNA recombinant system designed for the targeted editing of coding sequence in human genome.In certain embodiments, this recombinant system can be used with pool screening method.
[0240] For example, a library of candidate recombinant polypeptides and template guide RNAs (tgRNAs) can be introduced into mammalian cells to test the gene-editing capabilities of the candidates by pooled screening techniques. In certain embodiments, the library of candidate recombinant polypeptides is introduced into mammalian cells, followed by introduction of tgRNAs into the cells.
[0241] Representative, non-limiting examples of mammalian cells that can be used in the screen include HEK293T cells, U2OS cells, HeLa cells, HepG2 cells, Huh7 cells, K562 cells, or iPS cells.
[0242] The candidate engineered polypeptides can include 1) a Cas-nuclease, e.g., a wild-type Cas nuclease, e.g., a wild-type Cas9 nuclease, a mutant Cas nuclease, e.g., a Cas nickase, e.g., a Cas9 nickase such as Cas9 N863A nickase, or a Cas nuclease selected from Table 3 or Table 4, 2) a peptide linker, e.g., a sequence from Table D or Table 6, which can exhibit varying degrees of length, flexibility, hydrophobicity, and / or secondary structure; and 3) a reverse transcriptase (RT), e.g., an RT domain from Table D or Table 2. The candidate engineered polypeptide library includes a plurality of different candidate engineered polypeptides that differ from each other with respect to one, two, or all three of the Cas nuclease, peptide linker, or RT domain components, or a plurality of nucleic acid expression vectors encoding such candidate engineered polypeptides.
[0243] For screening of candidate recombinant polypeptides, a two-component system comprising a recombinant polypeptide component and a tgRNA component can be used. The recombinant component can include, for example, an expression vector, such as an expression plasmid or lentiviral vector encoding the candidate recombinant polypeptide, including a human codon-optimized nucleic acid encoding the candidate recombinant polypeptide, such as the Cas-linker-RT fusion described above. In certain embodiments, a lentiviral cassette is used that includes: (i) a promoter for expression in mammalian cells, such as a CMV promoter; (ii) a candidate recombinant library, such as a Cas-linker-RT fusion comprising a Cas nuclease from Table 3 or Table 4, a peptide linker from Table 6, and an RT from Table 2, e.g., a Cas-linker-RT fusion such as those in Table D; (iii) a self-cleaving polypeptide, such as a T2A peptide; (iv) a marker allowing selection in mammalian cells, such as a puromycin resistance gene; and (v) a termination signal, such as a polyA tail.
[0244] The tgRNA component can include a tgRNA or an expression vector, e.g., an expression plasmid that generates the tgRNA and drives expression of the tgRNA using, e.g., a U6 promoter, where the tgRNA is a non-coding RNA sequence that is recognized by Cas, localizing it to the genomic locus of interest, and that templates reverse transcription of the desired edit into the genome via the RT domain.
[0245] To prepare a pool of cells expressing recombinant polypeptide library candidates, mammalian cells, e.g., HEK293T or U2OS cells, can be transduced with a pooled recombinant polypeptide candidate expression vector preparation, e.g., a lentiviral preparation of the recombinant candidate polypeptide library. In certain embodiments, lentiviral plasmids are used, and HEK293 Lenti-X cells are seeded in 15 cm plates (approximately 12 x 10 cells) prior to lentiviral plasmid transfection. 6In such an embodiment, lentiviral plasmid transfection can be performed using Lentiviral Packaging Mix (Biosettia), and transfection of plasmid DNA for the recombinant candidate library can be performed using Lipofectamine 2000 and Opti-MEM medium according to the manufacturer's protocol. In such an embodiment, extracellular DNA can be removed by a complete medium change the next day, and virus-containing medium can be collected 48 hours later. The lentiviral medium can be concentrated using a Lenti-X Concentrator (TaKaRa Biosciences), and 5 mL lentiviral aliquots can be made and stored at -80°C. Lentiviral titer determination can be performed after selection, for example, by counting colony-forming units after puromycin selection.
[0246] To monitor gene editing of target DNA, mammalian cells, such as HEK293T or U2OS cells carrying target DNA, can be used. In other embodiments for monitoring gene editing of target DNA, mammalian cells, such as HEK293T or U2OS cells carrying a target DNA genomic landing pad, can be used. In certain embodiments, the target DNA genomic landing pad can contain a gene to be edited for the treatment of a disease or disorder of interest. In other specific embodiments, the target DNA is a genetic sequence that expresses a protein exhibiting a detectable characteristic that can be monitored to determine whether gene editing has occurred. For example, in certain embodiments, blue fluorescent protein (BFP)- or green fluorescent protein (GFP)-expressing genomic landing pads are used. In certain embodiments, mammalian cells, such as HEK293T or U2OS cells carrying target DNA, e.g., a target DNA genomic landing pad, are seeded into culture plates at 500x to 3000x cells per recombinant library candidate and transduced at a multiplicity of infection (MOI) of 0.2 to 0.3 to minimize the number of infections per cell. Puromycin (2.5 μg / mL) can be added 48 hours after infection to allow for selection of infected cells. In such an embodiment, the cells are placed under puromycin selection for at least 7 days and then scaled up for tgRNA introduction, e.g., tgRNA electroporation.
[0247] To confirm whether gene editing occurs, mammalian cells containing the target DNA to be edited can be infected with a candidate recombinant polypeptide library and then transfected with a tgRNA designed for use in editing the target DNA. The cells can then be analyzed, for example, by cell sorting and sequence analysis, to determine whether editing of the target locus occurred according to the designed results, or whether no editing or incomplete editing occurred.
[0248] In certain embodiments, to confirm whether genome editing occurs, BFP- or GFP-expressing mammalian cells, such as HEK293T or U2OS cells, may be infected with a recombinant library candidate and then transfected or electroporated at 250,000 cells / well with a tgRNA plasmid or RNA, e.g., 200 ng of a tgRNA plasmid designed to convert BFP to GFP or GFP to BFP, at a cell number that ensures >250×-1000× coverage per library candidate. In such embodiments, the genome editing ability of various constructs in this assay may be assessed by sorting cells by fluorescence-activated cell sorting (FACS) for the expression of color-converted fluorescent proteins (FPs) 4-10 days after electroporation. Cells are sorted and collected into distinct populations: non-edited cells (showing the original fluorescent protein signal), edited cells (showing the converted fluorescent protein signal), and incompletely edited cells (showing no fluorescent protein signal). A sample of unsorted cells can also be collected as an input population to determine candidate enrichment during analysis.
[0249] To determine whether the recombinant library candidates exhibit genome editing capabilities in the assay, genomic DNA (gDNA) is collected from the sorted cell populations and analyzed by sequencing the recombinant library candidates in each population. Briefly, the recombinant candidates are amplified from the genome using primers specific to the recombinant polypeptide expression vector, e.g., a lentiviral cassette, and amplified in a second round of PCR to dilute the genomic DNA, which can then be sequenced, for example, by a next-generation sequencing platform. After quality control of the sequencing reads, reads of at least about 1500 nucleotides, and generally no more than about 3200 nucleotides, are mapped to the recombinant polypeptide library sequence, and those containing a minimum of about 80% match with the library sequence are considered to have successfully aligned with a given candidate for this pooled screen. To identify candidates capable of gene editing in the assay, for example, editing BFP to GFP or GFP to BFP, the read count of each library candidate in the edited population is compared to its read count in the initial unsorted population.
[0250] For pooled screening, genetically modified candidates with genome editing capabilities are identified based on the enrichment of the edited (converted FP) population relative to the unsorted (input) cells. In some embodiments, an enrichment of at least 1.0, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or at least 100 times the input indicates potentially useful gene editing activity, e.g., at least 2-fold enrichment. In some embodiments, enrichment is converted to a log value by taking the log base 2 of the enrichment ratio. In some embodiments, a log enrichment score of at least 0, 1, 2, 3, 4, 5, 5.5, 6.0, 6.2, 6.3, 6.4, 6.5, or at least 6.6 indicates potentially useful gene editing activity, e.g., a log enrichment score of at least 1.0. In certain embodiments, the enrichment value observed for a transgenic candidate can be compared to the enrichment value observed under similar conditions using a reference, for example, element ID number: 17380.
[0251] In some embodiments, multiple tgRNAs can be used to screen recombinant candidate libraries.In certain embodiments, multiple tgRNAs can be used to optimize template / Cas-linker-RT fusion pairs, for example, for gene editing of specific target genes, for example, gene targets for disease treatment.In certain embodiments, pooling method for screening recombinant candidate can be carried out using many different tgRNAs in array format.
[0252] In some embodiments, multiple types of edits, for example, insertions, substitutions, and / or deletions of different lengths, may be used to screen a recombinant candidate library.
[0253] In some embodiments, multiple target sequences, e.g., different fluorescent proteins, may be used to screen a transgenic candidate library. In some embodiments, multiple target sequences, e.g., different fluorescent proteins, may be used to screen a transgenic candidate library. In some embodiments, multiple cell types, e.g., HEK293T or U2OS, may be used to screen a transgenic candidate library. One skilled in the art will understand that a given candidate may exhibit altered editing capabilities or increased or decreased observable or useful activity across different conditions, including tgRNA sequence (e.g., nucleotide modification, PBS length, RT template length), target sequence, target location, type of editing, location of mutation relative to the first strand nick of the transgenic polypeptide, or cell type. Thus, in some embodiments, a transgenic library candidate is screened across multiple parameters, e.g., using at least two different tgRNAs in at least two cell types, and gene editing activity is identified by enrichment in any single condition. In other embodiments, candidates with more robust activity across different tgRNAs and cell types are identified by enrichment in at least two conditions, e.g., all conditions screened. For clarity, candidates found to show little to no enrichment under any given condition are not presumed to be inactive across all conditions and can be screened using different parameters or reconstituted at the polypeptide level, for example, by exchanging, shuffling, or altering domains (e.g., RT domains), linkers, or other signals (e.g., NLS).
[0254] Exemplary Cas9-Linker-RT Fusion Sequences In some embodiments, the engineered polypeptide comprises a linker sequence and an RT sequence. In some embodiments, the engineered polypeptide comprises a linker sequence listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the engineered polypeptide comprises an RT domain amino acid sequence listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the recombinant polypeptide comprises a linker sequence listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and an amino acid sequence of an RT domain listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the recombinant polypeptide comprises (i) a linker sequence listed in a row of Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and (ii) an amino acid sequence of an RT domain listed in the same row of Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0255] Localization sequences for transgenic systems In certain embodiments, the gene editor system RNA further comprises a subcellular localization sequence, e.g., a nuclear localization sequence (NLS). In some embodiments, the engineered polypeptide comprises an NLS contained in SEQ ID NO: 4000 and / or SEQ ID NO: 4001, or an NLS having an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0256] The nuclear localization sequence can be an RNA sequence that promotes entry of the RNA into the nucleus. In certain embodiments, the nuclear localization signal is located on the template RNA. In certain embodiments, the recombinant polypeptide is encoded on a first RNA, the template RNA is a second, separate RNA, and the nuclear localization signal is located on the template RNA, but not on the RNA encoding the recombinant polypeptide. Without intending to be bound by theory, in some embodiments, the RNA encoding the recombinant polypeptide is targeted primarily to the cytoplasm to promote its translation, while the template RNA is targeted primarily to the nucleus to promote insertion into the genome. In some embodiments, the nuclear localization signal is located at the 3' end, 5' end, or within an internal region of the template RNA. In some embodiments, the nuclear localization signal is 3' to the heterologous sequence (e.g., directly 3' to the heterologous sequence) or 5' to the heterologous sequence (e.g., directly 5' to the heterologous sequence). In some embodiments, the nuclear localization signal is positioned outside the 5'UTR or outside the 3'UTR of the template RNA. In some embodiments, the nuclear localization signal is positioned between the 5'UTR and the 3'UTR, and optionally, the nuclear localization signal is not transcribed by the transgene (e.g., the nuclear localization signal is antisense-oriented or downstream of a transcription termination signal or polyadenylation signal). In some embodiments, the nuclear localization sequence is located within an intron. In some embodiments, multiple identical or different nuclear localization signals are present in an RNA, e.g., a template RNA. In some embodiments, the nuclear localization signal is less than 5 bp, 10 bp, 25 bp, 50 bp, 75 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1000 bp in length. Various RNA nuclear localization sequences can be used.For example, Lubelsky and Ulitsky, Nature 555(107-111), 2018, describe the RNA sequence that drives RNA to localize in the nucleus.In some embodiments, the nuclear localization signal is SINE-derived nuclear RNA localization (SIRLOIN) signal.In some embodiments, the nuclear localization signal binds to a nuclear-enriched protein. In some embodiments, the nuclear localization signal binds to an HNRNPK protein. In some embodiments, the nuclear localization signal is rich in pyrimidines, for example, a C / T-rich, C / U-rich, C-rich, T-rich, or U-rich region. In some embodiments, the nuclear localization signal is derived from a long untranslated RNA. In some embodiments, the nuclear localization signal is derived from the MALAT1 long untranslated RNA or the 600-nucleotide M region of MALAT1 (described in Miyagawa et al., RNA 18, (738-751), 2012). In some embodiments, the nuclear localization signal is derived from the BORG long untranslated RNA or is an AGCCC motif (described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014)). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a retrovirus.
[0257] In some embodiments, the polypeptides described herein comprise one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, e.g., nuclear localization sequences (NLSs). In some embodiments, the NLS is a bipartite NLS. In some embodiments, the NLS promotes entry of a protein comprising the NLS into the cell nucleus. In some embodiments, the NLS is fused to the N-terminus of a genetically engineered polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of a genetically engineered polypeptide. In some embodiments, the NLS is fused to the N-terminus or C-terminus of a Cas domain. In some embodiments, a linker sequence is disposed between the NLS and an adjacent domain of the genetically engineered polypeptide.
[0258] In some embodiments, the NLS has the amino acid sequence MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 4009), PKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 4010), RKSGKIAAIWKRPRKPKKKRKV (SEQ ID NO: 4011), KRTADGSEFESPKKKRKV (SEQ ID NO: 4642) ,KKTELQTTNAENKTKKL (SEQ ID NO: 4643) , or KRGINDRNFWRGENGRKTR (SEQ ID NO: 4644) ,KRPAATKKAGQAKKKK (SEQ ID NO: 4645) or functional fragments or variants thereof. Exemplary NLS sequences are also described in PCT / EP 2000 / 011690, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises an amino acid sequence as disclosed in Table 7. The NLSs in this table can be utilized in one or more copies in a polypeptide at one or more positions in the polypeptide to improve subcellular localization to the nucleus, for example, one, two, three, or more copies of the NLS within the N-terminal domain, between peptide domains, within the C-terminal domain, or in a combination of positions. Multiple unique sequences can be used in a single polypeptide. The sequences can be monopartite or bipartite in nature, for example, with one or two stretches of basic amino acids, or can be used as chimeric bipartite sequences. Sequence references correspond to UniProt accession numbers, except where indicated as SeqNLS, for sequences retrieved using a subcellular localization prediction algorithm (Lin et al. BMC Bioinformat 13:157 (2012), incorporated herein by reference in its entirety).
[0259] [Table 7-1]
[0260] [Table 7-2]
[0261] [Table 7-3]
[0262] [Table 7-4]
[0263] In some embodiments, the NLS is a bi-clad NLS. A bi-clad NLS typically comprises two basic amino acid clusters (which may be, for example, about 10 amino acids long) separated by a spacer sequence. A mono-clad NLS typically lacks a spacer. An example of a bi-clad NLS is the sequence KR[PAATKKAGQA]KKKK. (SEQ ID NO: 4645) (spacer in parentheses). Another exemplary bipartite NLS has the sequence PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 4016). Exemplary NLSs are described in WO2020051561 (incorporated herein by reference in its entirety, including the disclosure regarding nuclear localization sequences).
[0264] In certain embodiments, a gene editor system polypeptide (e.g., a transgenic polypeptide described herein) further comprises a subcellular localization sequence, e.g., a nuclear localization sequence and / or a nucleolar localization sequence. The nuclear localization sequence and / or nucleolar localization sequence can be an amino acid sequence that promotes entry of the protein into the nucleus and / or nucleolus, where it can promote integration of the heterologous sequence into the genome. In certain embodiments, a gene editor system polypeptide (e.g., a transgenic polypeptide described herein, by way of example) further comprises a nucleolar localization sequence. In certain embodiments, the transgenic polypeptide is encoded on a first RNA, the template RNA is a second separate RNA, and the nucleolar localization signal is encoded on the RNA encoding the transgenic polypeptide but not on the template RNA. In some embodiments, the nucleolar localization signal is located at the N-terminus, C-terminus, or within an internal region of the polypeptide. In some embodiments, multiple nucleolar localization signals, either the same or different, are used. In some embodiments, the nuclear localization signal is less than 5, 10, 25, 50, 75, or 100 amino acids in length. Nucleolar localization signals of various polypeptides can be used. For example, Yang et al., Journal of Biomedical Science 22, 33 (2015), describe a nuclear localization signal that also functions as a nucleolar localization signal. In some embodiments, the nucleolar localization signal can also be a nuclear localization signal. In some embodiments, the nucleolar localization signal can overlap with a nuclear localization signal. In some embodiments, the nucleolar localization signal can include a stretch of basic residues. In some embodiments, the nucleolar localization signal can be rich in arginine and lysine residues. In some embodiments, the nucleolar localization signal can be derived from a protein that is abundant in the nucleolus. In some embodiments, the nucleolar localization signal can be derived from a protein that is abundant in ribosomal RNA loci. In some embodiments, the nucleolar localization signal may be derived from a protein that binds to rRNA. In some embodiments, the nucleolar localization signal may be derived from MSP58. In some embodiments, the nucleolar localization signal may be a monoknot motif.In some embodiments, the nucleolar localization signal can be a bi-knot motif. In some embodiments, the nucleolar localization signal can be comprised of multiple mono- or bi-knot motifs. In some embodiments, the nucleolar localization signal can be comprised of a mixture of mono- and bi-knot motifs. In some embodiments, the nucleolar localization signal can be a double bi-knot motif. In some embodiments, the nucleolar localization motif can be KRASSQALGTIPKRRSSSRFIKRKK (SEQ ID NO: 4017). In some embodiments, the nucleolar localization signal can be derived from nuclear factor-κB-inducing kinase. In some embodiments, the nucleolar localization signal can be a RKKRKKK motif (SEQ ID NO: 4018) (described in Birbach et al., Journal of Cell Science, 117(3615-3624), 2004).
[0265] Evolved variants of genetically engineered polypeptides and systems In some embodiments, the present invention provides evolved variants of the genetically engineered polypeptides described herein. The evolved variants can, in some embodiments, be produced by mutagenizing a reference genetically engineered polypeptide or one of the fragments or domains contained therein. In some embodiments, one or more of the domains (e.g., the reverse transcriptase domain) are evolved. One or more of these evolved variant domains can, in some embodiments, evolve alone or together with other domains. One or more evolved variant domains can, in some embodiments, be combined with a non-evolved cognate component or an evolved variant of a cognate component (e.g., one that may have evolved in a parallel or sequential manner).
[0266] In some embodiments, the process of mutagenizing a reference engineered polypeptide, or a fragment or domain thereof, comprises mutagenizing the reference engineered polypeptide, or a fragment or domain thereof. In embodiments, the mutagenesis comprises, for example, a progressive evolution method (e.g., PACE) or a non-progressive evolution method (e.g., PANCE), as described herein. In some embodiments, the evolved engineered polypeptide, or a fragment or domain thereof, comprises one or more amino acid mutations introduced into its amino acid sequence compared to the amino acid sequence of the reference engineered polypeptide, or a fragment or domain thereof. In embodiments, the amino acid sequence mutation may comprise one or more mutated residues (e.g., conservative substitutions, non-conservative substitutions, or a combination thereof) within the amino acid sequence of the reference engineered polypeptide, for example, as a result of a change in the nucleotide sequence encoding the engineered polypeptide resulting in a change in a codon at any particular position within the coding sequence, a deletion of one or more amino acids (e.g., a truncated protein), an insertion of one or more amino acids, or any combination thereof. An evolved variant transgenic polypeptide can include variants in one or more components or domains of the transgenic polypeptide (eg, variants introduced into the reverse transcriptase domain).
[0267] In some aspects, the disclosure provides genetically engineered polypeptides, systems, kits, and methods that use or include evolved variants of genetically engineered polypeptides, e.g., use evolved variants of genetically engineered polypeptides, or genetically engineered polypeptides produced or producible by PACE or PANCE. In embodiments, the non-evolved reference genetically engineered polypeptide is a genetically engineered polypeptide disclosed herein.
[0268] The term "phage-assisted continuous evolution (PACE)," as used herein, generally refers to incremental evolution using phages as viral vectors. Examples of PACE technology include, for example, International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010, as WO 2010 / 028347; International PCT Application No. PCT / US2011 / 066747, filed December 22, 2011, and published June 28, 2012, as WO 2012 / 088381; U.S. Patent No. 9,023,594, issued May 5, 2015; U.S. Patent No. 9,771,574, issued September 26, 2017; U.S. Patent No. 9,771,574, issued July 19, 2016; No. 9,394,537, filed January 20, 2015, published September 11, 2015 as WO 2015 / 134121; U.S. Pat. No. 10,179,911, filed January 15, 2019; and International PCT Application No. PCT / US2016 / 027795, filed April 15, 2016, published October 20, 2016 as WO 2016 / 168631, each of which is incorporated herein by reference in its entirety.
[0269] The term "phage-assisted non-incremental evolution (PANCE)," as used herein, generally refers to non-incremental evolution using phages as viral vectors. Examples of PANCE technology are described, for example, in Suzuki T. et al., "Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase," Nat Chem Biol. 13(12):1261-1266 (2017), incorporated herein by reference in its entirety. Briefly, PANCE is a technique for rapid in vivo directed evolution using serial flask transfer of evolutionary selection phages (SPs) containing the gene of interest to be evolved into fresh whole host cells (e.g., E. coli cells). While the gene contained in the SP is progressively evolved, the gene inside the host cell can be held constant. Following phage propagation, an aliquot of the infected cells can be used to transfect the next flask containing the host E. coli. This process can be repeated and / or continued until the desired phenotype has evolved, for example, for as many transfers as desired.
[0270] Methods for applying PACE and PANCE to engineered polypeptides will be readily understood by those skilled in the art by reference to, inter alia, the aforementioned references. Further exemplary methods for directing the progressive evolution of genomic modified proteins or systems, e.g., in a population of host cells, using, for example, phage particles, can be applied to generate evolved variants of engineered polypeptides, or fragments or subdomains thereof. Non-limiting examples of such methods are described in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010, as WO 2010 / 028347; International PCT Application No. PCT / US2009 / 056194, filed December 22, 2011, and published June 28, 2012, as WO 2012 / 088381; , PCT / US2011 / 066747; U.S. Patent No. 9,023,594 issued May 5, 2015; U.S. Patent No. 9,771,574 issued September 26, 2017; U.S. Patent No. 9,394,537 issued July 19, 2016; WO 2015 / 134121 filed January 20, 2015, September 11, 2015 No. PCT / US2015 / 012022, published as International PCT Publication No. PCT / US2015 / 012022; U.S. Patent No. 10,179,911, issued January 15, 2019; International Patent Application No. PCT / US2019 / 37216, filed June 14, 2019, published January 31, 2019 as WO 2019 / 023680; International PCT Application No. PCT / US2016 / 027795, filed April 15, 2016, published October 20, 2016 as WO 2016 / 168631; and International Application No. PCT / US2019 / 47996, filed August 23, 2019, each of which is incorporated herein by reference in its entirety.
[0271] In some non-limiting exemplary embodiments, a method for evolving an evolved variant genetically engineered polypeptide, fragment, or domain thereof includes: (a) contacting a population of host cells with a population of viral vectors (starting genetically engineered polypeptide or fragment or domain thereof) containing a gene of interest, where (1) the host cells are suitable for infection with the viral vector; (2) the host cells express viral genes necessary for the production of viral particles; (3) the expression of at least one viral gene necessary for the production of infectious viral particles is dependent on the function of the gene of interest; and / or (4) the viral vector allows for the expression of a protein in the host cells and can be replicated by the host cells and packaged into viral particles. In some embodiments, the method includes (b) contacting the host cells with a mutagen using host cells containing mutations that enhance mutation rates (e.g., by delivering a mutant plasmid or some genomic modification—e.g., a damaged DNA proofreading polymerase, an SOS gene, e.g., UmuC, UmuD′, and / or RecA (such mutations can be under the control of an inducible promoter when bound to the plasmid), or a combination thereof). In some embodiments, the method includes (c) incubating the population of host cells under conditions that allow for viral replication and viral particle production, where host cells are removed from the host cell population and fresh, uninfected host cells are introduced into the population of host cells, thus replenishing the host cell population and forming a stream of host cells. In some embodiments, the cells are incubated under conditions that allow the gene of interest to acquire mutations. In some embodiments, the method further includes (d) isolating from the population of host cells a mutated version of the viral vector that encodes an evolved gene product (e.g., an evolved mutant genetically engineered polypeptide, or a fragment or domain thereof).
[0272] Those skilled in the art will appreciate various features that can be used within the above framework. For example, in some embodiments, the viral vector or phage is a filamentous phage, e.g., an M13 phage, e.g., an M13 selection phage. In certain embodiments, the gene required for the production of infectious viral particles is M13 gene III (gIII). In some embodiments, the phage may lack functional gIII but instead contain gI, gII, gIV, gV, gVI, gVII, gVIII, gIX, and gX. In some embodiments, the generation of infectious VSV particles includes the envelope protein VSV-G. In various embodiments, different retroviral vectors, e.g., murine leukemia virus vectors or lentiviral vectors, can be used. In some embodiments, retroviral vectors can be efficiently packaged, e.g., using VSV-G envelope proteins as a substitute for the virus's native envelope proteins.
[0273] In some embodiments, the host cells are incubated for a suitable number of viral life cycles, e.g., at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1250, at least 1500, at least 1750, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10000, or more consecutive viral life cycles, where an illustrative, non-limiting example for M13 phage is 10-20 minutes per viral life cycle. Similarly, conditions can be adjusted to control the residence time of the host cells in the population of host cells (e.g., about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 70, about 80, about 90, about 100, about 120, about 150, or about 180 minutes). The host cell population can be adjusted to control the density of host cells, or in some embodiments, the host cell density in the inflow, e.g., 10 3 cells / ml, approximately 10 4 cells / ml, approximately 10 5 cells / ml, approximately 5-10 5 cells / ml, approximately 10 6 cells / ml, approximately 5-10 6 cells / ml, approximately 10 7 cells / ml, approximately 5-10 7 cells / ml, approximately 10 8 cells / ml, approximately 5-10 8 cells / ml, approximately 10 9 cells / ml, approximately 5-10 9 cells / ml, approximately 10 10 cells / ml, or approximately 5-10 10 Cells / ml can be used to control some of the
[0274] Intein In some embodiments, as described in more detail below, an intein-N (intN) domain may be fused to the N-terminal portion of a first domain of a genetically engineered polypeptide described herein, and an intein-C (intC) domain may be fused to the C-terminal portion of a second domain of a genetically engineered polypeptide described herein, linking the N-terminal portion to the C-terminal portion, thereby linking the first and second domains. In some embodiments, the first and second domains are each independently selected from a DNA-binding domain, an RNA-binding domain, an RT domain, and an endonuclease domain.
[0275] Inteins can exist, for example, as self-splicing protein introns (e.g., peptides) that link flanking N-terminal and C-terminal exteins (e.g., fragments to be joined). Inteins can, in some cases, comprise a fragment of a protein that can be automatically excised to join the remaining fragment (extein) with a peptide bond in a process known as protein splicing. Inteins are also referred to as "protein introns." The process of inteins being automatically excised to join the remainder of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing."
[0276] In some embodiments, the inteins of a precursor protein (an intein-containing protein prior to intein-mediated protein splicing) are derived from two genes. Such inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). Thus, intein-based approaches can be used to link a first polypeptide sequence and a second polypeptide sequence together. For example, in cyanobacteria, DnaE, the catalytic subunit of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. When located as part of a first polypeptide sequence, an intein-N domain, such as that encoded by the dnaE-n gene, can link the first polypeptide sequence to a second polypeptide sequence, where the second polypeptide sequence contains an intein-C domain, such as that encoded by the dnaE-c gene. Thus, in some embodiments, a protein may be made by providing nucleic acids encoding a first and a second polypeptide sequence (e.g., a first nucleic acid molecule encoding the first polypeptide sequence and a second nucleic acid molecule encoding the second polypeptide sequence), which are introduced into a cell under conditions that allow for the production of the first and second polypeptide sequences and ligation of the first polypeptide sequence to the second polypeptide sequence via an intein-based mechanism.
[0277] The use of inteins to link heterologous protein fragments is described, for example, in Wood et al., J. Biol. Chem. 289(21);14512-9(2014), which is incorporated herein by reference in its entirety. For example, when fused to separate protein fragments, inteins IntN and IntC can recognize each other and splice from themselves and / or simultaneously ligate flanking N- and C-terminal exteins of the protein fragments to which they are fused, thereby reconstituting a full-length protein from the two protein fragments.
[0278] In some embodiments, synthetic inteins based on the dnaE intein, Cfa-N (e.g., split intein-N), and Cfa-C (e.g., split intein-C) intein pairs are used. Examples of such inteins are described, for example, in Stevens et al., J Am Chem Soc. 2016 Feb. 24;138(7):2162-5, incorporated herein by reference in its entirety. Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include the Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Pat. No. 8,394,604, incorporated herein by reference).
[0279] In some embodiments involving split Cas9s, the intein-N and intein-C domains may be fused to the N-terminal and C-terminal portions of split Cas9, respectively, for linking the N-terminal and C-terminal portions of split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., forming the structure N-[N-terminal portion of split Cas9]-[intein-N]-C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., forming the structure N-[intein-C]-[C-terminal portion of split Cas9]-C. The mechanism of intein-mediated protein splicing for linking intein-linked proteins (e.g., split Cas9) is described in Shah et al., Chem Sci. 2014;5(1):446-461 (incorporated herein by reference). Methods for designing and using inteins are known in the art and are described, for example, in WO2020051561, WO2014004336, WO2017132580, U.S. Patent Application Publication No. 20150344549, and U.S. Patent Application Publication No. 20180127780 (each of which is incorporated herein by reference in its entirety).
[0280] In some embodiments, split refers to division into two or more fragments. In some embodiments, a split Cas9 protein or split Cas9 comprises a Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a reconstituted Cas9 protein. In embodiments, the Cas9 protein is split into two fragments within a denatured region of the protein, as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or in Jiang et al. (2016) Science 351:867-871 and PDB file: 5F9R (each of which is incorporated herein by reference in its entirety). The denatured region can be determined by one or more protein structure determination techniques known in the art, including, but not limited to, X-ray crystallography, NMR spectroscopy, electron microscopy (e.g., cryoEM), and / or in silico protein modeling. In some embodiments, the protein is split into two fragments at any C, T, A, or S within the region of SpCas9 between amino acids A292-G364, F445-K483, or E565-T637, or at the corresponding position in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as protein cleavage.
[0281] In some embodiments, protein fragments range in length from about 2 to 1000 amino acids (e.g., 2 to 10, 10 to 50, 50 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 600, 600 to 700, 700 to 800, 800 to 900, or 900 to 1000 amino acids). In some embodiments, protein fragments range in length from about 5 to 500 amino acids (e.g., 5 to 10, 10 to 50, 50 to 100, 100 to 200, 200 to 300, 300 to 400, or 400 to 500 amino acids). In some embodiments, protein fragments range in length from about 20 to 200 amino acids (e.g., 20 to 30, 30 to 40, 40 to 50, 50 to 100, or 100 to 200 amino acids).
[0282] In some embodiments, a portion or fragment of the recombinant polypeptide is fused to an intein. A nuclease can be fused to the N-terminus or C-terminus of the intein. In some embodiments, a portion or fragment of the fusion protein is fused to an intein and also to an AAV capsid protein. The intein, nuclease, and capsid protein can be fused together in any configuration (e.g., nuclease-intein-capsid, intein-nuclease-capsid, capsid-intein-nuclease, etc.). In some embodiments, the N-terminus of the intein is fused to the C-terminus of the fusion protein, and the C-terminus of the intein is fused to the N-terminus of the AAV capsid protein.
[0283] In some embodiments, an endonuclease domain (e.g., a nickase Cas9 domain) is fused to intein-N, and a polypeptide containing an RT domain is fused to intein-C.
[0284] Exemplary nucleotide and amino acid sequences of intein-N domains and equivalent intein-C domains are set forth below. DnaE Intein-N DNA: [ka] DnaE Intein-N Protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN (SEQ ID NO: 4030) DnaE Intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT (SEQ ID NO: 4031) DnaE Intein-C Protein: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN (SEQ ID NO: 4032) Cfa-N DNA: [ka] Cfa-N protein: CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP (SEQ ID NO: 4034) Cfa-C DNA: [ka] Cfa-C protein: MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN (SEQ ID NO: 4036)
[0285] More domains: The recombinant polypeptide can bind to a target DNA sequence and a template nucleic acid (e.g., an RNA template), nick the target site, and write (e.g., reverse transcribe) the template into DNA, resulting in modification of the target site. In some embodiments, additional domains can be added to the polypeptide to enhance the efficiency of the process. In some embodiments, the recombinant polypeptide can include an additional DNA ligation domain to ligate the reverse-transcribed DNA to DNA at the target site. In some embodiments, the polypeptide can include a heterologous RNA-binding domain. In some embodiments, the polypeptide can include a domain with 5' to 3' exonuclease activity (e.g., where 5' to 3' exonuclease activity enhances repair of target site modifications, e.g., intended for modifications across the original genomic sequence). In some embodiments, the polypeptide can include a domain with 3' to 5' exonuclease activity, e.g., proofreading activity. In some embodiments, the writing domain, e.g., the RT domain, has 3' to 5' exonuclease activity, e.g., proofreading activity.
[0286] template nucleic acid The genetic engineering systems described herein can use template nucleic acid sequences to modify a host target DNA site. In some embodiments, the genetic engineering systems described herein transcribe an RNA sequence template into a host target DNA site using target-primed reverse transcription (TPRT). By directly recombining a DNA sequence into a host genome via reverse transcription of an RNA sequence template, the genetic engineering system can insert a sequence of interest into a target genome without the need to introduce an exogenous DNA sequence into the host cell (unlike, for example, CRISPR systems) and eliminate the exogenous DNA insertion step. The genetic engineering system can also delete sequences from the target genome or introduce replacements with a sequence of interest. Thus, the genetic engineering system provides a platform for the use of customized RNA sequence templates containing sequences of interest, e.g., sequences containing heterologous gene coding and / or functional information.
[0287] In some embodiments, the template nucleic acid comprises one or more sequences (eg, two sequences) that bind to the recombinant polypeptide.
[0288] In some embodiments, the template nucleic acid comprises a hybrid having both ribonucleotide and deoxyribonucleotide residues in the same strand.
[0289] In some embodiments, the systems or methods described herein include a single template nucleic acid (e.g., template RNA). In some embodiments, the systems or methods described herein include multiple template nucleic acids (e.g., template RNAs). For example, the systems described herein include a first RNA that includes (e.g., from 5' to 3') a sequence that binds to a genetically modified polypeptide (e.g., a DNA-binding domain and / or an endonuclease domain, e.g., a gRNA) and a sequence that binds to a target site (e.g., the second strand of a site in a target genome), and a second RNA (e.g., template RNA) that optionally includes (e.g., from 5' to 3') a sequence that binds to the genetically modified polypeptide (e.g., that specifically binds to an RT domain), a heterologous sequence of interest, and a PBS sequence. In some embodiments, when the system includes multiple nucleic acids, each nucleic acid includes a conjugation domain. In some embodiments, the conjugation domain allows for association of the nucleic acid molecules, e.g., by hybridization of complementary sequences. For example, in some embodiments, the first RNA comprises a first conjugation domain, the second RNA comprises a second conjugation domain, and the first and second conjugation domains are capable of hybridizing to each other, e.g., under stringent conditions. In some embodiments, stringent hybridization conditions include hybridization at about 65°C in 4x sodium chloride / sodium citrate (SSC), followed by a wash in 1x SSC at about 65°C.
[0290] In some embodiments, the template nucleic acid comprises RNA. In some embodiments, the template nucleic acid comprises DNA (e.g., single-stranded or double-stranded DNA). In some embodiments, the template nucleic acid comprises a hybrid having both ribonucleotide and deoxyribonucleotide residues in the same strand.
[0291] In some embodiments, the template nucleic acid comprises one or more (e.g., two) homology domains that are homologous to the target sequence, hi some embodiments, the homology domains are about 10-20, 20-50, or 50-100 nucleotides in length.
[0292] In some embodiments, the template RNA can include, for example, a gRNA sequence for guiding a transgenic polypeptide to a desired target site. In some embodiments, the template RNA includes (e.g., from 5' to 3'): (i) an optional gRNA spacer that binds to a target site (e.g., the second strand of a site in a target genome), (ii) an optional gRNA scaffold that binds to a polypeptide described herein (e.g., a transgenic polypeptide or a Cas polypeptide), (iii) a heterologous sequence of interest comprising a mutation region (optionally, the heterologous sequence of interest comprises, from 5' to 3', a first region of homology, a mutation region, and a second region of homology), and (iv) a primer binding site (PBS) sequence comprising a 3' target homology domain.
[0293] The template nucleic acid (e.g., template RNA) component of the genome editing systems described herein is typically capable of binding to a genetically modified polypeptide of the system. In some embodiments, the template nucleic acid (e.g., template RNA) has a 3' region capable of binding to a genetically modified polypeptide of the system. The binding region, e.g., the 3' region, can be, e.g., a structured RNA region having at least one, two, or three hairpin loops, capable of binding to a genetically modified polypeptide of the system. The binding region can associate the template nucleic acid (e.g., template RNA) with any of the polypeptide modules. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) can associate with an RNA-binding domain in a polypeptide. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) can associate with the reverse transcription domain of a genetically modified polypeptide (e.g., specifically bind to the RT domain). In some embodiments, the template nucleic acid (e.g., template RNA) can associate with the DNA-binding domain of a polypeptide, e.g., a gRNA can associate with a DNA-binding domain from Cas9. In some embodiments, the binding region may also provide DNA target recognition, e.g., a gRNA hybridizes to a target DNA sequence and binds to a polypeptide, e.g., a Cas9 domain. In some embodiments, a template nucleic acid (e.g., a template RNA) may be associated with multiple components of a polypeptide, e.g., a DNA-binding domain and a reverse transcription domain.
[0294] In some embodiments, the template RNA has a poly-A tail at the 3' end. In some embodiments, the template RNA does not have a poly-A tail at the 3' end.
[0295] In some embodiments, the template RNA can be customized to correct a given mutation in the genomic DNA of a target cell (e.g., ex vivo or in vivo, e.g., within a subject, e.g., within a target tissue or organ). For example, the mutation can be a disease-associated mutation relative to a wild-type sequence. Without intending to be bound by any particular theory, any given target site and editing will have a large number of possible template RNA molecules for use in a genetic engineering system that will result in a range of editing efficiencies and fidelity. To partially reduce this screening burden, an empirical parameter set can help ensure an optimal initial in silico design of the template RNA or a portion thereof. As a non-limiting example, the following design parameters can be utilized for a selected mutation: In some embodiments, the design begins by obtaining about 500 bp (e.g., up to 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, or 700 bp, and optionally at least 20, 30, 40, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, or 650 bp) of flanking sequence on either side of the mutation to serve as the target region. In some embodiments, the template nucleic acid comprises a gRNA. In some embodiments, the gRNA comprises a sequence that binds to the target site (e.g., a CRISPR spacer). In some embodiments, target site-binding sequences (e.g., CRISPR spacers) for use in targeting a template nucleic acid to a target region are selected by considering the use of a particular recombination polypeptide (e.g., comprising an endonuclease domain or writing domain, e.g., a CRISPR / Cas domain) (e.g., in the case of Cas9, a protospacer adjacent motif (PAM) of NGG immediately 3' to the 20-nucleotide gRNA binding region). In some embodiments, CRISPR spacers are first selected by ordering whether the PAM will be disrupted by the recombination system-induced editing. In some embodiments, disruption of the PAM can increase editing efficiency.In some embodiments, the PAM can also be disrupted during genetic recombination by introducing silent mutations (e.g., mutations that do not alter any amino acid residues encoded by the target nucleic acid sequence) into the target site (e.g., as part of or in addition to another modification to the target site in the genomic DNA). In some embodiments, CRISPR spacers are selected by ordering the sequences by proximity to the genomic site corresponding to their desired editing location. In some embodiments, the gRNA comprises a gRNA scaffold. In some embodiments, the gRNA scaffold used is a standard scaffold (e.g., for Cas9, 5'-GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC-3'). (SEQ ID NO: 4595) ) or may contain one or more nucleotide substitutions. In some embodiments, the heterologous sequence of interest has at least 90% identity, e.g., at least 90%, 95%, 98%, 99%, or 100% identity, or contains no more than 1, 2, 3, 4, or 5 positions (e.g., immediately 3' to the first strand nick, or up to 1, 2, 3, 4, or 5 nucleotides 3' to the first strand nick) that are non-identical to the target site 3' to the first strand nick, excluding any insertions, substitutions, or deletions that may be written into the target site by genetic recombination. In some embodiments, the 3' target homology domain has at least 90% identity, e.g., at least 90%, 95%, 98%, 99% or 100% identity, or includes no more than 1, 2, 3, 4 or 5 positions of non-identity to the target site 5' to the first strand nick (e.g., immediately 5' to the first strand nick, or up to 1, 2, 3, 4 or 5 nucleotides 3' to the first strand nick).
[0296] In some embodiments, the template nucleic acid is a template RNA. In some embodiments, the template RNA comprises one or more modified nucleotides. For example, in some embodiments, the template RNA comprises one or more deoxyribonucleotides. In some embodiments, regions of the template RNA are substituted with DNA nucleotides, e.g., to enhance the stability of the molecule. For example, the 3' end of the template may comprise DNA nucleotides, while the remainder of the template comprises RNA nucleotides that can be reverse transcribed. For example, in some embodiments, the heterologous sequence of interest is composed primarily or entirely of RNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% RNA nucleotides). In some embodiments, the PBS sequence is composed primarily or entirely of DNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% DNA nucleotides). In other embodiments, the heterologous sequence of interest for genome writing may comprise DNA nucleotides. In some embodiments, the DNA nucleotides in the template are replicated into the genome by a domain capable of DNA-dependent DNA polymerase activity. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in the polypeptide. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that is also capable of DNA-dependent DNA polymerization, e.g., second strand synthesis. In some embodiments, the template molecule is composed exclusively of DNA nucleotides. In some embodiments, the template nucleic acid comprises a hybrid having both ribonucleotide and deoxyribonucleotide residues in the same strand.
[0297] In some embodiments, the systems described herein include two nucleic acids that together comprise the sequence of a template RNA described herein. In some embodiments, the two nucleic acids are non-covalently associated with each other, e.g., directly associated with each other (e.g., by base pairing) or indirectly associated as part of a complex that includes one or more additional molecules.
[0298] The template RNAs described herein can include, from 5' to 3': (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous sequence of interest; and (4) a primer binding site (PBS) sequence. Each of these components is described in more detail herein.
[0299] gRNA spacer and gRNA scaffold The template RNAs described herein can include a gRNA spacer that directs the recombination system to the target nucleic acid, and a gRNA scaffold that facilitates association of the template RNA with the Cas domain of the recombination polypeptide. The systems described herein can also include a gRNA that is not part of the template nucleic acid. For example, a gRNA that includes a gRNA spacer and a gRNA scaffold but does not include a heterologous sequence of interest or a PBS sequence can be used to induce second strand nicking, e.g., as described herein in the section entitled "Second Strand Nicking."
[0300] In some embodiments, gRNAs are short synthetic RNAs composed of a scaffold sequence involved in the binding of CRISPR-associated proteins and a user-defined target sequence of approximately 20 nucleotides for a genomic target. The structure of a complete gRNA was described in Nishimasu et al., Cell 156, pp. 935-949 (2014). A gRNA (also referred to as sgRNA for single guide RNA) consists of sequences derived from a crRNA and a tracrRNA connected by an artificial tetraloop. The crRNA sequence can be divided into a guide (20 nt) and a repeat (12 nt) region, while the tracrRNA sequence can be divided into an anti-repeat (14 nt) and three tracrRNA stem loops (Nishimasu et al., Cell 156, pp. 935-949 (2014)). In practice, guide RNA sequences are generally designed to be 17 to 24 nucleotides (e.g., 19, 20, or 21 nucleotides) long and complementary to a target nucleic acid sequence. Custom gRNA generators and algorithms are commercially available for use in designing effective guide RNAs. In some embodiments, the gRNA comprises two RNA components from the native CRISPR system, such as a crRNA and a tracrRNA. As is well known in the art, the gRNA can also comprise a chimeric single RNA (sgRNA) that contains sequences from both the tracrRNA (to bind to the nuclease) and at least one crRNA (to guide the nuclease to the targeted sequence for editing / binding). Chemically modified sgRNAs have also proven effective for use with CRISPR-associated proteins; see, e.g., Hendel et al. (2015) Nature Biotechnol., 985-991. In some embodiments, the gRNA spacer comprises a nucleic acid sequence complementary to a DNA sequence associated with the target gene.
[0301] In some embodiments, a template nucleic acid, e.g., a region of a template RNA, comprising a gRNA adopts an underwound ribbon-like structure of the gRNA bound to a target DNA (e.g., as described in Mulepati et al., Science 19 Sep 2014: Vol. 345, Issue 6203, pp. 1479-1484). Without intending to be bound by any particular theory, it is believed that this non-canonical structure is facilitated by rotations every six nucleotides from the RNA-DNA hybrid. Thus, in some embodiments, a template nucleic acid, e.g., a region of a template RNA, comprising a gRNA, can tolerate increased mismatches with the target site at certain intervals, e.g., every sixth base. In some embodiments, a template nucleic acid, e.g., a region of a template RNA, comprising a gRNA that contains homology to a target site can have wobble positions at regular intervals, e.g., every sixth base, that do not require base pairing with the target site.
[0302] In some embodiments, the template nucleic acid (e.g., template RNA) comprises a gRNA spacer sequence of a length appropriate for the Cas9 domain of the recombinant polypeptide (Table 3), e.g., having at least 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 bases at the 5' end that are at least 80%, 85%, 90%, 95%, 99%, or 100% homologous to the target site.
[0303] In some embodiments, Cas9 derivatives with enhanced activity can be used in recombinant polypeptides. In some embodiments, Cas9 derivatives can include mutations that improve the activity of the HNH endonuclease domain, such as SpyCas9 R221K, N394K, or mutations that improve R-loop formation, such as SpyCas9 L1245V, or combinations of such mutations, such as SpyCas9 R221K / N394K, SpyCas9 N394K / L1245V, SpyCas9 R221K / L1245V, or SpyCas9 R221K / N394K / L1245V (see, e.g., Spencer and Zhang Sci Rep 7:16836 (2017)). The Cas9 derivatives and mutations contained therein are incorporated herein by reference). In some embodiments, Cas9 derivatives can include one or more types of mutations described herein, such as PAM-modifying mutations, protein-stabilizing mutations, activity-enhancing mutations, and / or mutations that partially or completely inactivate one or two endonuclease domains compared to the parent enzyme (e.g., one or more mutations that abolish endonuclease activity on one or both strands of target DNA, e.g., a nickase or catalytically inactive enzyme). In some embodiments, the Cas9 enzymes used in the systems described herein can include mutations that confer nickase activity to the enzyme (e.g., SpyCas9 N863A or H840A) in addition to mutations that improve catalytic efficiency (e.g., SpyCas9 R221K, N394K, and / or L1245V). In some embodiments, the Cas9 enzymes used in the systems described herein are SpyCas9 enzymes or derivatives that further include the N863A mutation, which confers nickase activity, in addition to the R221K and N394K mutations, which improve catalytic efficiency.
[0304] Table 8 defines the components for designing gRNAs and / or template RNAs and provides parameters for applying the Cas variants listed in Table 3 for genetic engineering. The cleavage site indicates the requirement for a validated or predicted protospacer adjacent motif (PAM), the location of the validated or predicted cleavage site (relative to the most upstream base of the PAM site). A gRNA for a given enzyme can be constructed by concatenating the crRNA, tetraloop, and tracrRNA sequences and adding a 5' spacer within the minimum and maximum lengths of the spacer that matches the protospacer at the target site. Furthermore, the predicted location of the ssDNA nick at the target is important for designing a PBS sequence in the template RNA that can anneal to the sequence immediately 5' of the nick to initiate target-primed reverse transcription. In some embodiments, a gRNA scaffold described herein comprises a nucleic acid sequence comprising, in the 5' to 3' direction, a crRNA of Table 8, a tetraloop from the same row of Table 8, and a tracrRNA from the same row of Table 8, or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gRNA or template RNA comprising the scaffold further comprises a gRNA spacer having a length within the spacer (minimum) and spacer (maximum) set forth in the same row of Table 8. In some embodiments, a gRNA or template RNA having a sequence according to Table 8 is included in a system further comprising a transgenic polypeptide, wherein the transgenic polypeptide comprises a Cas domain set forth in the same row of Table 8.
[0305] [Table 8-1]
[0306] [Table 8-2]
[0307] [Table 8-3]
[0308] [Table 8-4]
[0309] [Table 8-5]
[0310] It is further understood that terminal Us and Ts can be optionally added or removed from the tracrRNA sequence, and, when provided as RNA, can be modified or unmodified. While not intending to be bound by specific examples, alternative gRNA scaffold sequence forms to those exemplified in Table 8, e.g., alternative gRNA scaffold sequences with nucleotide additions, substitutions, or deletions, e.g., sequences with added or removed stem-loop structures, can also function with different Cas9 enzymes or derivatives thereof exemplified in Table 4. It is contemplated herein that gRNA scaffold sequences represent components of genetic engineering systems that can similarly be optimized for a given system, Cas-RT fusion polypeptide, instruction, target mutation, template RNA, or delivery vehicle.
[0311]
[0013] Where an RNA sequence (e.g., a template RNA sequence) is described herein as comprising a particular sequence (e.g., a sequence in Table 8 or a portion thereof) that includes thymine (T), it is of course understood that the RNA sequence can (and often does) include uracil (U) in place of T. For example, the RNA sequence can include U at every position shown as T in the sequences in Table 8. More specifically, the present disclosure provides RNA sequences according to all gRNA scaffold sequences in Table 8, wherein the RNA sequence has a U in place of each T in the sequences in Table 8.
[0312] Heterologous target sequence The template RNA described herein can include a heterologous sequence of interest that can be used as a template for reverse transcription by a recombinant polypeptide to write a desired sequence into a target nucleic acid. In some embodiments, the heterologous sequence of interest includes, from 5' to 3', a post-edited homology region, a mutation region, and a pre-edited homology region. Without intending to be bound by any particular theory, the RT performing the reverse transcription on the template RNA first reverse-transcribes the pre-edited homology region, then the mutation region, and then the post-edited homology region, thereby generating a DNA strand containing the desired mutation with homology regions on either side.
[0313] In some embodiments, the heterologous sequence of interest is at least 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 139, 142, 143, 2, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500 or 1,000 nucleotides (nt), or at least 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 kilobases in length. In some embodiments, the heterologous sequence of interest is 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78 , 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500, 1,000 or 2000 nucleotides (nt), or 20, 15, 10, 9, 8, 7, 6, 5, 4 or 3 kilobases in length.In some embodiments, the heterologous sequence of interest may have a length of 30-1000, 40-1000, 50-1000, 60-1000, 70-1000, 74-1000, 75-1000, 76-1000, 77-1000, 78-1000, 79-1000, 80-1000, 85-1000, 90-1000, 100-1000, 120-1000, 140-1000, 160-1000, 180-1000, 200-1000, 500-1000, 30-500, 40-500, 50-500, 60-500, 7 0~500, 74~500, 75~500, 76~500, 77~500, 78~500, 79~500, 80~500, 85~500, 90~500, 100~500, 120~500, 140~500, 160~500, 180~500, 200~500, 30~200, 40~200, 50~200, 60~200, 70~200, 74~200, 75~200, 76~200, 77~200, 78~200, 79~200, 80~200, 85~200, 90~200, 100~200, 120~ 200, 140-200, 160-200, 180-200, 30-100, 40-100, 50-100, 60-100, 70-100, 74-100, 75-100, 76-100, 77-100, 78-100, 79-100, 80-100, 85-100, or 90-100 nucleotides (nt) in length, or 1-20, 1-15, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, 2-20, 2-15, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5 , 2-4, 2-3, 3-20, 3-15, 3-10, 3-9, 3-8, 3-7, 3-6, 3-5, 3-4, 4-20, 4-15, 4-10, 4-9, 4-8, 4-7, 4-6, 4-5, 5-20, 5-15, 5-10, 5-9, 5-8, 5-7, 5-6, 6-20, 6-15, 6-10, 6-9, 6-8, 6-7, 7-20, 7-15, 7-10, 7-9, 7-8, 8-20, 8-15, 8-10, 8-9, 9-20, 9-15, 9-10, 10-15, 10-20 or 15-20 kilobases.In some embodiments, the heterologous target sequence is 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, or 10-20 nt in length, e.g., 10-80, 10-50, or 10-20 nt in length, e.g., about 10-20 nt in length. In some embodiments, the heterologous target sequence is 8-30, 9-25, 10-20, 11-16, or 12-15 nucleotides in length, e.g., 11-16 nt in length. Without intending to be bound by theory, in some embodiments, a larger insert size of edits, a larger region (e.g., distance between the first edit / replacement and the second edit / replacement in the target region), and / or a greater number of desired edits (e.g., mismatches of the heterologous target sequence relative to the target genome), may result in a longer optimal heterologous target sequence.
[0314] In certain embodiments, template nucleic acids include customized RNA sequence templates that can be identified, designed, modified, and constructed to contain sequences that modify or specify host genome function, such as by: introducing heterologous coding regions into a genome; influencing or causing structural / alternative splicing of exons, e.g., resulting in exon skipping of one or more exons; causing disruption of an endogenous gene, e.g., generating a gene knockout; causing transcriptional activation of an endogenous gene; causing epigenetic modulation of endogenous DNA; causing upregulation of one or more operably linked genes, e.g., resulting in gene activation or overexpression; or causing downregulation of one or more operably linked genes, e.g., generating a gene knockdown. In certain embodiments, customized RNA sequence templates can be modified to include sequences encoding exons and / or transgenes, with the design of binding sites for transcription factors activators, repressors, enhancers, etc., and combinations thereof. In some embodiments, the customized template may be modified to encode a nucleic acid or peptide tag expressed in an endogenous RNA transcript or endogenous protein operably linked to the target site. In other embodiments, the coding sequence may be further customized with a splice donor site, a splice acceptor site, or a polyA tail.
[0315] The template nucleic acid (e.g., template RNA) of the system typically includes a sequence of interest (e.g., a heterologous sequence of interest) for writing a desired sequence into target DNA. The sequence of interest can be a coding or non-coding sequence. The template nucleic acid (e.g., template RNA) can be designed to create an insertion, mutation, or deletion at a target DNA locus. In some embodiments, the template nucleic acid (e.g., template RNA) can be designed to cause an insertion in the target DNA. For example, the template nucleic acid (e.g., template RNA) can include a heterologous sequence, and reverse transcription will result in the insertion of the heterologous sequence into the target DNA. In other embodiments, the RNA template can be designed to introduce a deletion into the target DNA. For example, the template nucleic acid (e.g., template RNA) can match the target DNA upstream and downstream of the desired deletion, and reverse transcription will result in the replication of the upstream and downstream sequences from the template nucleic acid (e.g., template RNA) without the intervening sequence, e.g., causing a deletion of the intervening sequence. In other embodiments, a template nucleic acid (e.g., a template RNA) can be designed to introduce edits into a target DNA. For example, the template RNA can match the target DNA sequence with the exception of one or more nucleotides, and reverse transcription results in the replication of these edits into the target DNA, resulting in, for example, a mutation, such as a transition or transversion mutation.
[0316] In some embodiments, writing a sequence of interest into a target site results in nucleotide substitutions, e.g., where the full length of the sequence of interest corresponds to the matching length of the target site with one or more mismatched bases. In some embodiments, heterologous sequences of interest can be designed such that combinations of sequence variations can be present, e.g., simultaneous addition and deletion, addition and substitution, or deletion and substitution.
[0317] In some embodiments, the heterologous sequence of interest may contain an open reading frame or a fragment of an open reading frame. In some embodiments, the heterologous sequence of interest has a Kozak sequence. In some embodiments, the heterologous sequence of interest has an internal ribosome entry site. In some embodiments, the heterologous sequence of interest has a self-cleaving peptide such as a T2A or P2A site. In some embodiments, the heterologous sequence of interest has a start codon. In some embodiments, the template RNA has a splice acceptor site. In some embodiments, the template RNA has a splice donor site. Exemplary splice acceptor and splice donor sites are described in WO2016044416, which is incorporated by reference herein in its entirety. Exemplary splice acceptor site sequences are known to those of skill in the art. In some embodiments, the template RNA has a microRNA binding site downstream of the stop codon. In some embodiments, the template RNA has a polyA tail downstream of the stop codon of the open reading frame. In some embodiments, the template RNA comprises one or more exons. In some embodiments, the template RNA comprises one or more introns. In some embodiments, the template RNA comprises a eukaryotic transcription terminator. In some embodiments, the template RNA comprises an enhanced translation element or a translation-enhancing element. In some embodiments, the RNA comprises a human T-cell leukemia virus (HTLV-1) R region. In some embodiments, the RNA comprises a post-transcriptional regulatory sequence that enhances nuclear export, such as that of Hepatitis B Virus (HPRE) or Woodchuck Hepatitis Virus (WPRE).
[0318] In some embodiments, the heterologous sequence of interest may contain a non-coding sequence. For example, the template nucleic acid (e.g., template RNA) may include a regulatory element, such as a promoter or enhancer sequence or an miRNA binding site. In some embodiments, integration of the sequence of interest at the target site will result in up-regulation of the endogenous gene. In some embodiments, integration of the sequence of interest at the target site will result in down-regulation of the endogenous gene. In some embodiments, the template nucleic acid (e.g., template RNA) includes a tissue-specific promoter or enhancer, each of which may be unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter includes a TATA element. In some embodiments, the promoter includes a B recognition element. In some embodiments, the promoter has one or more binding sites for a transcription factor.
[0319] In some embodiments, the template nucleic acid (e.g., template RNA) comprises a site for coordinating epigenetic modifications. In some embodiments, the template nucleic acid (e.g., template RNA) comprises a chromatin insulator. For example, the template nucleic acid (e.g., template RNA) comprises a CTCF site or a site targeted for DNA methylation.
[0320] In some embodiments, the template nucleic acid (e.g., template RNA) comprises a gene expression unit comprised of at least one regulatory region operably linked to an effector sequence, which can be a sequence that is transcribed into RNA (e.g., a coding sequence or a non-coding sequence, e.g., a sequence encoding a microRNA).
[0321] In some embodiments, the heterologous sequence of interest of the template nucleic acid (e.g., template RNA) is inserted into the target genome into an endogenous intron. In some embodiments, the heterologous sequence of interest of the template nucleic acid (e.g., template RNA) is inserted into the target genome, thereby acting as a new exon. In some embodiments, the insertion of the heterologous sequence of interest into the target genome results in the replacement of a native exon or the skipping of a native exon.
[0322] In some embodiments, the heterologous sequence of interest of the template nucleic acid (e.g., template RNA) is inserted into the target genome in a genomic safe harbor site, such as the AAVS1, CCR5, ROSA26, or albumin locus. In some embodiments, genetic recombination is used to integrate the CAR into the T cell receptor alpha constant (TRAC) locus (Eyquem et al. Nature 543, 113-117 (2017)). In some embodiments, a genetic recombination system is used to integrate the CAR into the T cell receptor beta constant (TRBC) locus. Many other safe harbors have been identified by computational methods (Pellenz et al. Hum Gen Ther 30, 814-828 (2019)) and can be used for genetic recombination system-mediated integration. In some embodiments, the heterologous sequence of interest of the template nucleic acid (e.g., template RNA) is added to the genome within an intergenic or intragenic region. In some embodiments, the heterologous sequence of interest of the template nucleic acid (e.g., template RNA) is added to the genome 5' or 3' within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of the endogenous active gene. In some embodiments, the heterologous sequence of interest of the template nucleic acid (e.g., template RNA) is added to the genome 5' or 3' within 0.1 kb, 0.25 kb, 0.5 kb, 0.75 kb, 1 kb, 2 kb, 3 kb, 4 kb, 5 kb, 7.5 kb, 10 kb, 15 kb, 20 kb, 25 kb, 50, 75 kb, or 100 kb of the endogenous promoter or enhancer. In some embodiments, the heterologous target sequence of the template nucleic acid (e.g., template RNA) can be, for example, 50 to 50,000 base pairs (e.g., between 50 and 40,000 bp, between 500 and 30,000 bp, between 500 and 20,000 bp, between 100 and 15,000 bp, between 500 and 10,000 bp, between 50 and 10,000 bp, or between 50 and 5,000 bp).
[0323] A template nucleic acid (e.g., template RNA) can be designed to create an insertion, mutation, or deletion at a target DNA locus. In some embodiments, a template nucleic acid (e.g., template RNA) can be designed to cause an insertion into the target DNA. For example, the template nucleic acid (e.g., template RNA) can include a heterologous sequence of interest, and reverse transcription will result in the insertion of the heterologous sequence of interest into the target DNA. In other embodiments, an RNA template can be designed to write a deletion into the target DNA. For example, a template nucleic acid (e.g., template RNA) can match the target DNA upstream and downstream of the desired deletion, and reverse transcription will result in replication of the upstream and downstream sequences from the template nucleic acid (e.g., template RNA) without the intervening sequence, e.g., causing a deletion of the intervening sequence. In other embodiments, a template nucleic acid (e.g., template RNA) can be designed to write an edit into the target DNA. For example, the template RNA may match the target DNA sequence with the exception of one or more nucleotides, and reverse transcription will result in the replication of these edits into the target DNA, resulting in, for example, a mutation, such as a transition or transversion mutation.
[0324] In some embodiments, the preedited homology domain comprises a nucleic acid sequence that has 100% sequence identity to a nucleic acid sequence contained in the target nucleic acid molecule.
[0325] In some embodiments, the post-editing homology domain comprises a nucleic acid sequence that has 100% sequence identity to a nucleic acid sequence contained in the target nucleic acid molecule.
[0326] PBS sequence In some embodiments, the template nucleic acid (e.g., template RNA) comprises a PBS sequence. In some embodiments, the PBS sequence is located 3' to the heterologous sequence of interest and is complementary to a sequence adjacent to the site to be modified by the system described herein or contains no more than 1, 2, 3, 4, or 5 mismatches to a sequence complementary to a sequence adjacent to the site to be modified by the system / recombinant polypeptide. In some embodiments, the PBS sequence binds within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of a nick site in the target nucleic acid molecule. In some embodiments, binding of the PBS sequence to the target nucleic acid molecule allows for the initiation of target-primed reverse transcription (TPRT), for example, by the 3' homology domain, which acts as a primer for TPRT. In some embodiments, the PBS sequence is 3 to 5, 5 to 10, 10 to 30, 10 to 25, 10 to 20, 10 to 19, 10 to 18, 10 to 17, 10 to 16, 10 to 15, 10 to 14, 10 to 13, 10 to 12, 10 to 11, 11 to 30, 11 to 25, 11 to 20, 11 to 19, 11 to 18, 11 to 17, 11-16, 11-15, 11-14, 11-13, 11-12, 12-30, 12-25, 12-20, 12-19, 12-18, 12-17, 12-16, 12-15, 12-14, 12-13, 13-30, 13-25, 13-20, 13-19, 13-18, 13-17, 13-16, 1 3-15, 13-14, 14-30, 14-25, 14-20, 14-19, 14-18, 14-17, 14-16, 14-15, 15-30, 15-25, 15-20, 15-19, 15-18, 15-17, 15-16, 16-30, 16-25, 16-20, 16-19, 16-18, 16-17 , 17-30, 17-25, 17-20, 17-19, 17-18, 18-30, 18-25, 18-20, 18-19, 19-30, 19-25, 19-20, 20-30, 20-25, or 25-30 nucleotides in length, e.g., 10-17, 12-16, or 12-14 nucleotides in length. In some embodiments, the PBS sequence is 5-20, 8-16, 8-14, 8-13, 9-13, 9-12, or 10-12 nucleotides in length, e.g., 9-12 nucleotides in length.
[0327] The template nucleic acid (e.g., template RNA) may have some homology to the target DNA. In some embodiments, the template nucleic acid (e.g., template RNA) is positioned such that the PBS sequence domain can function as an annealing region to the target DNA, thereby priming the target DNA for reverse transcription of the template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid (e.g., template RNA) has at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, 200, or more bases of exact homology to the target DNA at the 3' end of the RNA. In some embodiments, the template nucleic acid (e.g., template RNA) has at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% homology to the target DNA over at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, 200 or more bases, e.g., at the 5' end of the template nucleic acid (e.g., template RNA).
[0328] Inducible gRNA In some embodiments, a gRNA described herein (e.g., a gRNA that is part of a template RNA or a gRNA used for second-strand nicking) has inducible activity. Inducible activity can be achieved by a template nucleic acid, e.g., a template RNA, that further includes a blocking domain (in addition to the gRNA), where the sequence of some or all of the blocking domain is at least partially complementary to some or all of the gRNA. Thus, the blocking domain is hybridizable or substantially hybridizable to some or all of the gRNA. In some embodiments, the blocking domain and inducibly active gRNA are disposed on a template nucleic acid, e.g., a template RNA, such that the gRNA can adopt a first conformation in which the blocking domain is hybridized or substantially hybridized to the gRNA and a second conformation in which the blocking domain is not hybridized or substantially hybridized to the gRNA. In some embodiments, in the first conformation, the gRNA is unable to bind to the engineered polypeptide (e.g., a template nucleic acid binding domain, a DNA binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)) or binds with a substantially reduced affinity compared to an otherwise similar template RNA lacking a blocking domain. In some embodiments, in the second conformation, the gRNA is able to bind to the engineered polypeptide (e.g., a template nucleic acid binding domain, a DNA binding domain, or an endonuclease domain (e.g., a CRISPR / Cas protein)). In some embodiments, whether the gRNA is in the first or second conformation can affect whether the DNA-binding or endonuclease activity of the engineered polypeptide (e.g., of the CRISPR / Cas protein that the engineered polypeptide comprises) is active.
[0329] In some embodiments, the gRNA associated with the second nick has inducible activity. In some embodiments, the gRNA associated with the second nick is induced after the template is reverse transcribed. In some embodiments, hybridization of the gRNA to the blocking domain can be disrupted using an aperture molecule. In some embodiments, the aperture molecule comprises an agent that binds to the gRNA or part or all of the blocking domain and inhibits hybridization of the gRNA to the blocking domain. In some embodiments, the aperture molecule comprises, for example, a nucleic acid comprising a sequence that is partially or fully complementary to the gRNA, the blocking domain, or both. By selecting or designing an appropriate aperture molecule, provision of the aperture molecule can promote a conformational change in the gRNA such that it can associate with a CRISPR / Cas protein and provide the relevant function of the CRISPR / Cas protein (e.g., DNA binding and / or endonuclease activity). Without intending to be bound by theory, provision of the aperture molecule at a selected time and / or location can enable spatial and temporal control of the activity of the gRNA, the CRISPR / Cas protein, or a genetic recombination system comprising them. In some embodiments, the aperture molecule is exogenous to the cell comprising the recombinant polypeptide and / or template nucleic acid. In some embodiments, the aperture molecule comprises an endogenous agent (e.g., endogenous to the cell comprising the recombinant polypeptide and / or template nucleic acid comprising a gRNA and a blocking domain). For example, the inducible gRNA, blocking domain, and aperture molecule can be selected such that the aperture molecule is an endogenous agent expressed in the target cell or tissue, e.g., thereby confirming activity of the recombinant system in the target cell or tissue. As a further example, the inducible gRNA, blocking domain, and aperture molecule can be selected such that the aperture molecule is absent or substantially not expressed in one or more non-target cells or tissues, e.g., thereby confirming activity of the recombinant system is absent or substantially absent, or present at a reduced level relative to the target cell or tissue, in one or more non-target cells or tissues.Exemplary blocking domains, aperture molecules, and their uses are described in PCT Publication WO 2020044039A1, which is incorporated herein by reference in its entirety. In some embodiments, a template nucleic acid, e.g., a template RNA, can include one or more sequences or structures for binding by one or more components of a genetically engineered polypeptide, e.g., a reverse transcriptase or an RNA-binding domain, and a gRNA. In some embodiments, the gRNA facilitates interaction with the template nucleic acid-binding domain (e.g., an RNA-binding domain) of the genetically engineered polypeptide. In some embodiments, the gRNA guides the genetically engineered polypeptide to a matching target sequence, e.g., in the genome of a target cell.
[0330] Circular RNA and ribozymes in genetic engineering systems It is contemplated that it may be useful to use circular and / or linear RNA states during formulation, delivery, or transgenic reactions within target cells. Accordingly, in some embodiments of any of the aspects described herein, the transgenic system comprises one or more circular RNAs (circRNAs). In some embodiments of any of the aspects described herein, the transgenic system comprises one or more linear RNAs. In some embodiments, a nucleic acid described herein (e.g., a nucleic acid molecule encoding a template nucleic acid, a transgenic polypeptide, or both) is a circRNA. In some embodiments, the circular RNA molecule encodes a transgenic polypeptide. In some embodiments, a circRNA molecule encoding a transgenic polypeptide is delivered to a host cell. In some embodiments, the circular RNA molecule encodes a recombinase, e.g., as described herein. In some embodiments, a circRNA molecule encoding a recombinase is delivered to a host cell. In some embodiments, a circRNA molecule encoding a transgenic polypeptide is linearized (e.g., within the host cell, e.g., in the nucleus of the host cell) prior to translation.
[0331] Circular RNAs (circRNAs) are found to occur naturally in cells and have diverse functions, including both non-coding and protein-coding roles in human cells. It has been shown that circRNAs can be modified by incorporating self-splicing introns into RNA molecules (or DNA encoding RNA molecules) to circularize the RNA, and that the modified circRNAs can have enhanced protein production and stability (Wesselhoeft et al. Nature Communications 2018). In some embodiments, a recombinant polypeptide is encoded as a circRNA. In certain embodiments, the template nucleic acid is DNA, such as dsDNA or ssDNA. In certain embodiments, the circDNA comprises a template RNA.
[0332] In some embodiments, the circRNA comprises one or more ribozyme sequences. In some embodiments, the ribozyme sequence is activated, e.g., for self-cleavage, e.g., thereby resulting in linearization of the circRNA, e.g., in a host cell. In some embodiments, the ribozyme is activated when the magnesium concentration reaches a level sufficient for cleavage, e.g., in a host cell. In some embodiments, the circRNA is maintained in a low magnesium environment before delivery to a host cell. In some embodiments, the ribozyme is a protein-responsive ribozyme. In some embodiments, the ribozyme is a nucleic acid-responsive ribozyme. In some embodiments, the circRNA comprises a cleavage site. In some embodiments, the circRNA comprises a second cleavage site.
[0333] In some embodiments, the circRNA is linearized in the nucleus of a target cell. In some embodiments, linearization of the circRNA in the nucleus of a cell involves components present in the nucleus of the cell, e.g., to activate a cleavage event. In some embodiments, a ribozyme, e.g., a ribozyme derived from a B2 or ALU element, responsive to a nuclear element, e.g., a protein that interacts with the genome, e.g., an epigenetic modifier, e.g., EZH2, is incorporated into the circRNA, e.g., in a recombinant DNA system. In some embodiments, nuclear localization of the circRNA results in increased autocatalytic activity of the ribozyme and linearization of the circRNA.
[0334] In some embodiments, the ribozyme is heterologous to one or more other components of the genetic engineering system. In some embodiments, the inducible ribozyme (e.g., in a circRNA described herein) is synthetically generated, for example, by using a protein ligand-responsive aptamer design. A system for using the satellite RNA of the tobacco ringspot virus hammerhead ribozyme with an MS2 coat protein aptamer has been described, resulting in activation of ribozyme activity in the presence of MS2 coat protein (Kennedy et al. Nucleic Acids Res 42(19):12306-12321 (2014) (incorporated herein by reference in its entirety). In several embodiments, such a system responds to a protein ligand localized in the cytoplasm or nucleus. In some embodiments, the protein ligand is not MS2. Methods for generating RNA aptamers for target ligands have been described, for example, based on systematic evolution of ligands by exponential enrichment (SELEX) (Tuerk and Gold, Science 249(4968):505-510(1990); Ellington and Szostak, Nature 346(6287):818-822(1990); each of which is incorporated herein by reference), and in some cases assisted by in silico design (Bell et al. PNAS 117(15):8486-8493; each of which is incorporated herein by reference). Thus, in some embodiments, aptamers for target ligands are generated and incorporated into synthetic ribozyme systems, e.g., to induce ribozyme-mediated cleavage and linearization of circRNAs in the presence of, for example, a protein ligand. In some embodiments, linearization of circRNAs is induced in the cytoplasm, for example, using an aptamer that associates with a ligand in the cytoplasm. In some embodiments, linearization of circRNAs is induced in the nucleus, for example, using an aptamer that associates with a ligand in the nucleus. In several embodiments, the ligand in the nucleus comprises an epigenetic modifier or a transcription factor.In some embodiments, the linearization-inducing ligand is present at higher levels in on-target cells than in off-target cells.
[0335] Furthermore, nucleic acid-responsive ribozyme systems may be used to linearize circRNAs. For example, biosensors that detect a defined target nucleic acid molecule to trigger ribozyme activation are described, for example, by Penchovsky (Biotechnology Advances 32(5):1015-1027(2014)), incorporated herein by reference). By these methods, the ribozyme naturally folds in an inactive state and is activated only in the presence of a defined target nucleic acid molecule (e.g., an RNA molecule). In some embodiments, the circRNA of the genetic engineering system comprises a nucleic acid-responsive ribozyme that is activated in the presence of a defined target nucleic acid, such as RNA, for example, mRNA, miRNA, guide RNA, gRNA, sgRNA, ncRNA, lncRNA, tRNA, snRNA, or mtRNA. In some embodiments, the nucleic acid that triggers linearization is present at higher levels in on-target cells than in off-target cells.
[0336] In some embodiments of any of the aspects herein, the genetic engineering system incorporates one or more ribozymes with induced specificity for a target tissue or cell of interest, e.g., a ribozyme that is activated by a ligand or nucleic acid that is present at higher levels in the target tissue or cell of interest. In some embodiments, the genetic engineering system incorporates a ribozyme with induced specificity for a subcellular compartment, e.g., the nucleus, nucleolus, cytoplasm, or mitochondria. In some embodiments, the ribozyme is activated by a ligand or nucleic acid that is present at higher levels in the target subcellular compartment. In some embodiments, the RNA component of the genetic engineering system is provided as a circRNA that is activated, e.g., by linearization. In some embodiments, linearization of a circRNA encoding a genetic engineering polypeptide activates the molecule for translation. In some embodiments, the signal that activates the circRNA component of the genetic engineering system is present at higher levels in on-target cells or tissues, e.g., the system is specifically activated in these cells.
[0337] In some embodiments, the RNA component of the genetic engineering system is provided as a circRNA that is inactivated by linearization. In some embodiments, the circRNA encoding the genetically engineered polypeptide is inactivated by cleavage and degradation. In some embodiments, the circRNA encoding the genetically engineered polypeptide is inactivated by cleavage that separates the translation signal from the coding sequence of the polypeptide. In some embodiments, the signal that inactivates the circRNA component of the genetic engineering system is present at higher levels in off-target cells or tissues, thereby specifically inactivating the system in these cells.
[0338] target nucleic acid site In some embodiments, after genetic recombination, the target site surrounding the edited sequence contains a limited number of insertions or deletions in about 50% or less than 10% of editing events, e.g., as determined by long-read amplicon sequencing of the target site, e.g., as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated herein by reference in its entirety). In some embodiments, the target site does not exhibit multiple contiguous editing events, e.g., head-to-tail or head-to-head duplications, e.g., as determined by long-read amplicon sequencing of the target site, e.g., as described in Karst et al. bioRxiv doi.org / 10.1101 / 645903 (2020) (incorporated herein by reference in its entirety). In some embodiments, the target site contains the integrated sequence corresponding to the template RNA. In some embodiments, the target site does not contain an insertion derived from the endogenous RNA in more than about 1% or 10% of events, e.g., as determined by long-read amplicon sequencing of the target site, e.g., as described in Karst et al. bioRxiv doi.org / 10.1101 / 645903(2020), which is incorporated by reference in its entirety. In some embodiments, the target site contains an integrated sequence that corresponds to the template RNA.
[0339] In certain embodiments of the invention, the host DNA binding site integrated by the genetic recombination system may be located within a gene, within an intron, within an exon, within an ORF, outside the coding region of any gene, within the regulatory region of a gene, or outside the regulatory region of a gene. In other embodiments, the polypeptide may bind to one or more host DNA sequences.
[0340] In some embodiments, a genetic engineering system is used to edit a target locus in multiple alleles. In some embodiments, the genetic engineering system is designed to edit a specific allele. For example, a genetic engineering polypeptide can be directed to a specific sequence present only on one allele, for example, comprising a template RNA that has homology to the target allele, e.g., a gRNA or annealing domain, but not to the second cognate allele. In some embodiments, the genetic engineering system can modify haplotype-specific alleles. In some embodiments, a genetic engineering system that targets a specific allele preferentially targets that allele, for example, with at least a 2-, 4-, 6-, 8-, or 10-fold preference for the target allele.
[0341] Second strand nicking In some embodiments, the genetic engineering systems described herein include a nickase activity (e.g., in a genetically engineered polypeptide) that nicks the first strand and a nickase activity (e.g., in a polypeptide separate from the genetically engineered polypeptide) that nicks the second strand of target DNA. As described herein, without intending to be bound by theory, it is believed that nicking the first strand of target site DNA provides a 3'OH that can be used by the RT domain to reverse transcribe a template RNA sequence, e.g., a heterologous sequence of interest. While not intending to be bound by theory, it is believed that introducing an additional nick into the second strand can bias the cellular DNA repair machinery to adopt a sequence based on the heterologous sequence of interest more frequently than the original genomic sequence. In some embodiments, the additional nick in the second strand is created by the same endonuclease domain (e.g., a nickase domain) as the nick in the first strand. In some embodiments, the same genetically engineered polypeptide serves the functions of both the nick in the first strand and the nick in the second strand. In some embodiments, the engineered polypeptide comprises a CRISPR / Cas domain, and the additional nick in the second strand is directed by an additional nucleic acid comprising, for example, a second gRNA that directs the CRISPR / Cas domain to nick the second strand. In other embodiments, the additional second strand nick is created by an endonuclease domain (e.g., a nickase domain) that is different from the nick in the first strand. In some embodiments, the different endonuclease domain is located in an additional polypeptide separate from the engineered polypeptide (e.g., the system further comprises an additional polypeptide). In some embodiments, the additional polypeptide comprises an endonuclease domain (e.g., a nickase domain) described herein. In some embodiments, the additional polypeptide comprises, for example, a DNA-binding domain described herein.
[0342] It is contemplated herein that the location of the second strand nick relative to the first strand nick can affect one or more of the following: the extent to which a desired recombinant DNA modification is obtained, the extent to which an unwanted double-strand break (DSB) occurs, the extent to which an unwanted insertion occurs, or the extent to which an unwanted deletion occurs. Without intending to be bound by any particular theory, second strand nicking can occur in two general orientations: an inward nick and an outward nick.
[0343] In some embodiments, in the inward nick orientation, the RT domain polymerizes (e.g., with a template RNA (e.g., a heterologous target sequence)) away from the second strand nick. In some embodiments, in the inward nick orientation, the position of the nick relative to the first strand and the position of the nick relative to the second strand are positioned between the first PAM site and the second PAM site (e.g., in situations where both nicks are created by a polypeptide (e.g., a genetically engineered polypeptide) comprising a CRISPR / Cas domain). In some embodiments, in the inward nick orientation, the position of the nick relative to the first strand and the position of the nick relative to the second strand are between the site where the polypeptide binds to the target DNA and the site where the additional polypeptide binds to the target DNA. In some embodiments, in the inward nick orientation, the position of the nick relative to the first strand is positioned on the same side of the binding site for the polypeptide and the additional polypeptide relative to the position of the nick relative to the first strand. In some embodiments, in the inward nick orientation, the position of the nick relative to the first strand and the position of the nick relative to the second strand are positioned between the PAM site and a site distant from the target site.
[0344] An example of a transgenic system that provides inward nick directionality includes a transgenic polypeptide comprising a CRISPR / Cas domain, a template RNA comprising a gRNA that directs nicking of target site DNA on the first strand, and a further nucleic acid comprising a further gRNA that directs nicking at a site spaced apart from the location of the first nick, where the location of the first nick and the location of the second nick are between the PAM sites where the two gRNAs direct the transgenic polypeptide. As a further example, another transgenic system that provides inward nick directionality includes a transgenic polypeptide comprising a zinc finger molecule and a first nickase domain, a further polypeptide comprising a CRISPR / Cas domain, and a further nucleic acid comprising a gRNA that directs the further polypeptide to nick at a site spaced apart from the target site DNA on the second strand, where the zinc finger molecule binds to the target DNA in a manner that directs the first nickase domain to nick the first strand of the target site; the location of the first nick and the location of the second nick are between the PAM site and the site where the zinc finger molecule binds. As a further example, another recombinant system providing inward nick directionality comprises an recombinant polypeptide comprising a zinc finger molecule and a first nickase domain, a TAL effector molecule and an additional polypeptide comprising a second nickase domain, wherein the zinc finger molecule binds to target DNA in a manner that directs the first nickase domain to nick a first strand of the target site; wherein the TAL effector molecule binds to a site spaced from the target site in a manner that directs the additional polypeptide to nick a second strand, and the location of the first nick and the location of the second nick are between the site where the TAL effector molecule binds and the site where the zinc finger molecule binds.
[0345] In some embodiments, in the outward nick orientation, the RT domain polymerizes toward the second strand nick (e.g., using a template RNA (e.g., a heterologous sequence of interest)). In some embodiments, in the inward nick orientation, when both the first and second nicks are created by a polypeptide (e.g., a genetically modified polypeptide) comprising a CRISPR / Cas domain, the first PAM site and the second PAM site are positioned between the position of the nick for the first strand and the position of the nick for the second strand. In some embodiments, in the inward nick orientation, the polypeptide (e.g., a genetically modified polypeptide) and the additional polypeptide bind to a site on the target DNA between the position of the nick for the first strand and the position of the nick for the second strand. In some embodiments, in the inward nick orientation, the position of the nick for the second strand is positioned on the opposite side of the binding site for the polypeptide and the additional polypeptide from the position of the nick for the first strand. In some embodiments, in the inward orientation, the PAM site and the site distant from the target site are positioned between the position of the nick for the first strand and the position of the nick for the second strand.
[0346] An example of a recombinant system that provides outward nick directionality includes an recombinant polypeptide that includes a CRISPR / Cas domain, a template RNA that includes a gRNA that directs nicking of target site DNA on the first strand, and an additional nucleic acid that includes an additional gRNA that directs nicking at a site spaced apart from the location of the first nick, where the location of the first nick and the location of the second nick are outside the PAM site of the sites to which the two gRNAs direct the recombinant polypeptide (i.e., the PAM site is between the location of the first nick and the location of the second nick). As a further example, another recombinant system that provides outward nick directionality includes an recombinant polypeptide comprising a zinc finger molecule and a first nickase domain, an additional polypeptide comprising a CRISPR / Cas domain, and an additional nucleic acid comprising a gRNA that directs the additional polypeptide to nick a site on a second strand away from the target site DNA, wherein the zinc finger molecule binds to the target DNA in a manner that directs the first nickase domain to nick the first strand of the target site; and the location of the first nick and the location of the second nick are outside the PAM site and the site where the zinc finger molecule binds (i.e., the PAM site and the site where the zinc finger molecule binds are between the location of the first nick and the location of the second nick). As a further example, another recombinant system providing outward nick directionality includes an recombinant polypeptide comprising a zinc finger molecule and a first nickase domain, a TAL effector molecule, and an additional polypeptide comprising a second nickase domain, wherein the zinc finger molecule binds to target DNA in a manner that directs the first nickase domain to nick a first strand of the target site; wherein the TAL effector molecule binds to a site spaced apart from the target site in a manner that directs the additional polypeptide to nick a second strand, and the positions of the first nick and the second nick are outside the sites where the TAL effector molecule binds and the zinc finger molecule binds (i.e., the sites where the TAL effector molecule binds and the zinc finger molecule bind are between the positions of the first nick and the second nick).
[0347] Without intending to be bound by any particular theory, in genetic recombination systems in which second-strand nicks are provided, outward nick orientation is believed to be preferred in some embodiments. As described herein, inward nicks can generate a greater number of double-strand breaks (DSBs) than outward nick orientation. DSBs can be recognized by DSB repair pathways in the cell's nucleus, resulting in unwanted insertions and deletions. Outward nick orientation can provide a reduced risk of DSB formation and a corresponding lower number of unwanted insertions and deletions. In some embodiments, the unwanted insertions and deletions are insertions and deletions not encoded by the heterologous sequence of interest, e.g., insertions or deletions generated by double-strand break repair pathways unrelated to the modification encoded by the heterologous sequence of interest. In some embodiments, the desired genetic modification comprises a modification (e.g., a substitution, insertion, or deletion) to the target DNA encoded by the heterologous sequence of interest (e.g., achieved by genetic recombination that writes the heterologous sequence of interest into the target site). In some embodiments, the first-strand nicks and second-strand nicks are outward-oriented.
[0348] Furthermore, the distance between the first-strand nick and the second-strand nick can affect one or more of the following: the extent to which the desired genetic recombination system DNA modification is achieved, the extent to which unwanted double-strand breaks (DSBs) occur, the extent to which unwanted insertions occur, or the extent to which unwanted deletions occur. Without intending to be bound by any particular theory, it is believed that second-strand nicks are advantageous because they bias DNA repair toward the integration of the heterologous sequence of interest into the target DNA, increasing as the distance between the first-strand nick and the second-strand nick decreases. However, it is believed that the risk of DSB formation also increases as the distance between the first-strand nick and the second-strand nick decreases. Accordingly, it is believed that the number of unwanted insertions and / or deletions may increase as the distance between the first-strand nick and the second-strand nick decreases. In some embodiments, the distance between the first-strand nick and the second-strand nick is selected to balance the benefit of biasing DNA repair toward the integration of the heterologous sequence of interest into the target DNA with the risk of DSB formation and unwanted deletions and / or insertions. In some embodiments, a system where the first strand nick and the second strand nick are separated by at least a threshold distance has an increased level of desired genetic recombination system modification results, a decreased level of unwanted deletions, and / or a decreased level of unwanted insertions, compared to an otherwise similar inward nick directional system where the first nick and the second nick are separated by less than a threshold distance. In some embodiments, the threshold distance is given below.
[0349] In some embodiments, the first nick and the second nick are at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides apart, In some embodiments, the first nick and the second nick are no more than 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, or 250 nucleotides apart. In some embodiments, the first and second nicks are 20 to 200, 30 to 200, 40 to 200, 50 to 200, 60 to 200, 70 to 200, 80 to 200, 90 to 200, 100 to 200, 110 to 200, 120 to 200, 130 to 200, 140 to 200, 150 to 200, 160 to 200, 170 to 200, 180 to 200, 190 to 200, 20 to 190, 30 to 190, 40 to 190 , 50~190, 60~190, 70~190, 80~190, 90~190, 100~190, 110~190, 120~190, 130~190, 140~190, 150~190, 160~190, 170~190, 180~190, 20~180, 30~180, 40~180, 50~180, 60~180, 70~180, 80~180, 90~180, 100~180, 110~180, 12 0~180, 130~180, 140~180, 150~180, 160~180, 170~180, 20~170, 30~170, 40~170, 50~170, 60~170, 70~170, 80~170, 90~170, 100~170, 110~170, 120~170, 130~170, 140~170, 150~170, 160~170, 20~160, 30~160, 40~160, 50~ 160, 60~160, 70~160, 80~160, 90~160, 100~160, 110~160, 120~160, 130~160, 140~160, 150~160, 20~150, 30~150, 40~150, 50~150, 60~150, 70~150, 80~150, 90~150, 100~150, 110~150, 120~150, 130~150, 140~150, 20~140,30~140, 40~140, 50~140, 60~140, 70~140, 80~140, 90~140, 100~140, 110~140, 120~140, 130~140, 20~130, 30~130, 40~130, 50~130, 60~130, 70~130, 80~130, 90~ 130, 100-130, 110-130, 120-130, 20-120, 30-120, 40-120, 50-120, 60-120, 70-120, 80-120, 90-120, 100-120, 110-120, 20-110, 30-110, 40-110, 50-110, 60-110 , 70~110, 80~110, 90~110, 100~110, 20~100, 30~100, 40~100, 50~100, 60~100, 70~100, 80~100, 90~100, 20~90, 30~90, 40~90, 50~90, 60~90, 70~90, 80~90, 20~80, The first nick and the second nick are separated by 30 to 80, 40 to 80, 50 to 80, 60 to 80, 70 to 80, 20 to 70, 30 to 70, 40 to 70, 50 to 70, 60 to 70, 20 to 60, 30 to 60, 40 to 60, 50 to 60, 20 to 50, 30 to 50, 40 to 50, 20 to 40, 30 to 40, or 20 to 30 nucleotides. In some embodiments, the first nick and the second nick are separated by 40 to 100 nucleotides.
[0350] Without intending to be bound by any particular theory, it is believed that for genetic recombination systems in which second-strand nicks are provided and inward nick directionality is selected, increasing the distance between the first-strand nick and the second-strand nick may be preferable. As described herein, inward nick directionality may generate a greater number of DSBs than outward nick directionality and may result in greater amounts of unwanted insertions and deletions than outward nick directionality, but increasing the distance between nicks may mitigate such increases in DSBs, unwanted deletions, and / or unwanted insertions. In some embodiments, inward nick directionality where the first and second nicks are at least a threshold distance apart has an increased level of desired genetic recombination system modification results, a decreased level of unwanted deletions, and / or a decreased level of unwanted insertions, compared to an otherwise similar inward nick direction system where the first and second nicks are less than a threshold distance apart. In some embodiments, the threshold distance is given below.
[0351] In some embodiments, the first strand nick and the second strand nick are inward-directed. In some embodiments, the first strand nick and the second strand nick are inward-directed and the first strand nick and the second strand nick are at least 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 350, 400, 450, or 500 nucleotides apart, e.g., at least 100 nucleotides apart (and optionally no more than 500, 400, 300, 200, 190, 180, 170, 160, 150, 140, 130, or 120 nucleotides apart). In some embodiments, the first strand nick and the second strand nick are inwardly directed, and the first strand nick and the second strand nick are 100-200, 110-200, 120-200, 130-200, 140-200, 150-200, 160-200, 170-200, 180-200, 190-200, 100-190, 110-190, 120-190, 130-190, 140-190, 150-190, 160-190, 170-190, 180-190, 100-180, 110-180, 120-180, 130-180, 140-180, 150-180 , 160-180, 170-180, 100-170, 110-170, 120-170, 130-170, 140-170, 150-170, 160-170, 100-160, 110-160, 120-160, 130-160, 140-160, 150-160, 100-150, 110-150, 120-150, 130-150, 140-150, 100-140, 110-140, 120-140, 130-140, 100-130, 110-130, 120-130, 100-120, 110-120, or 100-110 nucleotides apart.
[0352] Characteristics of chemically modified nucleic acids and nucleic acid termini Nucleic acids described herein (e.g., template nucleic acids, e.g., template RNA; or nucleic acids encoding recombinant polypeptides (e.g., mRNA); or gRNA) can contain unmodified or modified nucleobases. Naturally occurring RNA is synthesized from four basic ribonucleotides: ATP, CTP, UTP, and GTP, but can contain post-transcriptionally modified nucleotides. Furthermore, approximately 100 different nucleoside modifications have been identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucleic Acids Res 27:196-197). RNA can also contain non-naturally occurring, entirely synthetic nucleotides.
[0353] In some embodiments, the chemical modification is a modification of a polypeptide as described in WO 2016 / 183482, U.S. Patent Application Publication No. 20090286852, International Patent Publication No. WO 2012 / 019168, WO 2012 / 045075, WO 2012 / 135805, WO 2012 / 158736, WO 2013 / 039857, WO 2013 / 039861, WO 2013 / 052523, WO 2013 / 062526, WO 2013 / 052527, WO 2013 / 062528, WO 2013 / 062529 ... International Publication No. 13 / 090648, International Publication No. 2013 / 096709, International Publication No. 2013 / 101690, International Publication No. 2013 / 106496, International Publication No. 2013 / 130161, International Publication No. 2013 / 151669, International Publication No. 2013 / 151736, International Publication No. 2013 / 151672, International Publication No. 2013 / 151664, International Publication No. 2013 / 151665, International Publication No. 2013 / 151 668, WO 2013 / 151671, WO 2013 / 151667, WO 2013 / 151670, WO 2013 / 151666, WO 2013 / 151663, WO 2014 / 028429, WO 2014 / 081507, WO 2014 / 093924, WO 2014 / 093574, WO 2014 / 113089 Brochure, International Publication No. 2014 / 144711, International Publication No. 2014 / 144767, International Publication No. 2014 / 144039, International Publication No. 2014 / 152540, International Publication No. 2014 / 152030, International Publication No. 2014 / 152031, International Publication No. 2014 / 152027, International Publication No. 2014 / 152211, International Publication No. 2014 / 158795, International Publication No. 2014 / 159813,WO 2014 / 164253, WO 2015 / 006747, WO 2015 / 034928, WO 2015 / 034925, WO 2015 / 038892, WO 2015 / 048744, WO 2015 / 051214, WO 2015 / 051173, WO 2015 / 051169, WO 2015 / 058069, WO 2015 / 085318, WO 2015 / 089511, WO 2015 / 105926, WO 2015 /
[0023] The present invention relates to a method for producing a medicament for the treatment of a medicament comprising administering to a subject a medicament therapies, including but not limited to, those provided in WO 2016 / 0164674, WO 2015 / 196130, WO 2015 / 196128, WO 2015 / 196118, WO 2016 / 011226, WO 2016 / 011222, WO 2016 / 011306, WO 2016 / 014846, WO 2016 / 022914, WO 2016 / 036902, WO 2016 / 077125, or WO 2016 / 077123 (each of which is incorporated herein by reference in its entirety). It is understood that incorporation of chemically modified nucleotides into a polynucleotide can result in modifications being incorporated into the nucleobase, the backbone, or both, depending on the position of the modification in the nucleotide. In some embodiments, the backbone modifications are those provided in EP 2813570 (incorporated herein by reference in its entirety). In some embodiments, the modified caps are those provided in U.S. Patent Application Publication No. 20050287539 (incorporated herein by reference in its entirety).
[0354] In some embodiments, the chemically modified nucleic acid (e.g., RNA, e.g., mRNA) comprises one or more of ARCA:anti-reverse cap analog (m27.3'-OGP3G), GP3G (unmethylated cap analog), m7GP3G (monomethylated cap analog), m32.2.7GP3G (trimethylated cap analog), m5CTP (5'-methyl-cytidine triphosphate), m6ATP (N6-methyl-adenosine-5'-triphosphate), s2UTP (2-thio-uridine triphosphate), and Ψ (pseudouridine triphosphate).
[0355] In some embodiments, chemically modified nucleic acids include a 5' cap, such as: a 7-methylguanosine cap (e.g., an O-Me-m7G cap); a hypermethylated cap analog; an NAD+-derived cap analog (e.g., as described in Kiledjian, Trends in Cell Biology 28, 454-464 (2018)); or a modified, e.g., biotinylated cap analog (e.g., as described in Bednarek et al., Phil Trans R Soc B 373, 20180167 (2018)).
[0356] In some embodiments, the chemically modified nucleic acid may comprise a poly-A tail; a 16-nucleotide-long stem-loop structure flanked by five unpaired nucleotides (e.g., as described in Mannironi et al., Nucleic Acid Research 17, 9113-9126 (1989)); a triple helix structure (e.g., as described in Brown et al., PNAS 109, 19202-19207 (2012)); a tRNA, Y RNA, or vault RNA structure (e.g., as described in Labno et al., Biochemica et Biophysica Acta 1863, 3125-3147 (2016); incorporation of one or more deoxyribonucleotide triphosphates (dNTPs), 2'O-methylated NTPs, or phosphorothioate-NTPs; single nucleotide chemical modification (e.g., oxidation of the 3'-terminal ribose to a reactive aldehyde followed by conjugation of an aldehyde-reactive modified nucleotide); or chemical ligation to another nucleic acid molecule.
[0357] In some embodiments, the nucleic acid (e.g., template nucleic acid) may be, for example, dihydrouridine, inosine, 7-methylguanosine, 5-methylcytidine (5mC), 5'ribothymidine phosphate, 2'-O-methylribothymidine, 2'-O-ethylribothymidine, 2'-fluororibothymidine, C-5 propynyl-deoxycytidine (pdC), C-5 propynyl-deoxyuridine (pdU), C-5 propynyl-cytidine (pC), C-5 propynyl-uridine (pU), 5-methylcytidine, 5-methyluridine, 5-methyldeoxycytidine, 5-methyldeoxyuridine methoxy, 2,6-diaminopurine, , 5'-dimethoxytrityl-N4-ethyl-2'-deoxycytidine, C-5 propynyl-f-cytidine (pfC), C-5 propynyl-f-uridine (pfU), 5-methyl f-cytidine, 5-methyl f-uridine, C-5 propynyl-m-cytidine (pmC), C-5 propynyl-f-uridine (pmU), 5-methyl m-cytidine, 5-methyl m-uridine, LNA (locked nucleic acid), MGB (minor groove binder) pseudouridine (Ψ), 1-N-methylpseudouridine (1-Me-Ψ), or 5-methoxyuridine (5-MO-U).
[0358] In some embodiments, the nucleic acid comprises a backbone modification, e.g., a modification to a sugar or phosphate group in the backbone. In some embodiments, the nucleic acid comprises a nucleobase modification.
[0359] In some embodiments, a nucleic acid comprises one or more chemically modified nucleotides in Table 9, one or more chemical backbone modifications in Table 10, and one or more chemically modified caps in Table 11. For example, in some embodiments, a nucleic acid comprises two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of chemical modifications. By way of example, a nucleic acid can comprise two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of modified nucleobases, e.g., as described herein, e.g., in Table 9. Alternatively, or in combination, a nucleic acid can comprise two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of backbone modifications, e.g., as described herein, e.g., in Table 10. Alternatively, or in combination, a nucleic acid can comprise one or more modified caps, e.g., as described herein, e.g., in Table 11. For example, in some embodiments, a nucleic acid comprises one or more types of modified nucleobases and one or more types of backbone modifications; one or more types of modified nucleobases and one or more types of modified caps; one or more types of modified caps and one or more types of backbone modifications; or one or more types of modified nucleobases, one or more types of backbone modifications, and one or more types of modified caps.
[0360] In some embodiments, a nucleic acid comprises one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more) modified nucleobases. In some embodiments, all of the nucleobases of a nucleic acid are modified. In some embodiments, the nucleic acid is modified at one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more) positions in the backbone. In some embodiments, all backbone positions of the nucleic acid are modified.
[0361] [Table 9-1]
[0362] [Table 9-2]
[0363] [Table 10]
[0364] [Table 11]
[0365] The nucleotides comprising the template of the genetic recombination system can be natural or modified bases, or a combination thereof. For example, the template can include pseudouridine, dihydrouridine, inosine, 7-methylguanosine, or other modified bases. In some embodiments, the template can include locked nucleic acid nucleotides. In some embodiments, the modified bases used in the template do not inhibit reverse transcription of the template. In some embodiments, the modified bases used in the template can improve reverse transcription, for example, specificity or fidelity.
[0366] In some embodiments, the RNA component of the system (e.g., the template RNA or the gRNA) comprises one or more nucleotide modifications. In some embodiments, the modification pattern of the gRNA can significantly affect in vivo activity compared to an unmodified or end-modified guide (e.g., as shown in Figure 1D from Finn et al. Cell Rep 22(9):2227-2235 (2018) (incorporated herein by reference in its entirety)). Without intending to be bound by any particular theory, this process may be due, at least in part, to the stabilization of the RNA imparted by the modifications. Non-limiting examples of such modifications may include 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), 2'-fluoro (2'-F), internucleotide phosphorothioate (PS) linkages, GC substitutions, and inverted internucleotide abasic linkages and their equivalents.
[0367] In some embodiments, the template RNA (e.g., the portion thereof that binds to the target site) or the guide RNA comprises a 5'-end region. In some embodiments, the template RNA or the guide RNA does not comprise a 5'-end region. In some embodiments, the 5'-end region comprises a gRNA spacer region, e.g., as described for sgRNAs in Briner AE et al., Molecular Cell 56:333-339 (2014) (incorporated by reference herein in its entirety; e.g., applicable herein to all guide RNAs). In some embodiments, the 5'-end region comprises a 5'-end modification. In some embodiments, the 5'-end region may be associated with a crRNA, trRNA, sgRNA, and / or dgRNA with or without a spacer region. The gRNA spacer region may, in some cases, comprise a guide region, a guide domain, or a target domain.
[0368] In some embodiments, a template RNA (e.g., a portion thereof that binds to a target site) or guide RNA described herein comprises any of the sequences set forth in Table 4 of WO2018107028A1 (incorporated herein by reference in its entirety). In some embodiments, where a sequence represents a guide and / or spacer region, the composition may or may not include this region. In some embodiments, the guide RNA comprises one or more modifications of any of the sequences set forth in Table 4 of WO2018107028A1, e.g., as identified by a SEQ ID NO: therein. In some embodiments, the nucleotides may be the same or different, and / or the modification pattern shown may be identical to or similar to the modification pattern of a guide sequence as set forth in Table 4 of WO2018107028A1. In some embodiments, the modification pattern includes the relative position and identity of the modifications on the gRNA or regions of the gRNA (e.g., the 5' end region, the more downstream stem region, the bulge region, the more upstream stem region, the nexus region, the hairpin 1 region, the hairpin 2 region, the 3' end region). In some embodiments, the modification pattern includes at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the modifications of any one of the sequences set forth in the sequence column of Table 4 of WO2018107028A1, and / or across one or more regions of that sequence. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical to the modification pattern of any one of the sequences shown in the sequence column of Table 4 of WO2018107028A1.In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical across one or more regions of the sequence shown in Table 4 of WO2018107028A1, e.g., in the 5'-terminal region, the more downstream stem region, the bulge region, the more upstream stem region, the nexus region, the hairpin 1 region, the hairpin 2 region, and / or the 3'-terminal region. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical to the modification pattern of the sequence across the 5'-terminal region. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical across the more downstream stem. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical across the bulge. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical across the more upstream stem. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical across the nexus. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical across hairpin 1. In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% identical across hairpin 2.In some embodiments, the modification pattern is at least 50%, 55%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical across the 3' end. In some embodiments, the modification pattern differs from the modification pattern of a sequence in Table 4 of WO2018107028A1, or a region of such a sequence (e.g., the 5' end, the more downstream stem, the bulge, the more upstream stem, the nexus, hairpin 1, hairpin 2, the 3' end), e.g., by 0, 1, 2, 3, 4, 5, 6, or more nucleotides. In some embodiments, the gRNA comprises a modification that differs from the modification of a sequence in Table 4 of WO2018107028A1, e.g., by 0, 1, 2, 3, 4, 5, 6, or more nucleotides. In some embodiments, the gRNA includes a modification that differs from a modification in a region of the sequence in Table 4 of WO2018107028A1 (e.g., the 5' end, the more downstream stem, the bulge, the more upstream stem, the nexus, hairpin 1, hairpin 2, the 3' end), e.g., by 0, 1, 2, 3, 4, 5, 6, or more nucleotides.
[0369] In some embodiments, the template RNA (e.g., the portion thereof that binds to the target site) or the gRNA comprises 2'-O-methyl (2'-O-Me) modified nucleotides. In some embodiments, the gRNA comprises 2'-O-(2-methoxyethyl) (2'-O-moe) modified nucleotides. In some embodiments, the gRNA comprises 2'-fluoro (2'-F) modified nucleotides. In some embodiments, the gRNA comprises phosphorothioate (PS) internucleotide linkages. In some embodiments, the gRNA comprises a 5'-end modification, a 3'-end modification, or a 5'- and 3'-end modification. In some embodiments, the 5'-end modification comprises a phosphorothioate (PS) internucleotide linkage. In some embodiments, the 5'-end modification comprises 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modified nucleotides. In some embodiments, the 5'-end modification comprises at least one phosphorothioate (PS) linkage and one or more of 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modified nucleotides. The end modification may comprise a phosphorothioate (PS), 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modification. Equivalent end modifications are also encompassed by the embodiments described herein. In some embodiments, the template RNA or gRNA comprises an end modification in combination with modifications of one or more regions of the template RNA or gRNA. Further exemplary modifications and methods for protecting RNA, e.g., gRNA, and its formula, are described in International Publication No. WO2018126176A1, which is incorporated herein by reference in its entirety.
[0370] In some embodiments, structure-guided and phylogenetic approaches are used to introduce modifications (e.g., 2'-OMe-RNA, 2'-F-RNA, and PS modifications) into template or guide RNAs, as described, for example, in Mir et al. Nat Commun 9:2641 (2018), incorporated herein by reference in its entirety. In some embodiments, the incorporation of 2'-F-RNA increases the thermal and nuclease stability of RNA:RNA or RNA:DNA duplexes, for example, while minimally interfering with C3'-endo sugar puckering. In some embodiments, 2'-F may be better tolerated than 2'-OMe at positions where 2'-OH is important for RNA:DNA duplex stability. In some embodiments, the crRNA includes one or more modifications that do not reduce Cas9 activity, such as C10, C20, or C21 (fully modified), as described in Supplementary Table 1 of Mir et al. Nat Commun 9:2641 (2018), which is incorporated by reference in its entirety. In some embodiments, the tracrRNA includes one or more modifications that do not reduce Cas9 activity, such as T2, T6, T7, or T8 (fully modified), as described in Supplementary Table 1 of Mir et al. Nat Commun 9:2641 (2018). In some embodiments, a crRNA including one or more modifications (e.g., as described herein) can be paired with a tracrRNA including one or more modifications, such as C20 and T2. In some embodiments, the gRNA comprises, for example, a chimera of a crRNA and a tracrRNA (e.g., Jinek et al. Science 337(6096):816-821(2012)). In embodiments, modifications of the crRNA and tracrRNA are mapped onto a single guide chimera, for example, to generate a modified gRNA with enhanced stability.
[0371] In some embodiments, gRNA molecules can be modified by the addition or subtraction of naturally occurring components, such as hairpins. In some embodiments, gRNAs can include gRNAs lacking one or more 3' hairpin elements, as described, for example, in International Publication No. WO 2018106727 (incorporated herein by reference in its entirety). In some embodiments, gRNAs can contain added hairpin structures, such as those added within the spacer region shown to increase the specificity of CRISPR-Cas systems in the teachings of Kocak et al. Nat Biotechnol 37(6):657-666 (2019). Further modifications, including examples of shortened gRNAs and specific modifications that improve in vivo activity, can be found in U.S. Patent Application Publication No. 20190316121 (incorporated herein by reference in its entirety).
[0372] In some embodiments, modifications to the template RNA are found using structure-guided and phylogenetic approaches (e.g., as described in Mir et al. Nat Commun 9:2641 (2018); incorporated herein by reference in its entirety). In embodiments, the modifications are identified by the inclusion or exclusion of the guide region of the template RNA. In some embodiments, the structure of a polypeptide bound to the template RNA can be used to determine nucleotides in contact with non-proteins of the RNA, and then, for example, to select for modifications that have a lower risk of disrupting the association of the RNA with the polypeptide. Additionally, secondary structures in the template RNA can be predicted in silico using software tools, for example, the RNA structure tool available at rna.urmc.rochester.edu / RNAstructureWeb (Bellaousov et al. Nucleic Acids Res 41:W471-W474 (2013); incorporated herein by reference in its entirety), to determine secondary structures for selecting modifications, e.g., hairpins, stems, and / or bulges.
[0373] Preparation of compositions and systems As will be appreciated by those skilled in the art, methods for designing and constructing nucleic acid constructs and proteins or polypeptides (such as the systems, constructs, and polypeptides described herein) are routine in the art. Generally, recombinant methods may be used. For general information, see Smales & James (Eds.), Therapeutic Proteins: Methods and Protocols (Methods in Molecular Biology), Humana Press (2005); and Crommelin, Sindelar & Meibohm (Eds.), Pharmaceutical Biotechnology: Fundamentals and Applications, Springer (2013). Methods for designing, preparing, evaluating, purifying, and manipulating nucleic acid compositions are described in Green and Sambrook (Eds.), Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).
[0374] The present disclosure provides, in part, nucleic acids, e.g., vectors, encoding a recombinant polypeptide described herein, a template nucleic acid described herein, or both. In some embodiments, the vector comprises a selectable marker, e.g., an antibiotic resistance marker. In some embodiments, the antibiotic resistance marker is a kanamycin resistance marker. In some embodiments, the antibiotic resistance marker does not confer resistance to a beta-lactam antibiotic. In some embodiments, the vector does not comprise an ampicillin resistance marker. In some embodiments, the vector comprises a kanamycin resistance marker and does not comprise an ampicillin resistance marker. In some embodiments, the vector encoding the recombinant polypeptide integrates into the target cell genome (e.g., upon administration to a target cell, tissue, organ, or subject). In some embodiments, the vector encoding the recombinant polypeptide does not integrate into the target cell genome (e.g., upon administration to a target cell, tissue, organ, or subject). In some embodiments, the vector encoding the template nucleic acid (e.g., template RNA) does not integrate into the target cell genome (e.g., upon administration to a target cell, tissue, organ, or subject). In some embodiments, when the vector is integrated into the target site in the target cell genome, the selectable marker is not integrated into the genome. In some embodiments, when the vector is integrated into the target site in the target cell genome, genes or sequences involved in vector maintenance (e.g., plasmid maintenance genes) are not integrated into the genome. In some embodiments, when the vector is integrated into the target site in the target cell genome, import regulatory sequences (e.g., inverted terminal sequences, e.g., from AAV) are not integrated into the genome. In some embodiments, administering a vector (e.g., encoding a recombinant polypeptide described herein, a template nucleic acid described herein, or both) to a target cell, tissue, organ, or subject results in integration of a portion of the vector into one or more target sites in the genome of the target cell, tissue, organ, or subject.In some embodiments, less than 99, 95, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, 4, 3, 2, or 1% of the target sites containing integrated material (e.g., no target sites) contain a selectable marker (e.g., an antibiotic resistance gene), an import regulatory sequence (e.g., an inverted terminal end sequence, e.g., from an AAV), or both, from the vector.
[0375] Exemplary methods for producing pharmaceutical proteins or polypeptides described herein include expression in mammalian cells, although recombinant proteins can also be produced using insect cells, yeast, bacteria, or other cells under the control of an appropriate promoter. Mammalian expression vectors can include non-transcriptional elements, such as an origin of replication, a suitable promoter, and other 5'- or 3'-flanking non-transcribed sequences, as well as 5'- or 3'-untranslated sequences, such as necessary ribosome binding sites, polyadenylation sites, splice donor and acceptor sites, and termination sequences. DNA sequences derived from the SV40 viral genome, such as the SV40 origin, early promoter, splice, and polyadenylation sites, can be used to provide other genetic elements required for expression of heterologous DNA sequences. Cloning and expression vectors suitable for use in bacterial, fungal, yeast, and mammalian cell hosts are described in Green & Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).
[0376] Various mammalian cell culture systems can be used to express and produce recombinant proteins. Examples of mammalian expression systems include CHO, COS, HEK293, HeLA, and BHK cell lines. Host cell culture methods for protein therapeutic production are described in Zhou and Kantardjieff (Eds.), Mammalian Cell Cultures for Biologics Manufacturing (Advances in Biochemical Engineering / Biotechnology), Springer (2014). The compositions described herein can include a vector, e.g., a viral vector, e.g., a lentiviral vector, encoding the recombinant protein. In some embodiments, the vector, e.g., the viral vector, can include a nucleic acid encoding the recombinant protein.
[0377] Purification of protein therapeutics is described in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013); and Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).
[0378] The present disclosure also provides compositions and methods for generating template nucleic acid molecules (e.g., template RNA) with specificity for a recombinant polypeptide and / or a genomic target site. In one aspect, the method includes generating an RNA segment that includes an upstream homology segment, a heterologous sequence of interest segment, a recombinant polypeptide binding motif, and a gRNA segment.
[0379] Purpose therapeutic use In some embodiments, the genetic engineering systems described herein can be used to modify cells (e.g., animal cells, plant cells, or fungal cells). In some embodiments, the genetic engineering systems described herein can be used to modify mammalian cells (e.g., human cells). In some embodiments, the genetic engineering systems described herein can be used to modify cells from livestock animals (e.g., cows, horses, sheep, goats, pigs, llamas, alpacas, camels, yaks, chickens, ducks, geese, or ostriches). In some embodiments, the genetic engineering systems described herein can be used as experimental or research tools or in experimental or research methods, for example, to modify animal cells, e.g., mammalian cells (e.g., human cells), plant cells, or fungal cells.
[0380] Transgenic systems can address therapeutic needs by incorporating a coding gene into an RNA sequence template, for example, by providing expression of a therapeutic transgene in individuals with a loss-of-function mutation, by replacing a gain-of-function mutation with a normal transgene, by providing regulatory sequences to eliminate gain-of-function mutation expression, and / or by controlling expression of operably linked genes, transgenes, and the system. In certain embodiments, the RNA sequence template encodes a promoter region specific to the therapeutic needs of the host cell, e.g., a tissue-specific promoter or enhancer. In yet other embodiments, a promoter can be operably linked to the coding sequence.
[0381] In some embodiments, the insertion, deletion, substitution, or a combination thereof increases or decreases expression (e.g., transcription or translation) of the target gene. In some embodiments, the insertion, deletion, substitution, or a combination thereof increases or decreases expression (e.g., transcription or translation) of the target gene by modifying, adding, or deleting sequences in a promoter or enhancer, such as sequences that bind to transcription factors. In some embodiments, the insertion, deletion, substitution, or a combination thereof alters translation of the target gene (e.g., alters the amino acid sequence), inserts or deletes start or stop codons, alters or restores the translation frame of the gene. In some embodiments, the insertion, deletion, substitution, or a combination thereof alters splicing of the target gene, for example, by inserting, deleting, or modifying a splice acceptor or donor site. In some embodiments, the insertion, deletion, substitution, or a combination thereof alters the half-life of the transcript or protein. In some embodiments, the insertion, deletion, substitution, or a combination thereof alters, increases, or decreases the activity of the target gene, e.g., the protein encoded by the target gene.
[0382] Compensatory edit In some embodiments, a compensatory edit can be introduced using the systems or methods provided herein. In some embodiments, the compensatory edit is present in a location of the gene associated with the disease or disorder that is different from the location of the disease-causing mutation. In some embodiments, the compensatory mutation is not present in the gene containing the causative mutation. In some embodiments, the compensatory edit can neutralize or compensate for the disease-causing mutation. In some embodiments, the compensatory edit can be introduced by the systems or methods provided herein to suppress or reverse the effect of a variant of the disease-causing mutation.
[0383] AdjustmentsEdit In some embodiments, a regulatory edit can be introduced using the systems or methods provided herein. In some embodiments, the regulatory edit is introduced into a regulatory sequence of a gene, such as a gene promoter, a gene enhancer, a gene repressor, or a sequence that controls gene splicing. In some embodiments, the regulatory edit increases or decreases the expression level of a target gene. In some embodiments, the target gene is the same as a gene containing a disease-causing mutation. In some embodiments, the target gene is different from the gene containing the disease-causing mutation.
[0384] Repeat expansion disease In some embodiments, the systems or methods provided herein can be used to treat repeat expansion diseases. In some embodiments, the systems or methods provided herein, for example, those comprising genetically engineered polypeptides, can be used to treat repeat expansion diseases by resetting the number of repeats at a genetic locus according to a customized RNA template.
[0385] Treatment indications In some embodiments, the systems or methods provided herein can be used to treat any of the indications in Tables 12-15 below. For example, in some embodiments, the genetic modification system modifies a target site in genomic DNA in a cell, where the target site is present in a gene in any of Tables 12-15, e.g., in a subject with a corresponding indication listed in any of Tables 12-15. In some embodiments, the cell is a liver cell, and the target site is present in a gene in Table 12, e.g., in a subject with a corresponding indication listed in Table 12. In some embodiments, the cell is a hematopoietic stem cell (HSC), and the target site is present in a gene in Table 13, e.g., in a subject with a corresponding indication listed in Table 13. In some embodiments, the cell is a central nervous system (CNS) cell, and the target site is present in a gene in Table 14, e.g., in a subject with a corresponding indication listed in Table 14. In some embodiments, the cell is an ocular cell, and the target site is present in a gene in Table 15, e.g., in a subject with a corresponding indication listed in Table 15. In some embodiments, the target site is present in the coding region of the gene. In some embodiments, the target site is present in a promoter. In some embodiments, the target site is present in the 5' UTR or 3' UTR of a gene in any of Tables 12-15. In some embodiments, the target site is present in an intron or exon of the gene. In some embodiments, the genetic modification system corrects a mutation in the gene. In some embodiments, the genetic modification polypeptide inserts a sequence that was deleted from the gene (e.g., through a disease-causing mutation). In some embodiments, the genetic modification system deletes a sequence that was duplicated in the gene (e.g., through a disease-causing mutation). In some embodiments, the genetic modification system replaces a mutation (e.g., a disease-causing mutation) with the corresponding wild-type sequence. In some embodiments, the mutation is a substitution, insertion, deletion, or inversion.
[0386] [Table 12-1]
[0387] [Table 12-2]
[0388] [Table 13]
[0389] [Table 14]
[0390] [Table 15]
[0391] Application to plants In some embodiments, the systems or methods provided herein can be used to modify plants or plant parts (e.g., leaves, roots, flowers, fruits, or seeds), for example, to increase the fitness of the plant.
[0392] Delivery to plants Provided herein are methods for delivering the genetically modified systems described herein to plants, including methods for delivering the genetically modified systems to plants by contacting the plant, or a portion thereof, with the genetically modified system. These methods are useful for modifying plants, for example, to enhance the fitness of the plant.
[0393] More specifically, in some embodiments, a nucleic acid described herein (e.g., a nucleic acid encoding a transgenic system) can be encoded within a vector and inserted adjacent to a plant promoter, such as the maize ubiquitin promoter (ZmUBI), in a plant vector (e.g., pHUC411). In some embodiments, a nucleic acid described herein is introduced into a plant (e.g., japonica rice) or plant part (e.g., plant callus) via agrobacteria. In some embodiments, the systems and methods described herein can be used in plants by replacing a plant gene (e.g., hygromycin phosphotransferase (HPT)) with a null allele (e.g., containing a base substitution in the start codon). Systems and methods for modifying plant genomes are described in Xu et al., "Development of plant prime-editing systems for precise genome editing," 2020, Plant Communications.
[0394] In one aspect, provided herein is a method of increasing the fitness of a plant, the method comprising delivering to a plant a genetically modified system described herein (e.g., in an effective amount and for a period of time) to increase the fitness of the plant relative to an untreated plant (e.g., a plant that has not been delivered the genetically modified system).
[0395] The increased fitness of the plant as a result of the delivery of the genetically modified system can be maintained in several ways, for example, to achieve improved plant production, such as increased yield, improved plant vigor or quality of the product harvested from the plant, improved pre- or post-harvest characteristics deemed desirable for agriculture or horticulture (e.g., taste, appearance, shelf life), or improved characteristics that are otherwise beneficial to humans (e.g., reduced allergen production). Improved plant yield refers to an increase in the yield of a product of the plant (e.g., as measurable by plant biomass, grain, seed, or fruit yield, protein content, carbohydrate or oil content, or leaf area) by a measurable amount relative to the yield of the same product of a plant produced under the same conditions but without application of the composition, or compared to the application of a conventional plant modifier. For example, yield may be increased by at least about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100%, or more than 100%. In some cases, the method is effective to increase yield by about 2-fold, 5-fold, 10-fold, 25-fold, 50-fold, 75-fold, 100-fold, or more than 100-fold compared to untreated plants. Yield can be expressed in terms of plant weight or volume, or plant product based on some standard. Standards can be expressed in terms of time, cultivated area, weight of plant produced, or amount of raw material used. For example, such methods can increase the yield of plant tissues such as, but not limited to, seeds, fruits, grains, nuts, tubers, roots, and leaves.
[0396] Increased fitness of a plant as a result of delivery of a genetically modified system can also be measured by other means, for example, an increase or improvement in vigor assessment, canopy (number of plants per area unit), plant height, stem condition, stem length, number of leaves, leaf size, canopy, appearance (e.g., dark green leaf color, etc.), root assessment, emergence, protein content, increased tillering, increased leaf size, increased leaves, reduced basal leaf emergence, stronger shoots, reduced fertilizer requirements, reduced seed requirements, more productive shoots, earlier flowering, earlier grain or seed maturity, reduced plant verse (lodging), increased shoot growth, earlier germination, or any combination of these factors, by a measurable or discernible amount compared to the same factor in a plant product produced under the same conditions but without administration of the composition or with application of a conventional plant modifier.
[0397] Accordingly, provided herein are methods of modifying plants, comprising delivering an effective amount of any of the genetically modified systems prov...
Claims
1. A recombinant polypeptide comprising: Cas domain; a polymerase (Pol) domain of Table 1 or Table 23 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto, wherein the Pol domain is C-terminal to the Cas domain; and a linker disposed between the Pol domain and the Cas domain A recombinant polypeptide comprising: (a) the linker has a sequence from Table 6 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto; and / or 2. The recombinant polypeptide of claim 1, wherein (b) the Pol domain has a sequence with at least 90% identity to the Pol domain of Table 1 or 23.
3. 2. The genetically engineered polypeptide of claim 1, wherein the Cas domain comprises a sequence in Table 4 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
4. The Cas domain is (a) a Cas nickase domain; (b) a Cas9 nickase domain; and / or (c) the recombinant polypeptide of claim 1, comprising an N863A mutation.
5. 2. The recombinant polypeptide of claim 1, which comprises an NLS, for example, two NLS.
6. comprising an NLS N-terminal to the Cas9 domain; and / or The recombinant polypeptide of claim 1 , comprising an NLS C-terminal to the Pol domain.
7. 10. A nucleic acid (e.g., DNA or RNA, e.g., mRNA) encoding the recombinant polypeptide of claim 1.
8. A cell comprising a recombinant polypeptide according to claim 1 or a nucleic acid according to claim 7.
9. 1. A system comprising: i) a recombinant polypeptide according to claim 1, and ii) a template nucleic acid, a) a gRNA spacer complementary to a portion of the target nucleic acid sequence; b) a gRNA scaffold that binds to the Cas domain of the genetically engineered polypeptide; c) a heterologous sequence of interest; and d) Primer binding site sequence (PBS sequence) A template nucleic acid comprising A system including:
10. The template nucleic acid is (a) RNA; (b) DNA; and / or (c) a nucleic acid sequence listed in Table 24 or a nucleic acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. The system of claim 9 , comprising:
11. The template nucleic acid sequence is, for example, as follows in order from 5' to 3': (i) a spacer sequence listed in Table 24 or a nucleic acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto; (ii) a scaffold sequence listed in Table 24 or a nucleic acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto; (iii) a PBS sequence listed in Table 24, or a nucleic acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; and / or (iv) the template nucleic acid comprises an entire template molecule sequence listed in Table 24, or a nucleic acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; The system of claim 10.
12. (a) the gRNA spacer and the gRNA scaffold comprise RNA; (b) the heterologous sequence of interest comprises DNA and the PBS sequence comprises RNA, or the heterologous sequence of interest and the PBS sequence comprise DNA; and / or (c) the recombinant polypeptide is (i) comprises an amino acid sequence listed in Table 23, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% sequence identity thereto; (ii) comprising the amino acid sequence of any one of nCas9-UL-Polθ_L, nCas9-UL-Polθ_M, nCas9-UL-Polθ_4x0q, nCas9-FL-Polθ_M, or nCas9-FL-Polθ_4x0q, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; (iii) comprises the amino acid sequence of a Cas domain of an engineered polypeptide listed in Table 23, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; (iv) comprising the Cas domain amino acid sequence of any one of nCas9-UL-Polθ_L, nCas9-UL-Polθ_M, nCas9-UL-Polθ_4x0q, nCas9-FL-Polθ_M, or nCas9-FL-Polθ_4x0q, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; (v) comprises a Pol domain amino acid sequence of a recombinant polypeptide listed in Table 23, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; (vi) comprising the Pol domain amino acid sequence of any one of nCas9-UL-Polθ_L, nCas9-UL-Polθ_M, nCas9-UL-Polθ_4x0q, nCas9-FL-Polθ_M, or nCas9-FL-Polθ_4x0q, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; (vii) the recombinant polypeptide comprises the Cas domain amino acid sequence and the Pol domain amino acid sequence of an recombinant polypeptide listed in Table 23, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto; and / or (viii) The system of claim 9, wherein the recombinant polypeptide comprises the Cas domain amino acid sequence and the Pol domain amino acid sequence of any one of nCas9-UL-Pol θ_L, nCas9-UL-Pol θ_M, nCas9-UL-Pol θ_4x0q, nCas9-FL-Pol θ_M, or nCas9-FL-Pol θ_4x0q, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
13. A pharmaceutical composition for use in modifying a target nucleic acid in a cell (e.g., a human cell), comprising the system of any one of claims 9 to 12 or a nucleic acid encoding the same.
14. Use of a system described in any one of claims 9 to 12 or a nucleic acid encoding the same in the manufacture of a drug for modifying a target nucleic acid in a cell (e.g., a human cell).
15. 1. A pharmaceutical composition for use in treating a subject having a disease or condition associated with a genetic defect, comprising: A pharmaceutical composition for use, comprising the system according to any one of claims 9 to 12, the polypeptide according to any one of claims 1 to 6, or DNA encoding the polypeptide, or the nucleic acid according to claim 7. (a) the disease or condition associated with a genetic defect is an indication listed in any of Tables 12 to 15, and / or the genetic defect is a defect in a gene listed in any of Tables 12 to 15, and / or (b) the subject is a human patient,
17. Use of a system described in any one of claims 9 to 12, a polypeptide described in any one of claims 1 to 6, or DNA encoding the same, or a nucleic acid described in claim 7 in the manufacture of a medicament for treating a subject having a disease or symptom associated with a genetic defect. (a) the disease or condition associated with a genetic defect is an indication listed in any of Tables 12 to 15, and / or the genetic defect is a defect in a gene listed in any of Tables 12 to 15, and / or (b) the subject is a human patient.