Synthetic genome editing system
Patent Information
- Application Number
- JP2024521748
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2022-10-07
- Publication Date
- 2025-10-15
AI Technical Summary
Existing RNA-guided CRISPR-Cas9 systems for genome editing are limited by size, which restricts efficient AAV packaging, and immunogenicity is a concern for in vivo therapeutic applications, necessitating a more efficient and smaller system for targeted DNA modification.
A synthetic modular system comprising a targeting nucleic acid and a modular polypeptide with a DNA binding domain (DBD) that recognizes predefined sequences, allowing for site-specific single-strand or double-strand breaks, utilizing a smaller artificial nickase or nuclease for efficient genome editing.
The system enables efficient and targeted genome editing with reduced size, facilitating delivery through various vectors and minimizing immunogenicity, thereby enhancing therapeutic applications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present invention provides a synthetic modular system for DNA modification that can be used for genome editing, comprising a targeting nucleic acid that has both DNA targeting ability and ability to bind to the recognition module of a modular polypeptide, and the modular polypeptide also comprises an effector component as a separate module. The effector component can be a protein component for modifying the structure or regulation of a gene, either alone or, for example, after dimerization. Desirably, the effector component can be an artificial nickase that comprises a multimer of linked, self-assembling short peptides, as also taught herein. By further including a short peptide DNA binding component that lacks enzymatic activity and recognizes a predetermined sequence in the target in the modular polypeptide component, a completely synthetic gene modification tool is obtained that is much smaller than the CRISPR-Cas9 single guide RNA (sgRNA) gene modification system, thereby facilitating delivery to cells. The polypeptide component can be delivered together with the targeting nucleic acid(s) in a single AAV vector. [Background technology]
[0002] The RNA-guided CRISPR-Cas9 system and other RNA-guided CRISPR-Cas systems have revolutionized the field of gene editing since the first studies revealed the ability of sgRNA to direct CRISPR-Cas9 double-stranded DNA breaks by protospacer adjacent motif (PAM) recognition in target DNA. However, several drawbacks remain to be addressed for using such RNA-guided DNA targeting systems for all desired purposes, especially therapeutic genome modification. One such recognized disadvantage is the size that limits efficient AAV packaging. Streptococcus pyogenes Cas9 (SpCas9) with sgRNA (the most widely used Cas9 version) is 4.1 kbp (1368 amino acids) in size, and the more recently identified smaller cas enzyme is still a nuclease of significant size and is far from optimal for the desired vector delivery. The overall size of AAV can only package heterologous sequences up to 5 kb.
[0003] Furthermore, immunogenicity is a consideration for in vivo therapeutic applications in humans. Streptococcus pyogenes Cas9-reactive T cells have been reported in the adult human population.
[0004] WO 2013 / 088446 of Targetgene Biotechnologies Ltd. proposes an alternative DNA editing system that uses a Fok1 nuclease domain linked to a specificity-conferring nucleic acid (SCNA) via a linking domain. The SCNA comprises a nucleotide sequence complementary to the sequence of the target nucleic acid and a recognition component that can specifically link the SCNA to the linking domain, for example, the recognition component of the SCNA can be a naturally occurring RNA aptamer that recognizes one of a specific binding partner pair or its cognate polypeptide binding component. The pairing of such an artificial nucleoprotein complex on the target DNA is required to allow dimerization of the Fok1 nuclease domain. Furthermore, the Fok1 nuclease thus provided must be able to effectively interact with DNA to cleave double-stranded DNA at the target site. However, WO 2013 / 088446 only provides prophetic examples of therapeutic utility, leaving doubts about the efficiency of such a system. Summary of the Invention
[0005] The present inventors aimed to provide a fully synthetic modular system for gene modification, in this case as an alternative to the use of the CRISPR-Cas / sgRNA system, which in nuclease mode provides good efficiency for DNA targeting double-stranded or single-stranded cleavage at a predefined target site, and can be designed to provide a gene editing system with a much smaller size. The small size achievable using the currently taught artificial nickases provides the opportunity to choose a wide variety of delivery means, including single AAV vector delivery with one or more targeting RNAs.
[0006] In one aspect, the invention provides a nucleoprotein complex for use in modifying a target nucleic acid, e.g., a target DNA, comprising: (A) a targeting nucleic acid; and (B) a modular polypeptide component, The targeting nucleic acid (i) a targeting nucleic acid element (also referred to herein as a targeting element, TE) that is complementary to a region of the target; (ii) a recognition element (e.g., an RNA scaffold, RS) that specifically interacts with a nucleic acid recognition module (e.g., an RNA scaffold binding domain, RSBD) of the modular polypeptide component; and (iii) a connecting sequence linking the targeting element and the recognition element; The modular polypeptide components are arranged as linked, separate functional modules: (a) the nucleic acid recognition module lacking enzymatic activity; (b) a DNA binding domain (DBD) that recognizes a predetermined sequence in a target, where module (a) and said DBD are not found in conjunction in nature; and (c) an effector component for use in modifying a target, whereby a site-specific modification of a targeting nucleic acid occurs directed by the targeting element and the DBD; and (d) optionally or if necessary, including a first linker and / or a second linker by which modules (a), (b) and (c) are joined; The present invention provides a nucleoprotein complex, wherein the DBD is a polypeptide sequence of 70 to 75 amino acid residues or less, for example, 30-mer or less, 25-mer or less, 20-mer or less, preferably 15-mer or less.
[0007] Without being bound by theory, the intended purpose of DBD is to destabilize the structure of targeted dsDNA when bound to a predefined sequence (PDS) in such a way as to promote the access of the targeting nucleic acid sequence to the dsDNA structure for hybridization. Such functional ability can be assessed, for example, as illustrated in Example 1, using a dsDNA containing a selected PDS and T7 endonuclease 1 (T7E1) to evaluate the DNA structure. T7E1 is well known as a structure-selecting enzyme that detects structural anomalies in dsDNA. Other methods of detecting desired structural deformations when DBD polypeptide binds to PDS are recognized.
[0008] Although the DBD may have a binding affinity for a single predefined sequence, it may have the ability to bind to more than one predefined sequence. The prerequisite is that the DBD cooperates with the targeting element of the targeting nucleic acid to direct site-specific modification. However, the DBD is considered to be more than just a secondary tool for positioning at the required target site, and is considered to have a structural role in facilitating the hybridization of the targeting element, thereby facilitating the overall desired target modification. Thus, it is recognized that the provision of a short DBD polypeptide sequence in this case cannot be compared to, for example, using zinc finger proteins or TALE arrays only to position an effector or part of an effector to a specific DNA sequence.
[0009] Typically, the effector component of the modular polypeptide of the invention will be a terminal module. Preferably, the nucleic acid recognition module (e.g., RSBD) will also be a terminal module. Thus, preferably, the modular polypeptide component will be: (a) the nucleic acid recognition module; (b) a DNA binding domain (DBD) directly linked to (a) that recognizes a predetermined sequence in a target, wherein module (a) and the DBD are not found linked in nature; (c) an effector component directly linked to (b) for use in modifying a target, whereby site-specific modification of a targeting nucleic acid is directed by a targeting element and the DBD; and (d) optionally or if necessary, a first linker and / or a second linker that provide a direct linkage of the DBD to the effector component and the nucleic acid recognition module, respectively; Includes.
[0010] The modules of modular polypeptide components are generally linearly linked.Therefore, modules (a), (b) and (c) can be linearly joined, for example, to (a) at N-terminus and to effector component at C-terminus, or vice versa.In this way, the interaction between targeting nucleic acid and recognition module of modular polypeptide components is separated from effector component.
[0011] The targeting nucleic acid may be RNA, and preferably is entirely RNA. It is recognized that the nucleic acid targeting element is comparable in functional purpose to CRISPR-Cas9 gRNA. As mentioned above, conveniently and preferably, the targeting element may be joined via a connector with an RNA motif that provides an RNA scaffold (possibly alternatively called an RNA aptamer) that binds to a cognate protein or peptide module in a modular polypeptide. In this case, the RNA scaffold recognition module of the modular polypeptide component may also be called an RNA scaffold binding domain (RSBD). Thus, a nucleoprotein complex may conveniently comprise (i) a nucleic acid component (NAC) that can be expressed as RNA from a coding nucleic acid sequence, and (ii) a modular polypeptide component that can also be expressed from a single coding nucleic acid sequence.
[0012] Thus, in a preferred embodiment, a nucleoprotein complex of the invention for use in modifying a target nucleic acid, e.g., a target DNA, comprises (A) a complete nucleic acid component (NAC) and (B) a modular polypeptide component, NAC: (i) a targeting element (TE) that is complementary to a region of the target; (ii) an RNA scaffold (RS) that specifically binds to an RNA scaffold binding domain (RSBD) of said modular polypeptide component; and (iii) a connecting sequence linking the TE and RS wherein the modular polypeptide components comprise, as linked separate functional modules: (a) the RSBD; (b) a DNA binding domain (DBD) linked to (a), which recognizes a predetermined sequence in a target, module (a) and said DBD are not found linked in nature, (c) an effector component linked to (b) for use in modifying a target, whereby site-specific modification of NAC is directed by the targeting element and the DBD; and (d) optionally or if necessary, a first linker and / or a second linker linking the DBD to the effector component and RSBD, respectively; Including, As discussed above, the DBD is a polypeptide sequence of 70-75 or fewer amino acid residues, for example, a 30-mer or fewer, a 25-mer or fewer, a 20-mer or fewer, preferably a 15-mer or fewer. See FIG.
[0013] It will be appreciated that separate functional modules of a modular polypeptide component can be conveniently encoded by a single nucleic acid such that they are independently permuted and joined in a limited linear sequence.
[0014] The target can be any target nucleic acid, RNA or DNA, including viral nucleic acid.The target can be single-stranded DNA or more generally double-stranded DNA (dsDNA).The above-mentioned nucleoprotein complex of the present invention can be particularly preferred as an alternative to CRISPR-Cas system for delivery to cells to achieve genomic DNA modification, especially in eukaryotic cells, including plant cells and human cells.
[0015] As noted above, without being bound by theory, it is believed that the DBD, through its specific recognition of a defined sequence in double-stranded DNA, assists in the melting and / or unwinding of the double helix, thereby assisting in the targeting and action of the effector component at the desired site by the target nucleic acid.
[0016] A nuclear localization signal (NLS) and / or an organelle localization signal, e.g., a mitochondrial or chloroplast localization signal, may also be provided as part of the modular polypeptide component for efficient transport into the nucleus of a eukaryotic cell or into a desired cellular organelle. Such signal sequences are generally provided at the N- or C-terminus. Such signal sequences may be connected to one or more additional sequences, e.g., a detection tag that aids in detection, such as an epitope tag, provided that the required function of the modular polypeptide component is maintained; see FIG. 25.
[0017] Effector component does not specifically bind to said predetermined sequence, and generally lacks target binding ability.However, it is recognized that effector component only needs not to interfere with the site-specific targeting of desired modification by nucleic acid targeting element and DBD.This does not have to exclude effector component that has recognition ability for some target, for example, DNA.
[0018] As mentioned above, target modification may be a structural or chemical modification, e.g., a nucleotide sequence modification, or a regulatory modification, e.g., transcriptional activation.
[0019] The effector component may be any type of effector known to modify DNA, including, for example, a Fok1 nuclease domain (i.e., Fok1 nuclease minus its DNA binding domain) that can form a functional nuclease upon dimerization, for example, by appropriately directing two nucleoprotein complexes of the invention to a target DNA region. However, as mentioned above, the effector component may preferably include an artificial nuclease formed from linked self-assembling short peptides, as further discussed below. Such an artificial nuclease may function as an artificial nickase in that it cuts only one strand of double-stranded DNA at the site to which it is targeted. The effector component may be a fusion protein in which the nickase is linked to or replaced by another functional component, for example, a base editor or a reverse transcriptase.
[0020] As mentioned above, the DBD will be a short peptide, typically 20-mer or less, e.g., 16-mer to 18-mer or less, preferably 15-mer or less, selected for binding affinity to a predetermined sequence in the target, e.g., 2-7, preferably 3-6 nucleotides, e.g., the 6 nucleotide sequence 5'GAGGTC3' in the dsDNA target exemplified herein. The DBD is expected to aid in genome scanning and establish initial contacts next to the desired modification site, e.g., the cleavage site. It is further emphasized that binding of the DBD to the target DNA is expected to trigger a conformational change and subsequent melting at the adjacent site. Such unwinding should help the TE to interrogate adjacent sequences for complementarity and, if the TE is an RNA sequence, to form an RNA-DNA complex (R-loop). Thus, the DBD component of the genome editing tool of the present invention is considered to be an important module for improving genome modification efficiency, and can be provided without compromising the desire for a short total polypeptide component length, preferably suitable for expression from an AAV vector that also expresses at least one targeting RNA.
[0021] The modular polypeptide components of the genome editing tools of the present invention may therefore comprise fully synthetic short peptides with linkers that provide nickase or nuclease activity by the artificial nuclease at the target DNA site. Such genome editing tools using modular polypeptide components have been termed ApGet, which stands for "artificial peptidic genome editing tool". Such polypeptide components include: The following components fused N- to C-terminally: - RSBD of 22 amino acids corresponding to the lambda N22 protein,
[0022] [ka] - Linker
[0023] [ka] That is, (GGGGS)7 - DBD that binds to a dsDNA sequence represented by the 5' to 3' sequence 5'GAGGTC3'
[0024] [ka] - Linker
[0025] [ka] and - Nuclease Module
[0026] [ka] Minus, for example, a short NLS and any additional N- or C-terminal sequences that may be provided as described above, perhaps to provide a detection tag. Consists of: ApGet polypeptide of SEQ ID NO:1:
[0027] [ka] The sequence may be 220 or fewer amino acids (minus any cleavable tag, localization signal or any other sequence other than components (a)-(d)), exemplified by:
[0028] The nuclease (shown in bold in SEQ ID NO:1 above) has 10 identical 7-mer peptide units (IEIDIHI; SEQ ID NO:7), all of which are linked in the same N-to-C-terminal direction by a 4-mer linker (referred to herein as an example of a non-inverted decamer) to give a beta turn and additional N- and C-terminal flanking sequences. This allows folding, whereby a decamer of identical peptide units can give rise to a secondary structure resembling an antiparallel beta sheet in which successive monomer units appear oriented in opposite directions; see FIG. 8b and FIG. 9. It has been shown that the modular polypeptide building block of SEQ ID NO:1, as part of the DNA editing tool of the present invention, can cleave a single strand of dsDNA at a targeted site. Thus, the artificial nuclease module can be considered to give the function of a nickase, and thus may be referred to as an artificial nickase. It is expected that the more compact such genome editing tools can be achieved, for example, nucleases with less diversity of self-assembling peptides and / or variation in other components will be provided. This includes any different peptide sequences of the same pattern and combinations of different peptide sequences.
[0029] The artificial nuclease of SEQ ID NO:1 may be replaced with another effector component, for example, a transcription activator such as VP64 or another effector component discussed below. Furthermore, it is recognized that one or both of the flexible linkers may be replaced by alternative linkers. Thus, for example, linker 1 between the artificial nuclease (or alternative effector component) and the DBD may be replaced by a linker designated as linker L1a (SEQ ID NO:69). Linker 2 between the RSBD and the DBD may be replaced by a shorter linker, for example, exemplified by linker 2b (SEQ ID NO:70). The linkers in the modular polypeptide components of the invention may be miscellaneous and optimized for any pair of DBD binding site and target cleavage site. It is recognized that such variations are a matter of length requirement calculation based on known target sequence information and appropriate testing. Such linkers may also be designed to increase protein stability and / or solubility.
[0030] Although the lambda N22 protein was selected as the RSBD for the illustrative ApGet presented above as SEQ ID NO:1 on the basis that it is only a short 22 amino acid sequence with a well-known ability to recognize the 19 nucleotide box B lambda phage sequence (Baron-Benhamou et al. (2004) Methods Mol. Biol. 257, 135-54), it is recognized that it may additionally or alternatively be replaced by another RSBD, including an RNA scaffold, e.g., any variant thereof that retains the desired binding affinity to the same box B sequence. It is further recognized that RSs may be used in tandem to facilitate binding of multiple RSBDs. The lambda N22 protein RSBD paired with the box B sequence provided as the RS of the targeting nucleic acid may be replaced by any of several other well-known RNA aptamer-polypeptide binding domain pairings discussed further herein below.
[0031] The small size of the polypeptide components of ApGet is expected to produce targeted DNA strand breaks and improve the DNA scanning and recognition efficiency of the system compared to known gene editing systems that use larger molecular weight proteins such as ZFN, TALEN, meganucleases and CRISPR-Cas entities, which have a lower diffusion rate.
[0032] The examples provided herein further show that ApGet minus any nuclease or other effector component, i.e., just containing a nucleic acid recognition module, e.g., RSBD, linked to DBD and binding to a suitable targeting nucleic acid, may be useful for inhibiting transcription at a target site. Such a system is referred to herein as the ApGet-i system. In some cases, it may be chosen to express the ApGet-i polypeptide with a final C-terminal linker sequence, e.g., the linker of SEQ ID NO: 5 above. However, for transcription inhibition, the ApGet-i polypeptide minus such a linker is preferred.
[0033] It will be appreciated that the modular polypeptide components for the nucleoprotein complexes of the invention may be initially expressed as part of a longer polypeptide with N- and / or C-terminal extensions to aid, for example, in solubility and / or purification and / or detection in an expression system, e.g., a host cell such as E. coli or an in vitro expression system. Such extensions may include a protease cleavage site, e.g., a TEV protease cleavage site, whereby protease cleavage removes undesirable sequences in the final modular polypeptide for targeted modification; see Figures 10 and 31.
[0034] In some cases, one may choose to express modular polypeptide components that contain a protein transduction domain (PTD), sometimes referred to as a cell penetrating peptide (CPP) or domain that facilitates delivery across a membrane, e.g., into a target cell. Again, such a domain may be provided with a linker that provides a protease cleavage site. Such a PTD may be located at the N-terminus or C-terminus. In general, PTDs can be classified into three types: cationic peptides, e.g., 6-12 amino acids in length, containing primarily arginine, ornithine and / or lysine residues; hydrophobic peptides, such as leader sequences of secreted growth factors and cytokines; and cell-specific peptides. Exemplary PTDs include, but are not limited to, PTDs of the HIV TAT protein, polyarginine peptide sequences, VP22 domains (Zender et al. (2002) Cancer Gene Ther. 9,, 486-96), PDX1 protein transduction domain (Noguchi et al. (2003) Diabetes 52, 1732-1737), truncated human calcitonin peptides (Trehin et al. (2004) Pharma. Res. 21 1248-1256), and polylysine sequences (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97, 13003-13008). Activatable CPPs (ACPPs) can be provided that contain a polycationic CPP (e.g., Arg9 or R9) connected via a cleavable linker to a matching polyanion (e.g., Glu9 or E9) that reduces the net charge to near zero, thereby inhibiting attachment and uptake into cells (Aguilera et al. (2009) Integr. Biol. (Camb) 1, 371-381). Cleavage of the linker releases the polyanion, thus locally unmasking the polyarginine and activating its inherent adhesiveness to facilitate membrane transport.
[0035] The nucleoprotein complexes of the invention can be delivered to the host cell as active complexes, for example, by electroporation, or together with either or both of the polypeptide component expressed by the polynucleotide and the targeting nucleic acid.
[0036] Thus, in a further aspect, a nucleic acid or a combination of nucleic acids is provided for providing one or more nucleoprotein complexes of the present invention in a host cell.As mentioned above, the nucleic acid can preferably be a vector, for example, an AAV vector, which can express both modular polypeptide components and targeting nucleic acids, including artificial nickases of the present invention.More than one targeting nucleic acid can be provided to a host cell to target more than one site.For example, a pair of nucleoprotein complexes of the present invention, including artificial nickases, can be provided by vector delivery to a cell, possibly by a single vector, together with sequence templates for homologous recombination.
[0037] The application of the nucleoprotein complex(es) of the present invention extends to a full set of applications for DNA modification that may use CRISPR-Cas / sgRNA. Thus, as a further aspect of the present invention, a method for modifying one or more target nucleic acid sequences using one or more nucleoprotein complexes of the present invention or one or more nucleic acids for their provision in a host cell is provided, with the proviso that such a method as claimed does not extend to a method for modifying human germline identity or a medical method performed on the human or animal body itself. As mentioned above, desirably, a single modular polypeptide component for the provision of one or more nucleoprotein complexes can be provided in a cell by expression from a vector. The same vector can preferably also express the required one or more targeting nucleic acids.
[0038] As yet a further aspect of the present invention, there is provided a combination of (i) at least one modular polypeptide component for the genome modification tool of the present invention, or a polynucleotide encoding the same, and (ii) one or more targeting nucleic acids that can be linked to said polypeptide component(s), or one or more polynucleotides encoding the same, for use in the methods of the present invention discussed above or for use in therapeutic treatments. As mentioned above, both the polypeptide component and the one or more targeting nucleic acids can preferably be expressed from a single vector.
[0039] It is recognized that the artificial nuclease for use in the ApGet nucleoprotein complex discussed herein represents another aspect of the present invention. The use of self-assembling peptide clusters incorporating catalytic residues has been used previously to generate artificial enzymes that function as esterases or provide some other enzymatic reactions. However, in this case, the inventors have obtained for the first time an artificial nuclease that can provide nickase function in dsDNA using linked self-assembling short peptides, an approach designed to perform the required function while minimizing size and thereby providing an artificial enzyme that is well suited to provide an effector module in a genome editing tool. Linking the individual peptide units in this way also helps to control the degree of self-assembly.
[0040] Such artificial nucleases and genome-modified nucleoprotein complexes of the present invention are discussed in further detail below with reference to the figures that follow. [Brief description of the drawings]
[0041] [Figure 1]Schematic diagram of genome editing tool of the present invention consisting of a nucleic acid component (NAC) and a modular polypeptide component. NAC is a targeting RNA, where TE is a targeting element that provides a complementary sequence to the target strand region. RS is an RNA scaffold part of NAC that binds to an RNA scaffold binding domain (RSBD), whereby NAC interacts with the modular polypeptide component. TE and RS form a single nucleic acid with a connecting sequence that joins two functional binding sequences so that the required binding can occur. DBD is the DNA binding domain of the modular polypeptide component, which recognizes a predefined sequence (PDS) of the target DNA, L1 is a linker that connects the DBD to an effector protein for genome modification, such as a nuclease (endonuclease for double-stranded break or nickase for single-stranded break), a transcription activator or repressor, or a methylase, and L2 is a linker that connects the DBD and the RSBD. The target sequence can be immediately adjacent to the PDS or spaced from the PDS. [Diagram 2] More specifically, it is a schematic diagram of the ApGet system, in which the modular polypeptide components include fusion artificial nucleases. The TE is a targeting element of the targeting RNA that hybridizes to a sequence in the target DNA region for nickase action. The recognition element of the targeting RNA is joined to the TE by a connecting sequence, providing an RNA scaffold (RS) that binds to the RNA scaffold binding domain (RSBD). The DBD is a DNA binding domain that binds to a predefined sequence (PDS) on one strand of the target dsDNA. L1 and L2 are peptide linkers. Binding of the DBD to the PDS promotes hybridization of the TE to its complementary sequence. [Figure 3-1]Figure 3a-c: Evidence of specific binding and unwinding of dsDNA targets by selected phage clones exhibiting a 12-mer polypeptide. The dsDNA targets exhibited two PDS sequences: 5'-GAGGTC-3' (PDS1) and 5'-ACGGGT-3' (PDS2). Selected phage clones were incubated with dsDNA containing the PDS1 (reactions 3, 4, 5, 10, 11 and 12) or PDS2 (reactions 1, 2, 8 and 9) sequences and with T7E1 nuclease (reactions 1-5) or Surveyor nuclease (reactions 8-12). Mock reactions were incubations of dsDNA with T7E1 (reactions 6 and 7) or Surveyor (reactions 13 and 14) without the selected phage clones. When phage clones bind to dsDNA, partial unwinding can occur, generating structures resembling heteroduplex DNA and providing suitable substrates for endonucleolytic cleavage by T7E1 and Surveyor nucleases, resulting in faster migrating DNA bands (see arrows in Figure 3a). Quantification of T7E1 and Surveyor-mediated cleavage was performed by densitometric analysis using ImageJ software. The results are shown in Figures 3b and 3c, respectively. [Figure 3-2] Same as above. [Figure 4] Cleavage assay results using T7E1 enzyme for binding of chemically synthesized peptides 1, 2, 3 and 4 (Table 2 in Example 1) to two 6 nt sequences as PDS1 and PDS2 named above to dsDNA target and unwinding of dsDNA target. Arrows indicate bands that increase upon T7E1 endonucleolytic cleavage of dsDNA target. These results led to the selection of peptide 1 sequence (SEQ ID NO: 4, also shown as SEQ ID NO: 37 in Table 2, incorporating a GGS linker sequence) as the DNA binding domain for 6 nt PDS 5'-GAGGTC-3'. [Diagram 5]Schematic diagram illustrating FRET experiments to evaluate specific binding and unwinding of dsDNA targets by peptide binding to PDS along with RNA hybridization. Results are also shown for binding of peptide 1 (SEQ ID NO: 37 shown in Table 2) to PDS1 with the sequence GAGGTC 5' to 3'. FRET signal was determined upon incubation of PDS1 containing dsDNA target with 5'FAM-RNA with (2, 4) or without (1, 3) peptide. Binding of peptide 1 has been shown to promote hybridization of RNA to the sequence adjacent to PDS1, as evidenced by an increase in FRET signal. [Figure 6] Schematic diagram of the genome editing tool of the present invention showing possible alternative strand hybridization of the targeting element (TE) of the targeting RNA. RS is an RNA scaffold for linking the targeting RNA to the modular polypeptide component via the RNA scaffold binding domain (RSBD). DBD is a DNA binding domain that binds to a short predefined sequence (PDS) to promote heteroduplex formation. DBD is linked to a C-terminal effector protein via a linker for nucleic acid modification. [Figure 7] Luminescence signal of Thioflavin T (Tht) in the presence of the 7 amino acid peptide LELDLHL (bD peptide; SEQ ID NO: 8), indicating the formation of amyloid β-sheet structure. [Figure 8]FIG. 2 is a design for an artificial nuclease using (a) a pentamer or (b) a decamer of the same beta-sheet forming peptide linked to adjacent peptide units with increasing positive charge due to replacement of hydrophobic residues with positively charged residues such as lysine or arginine. H=hydrophobic residue. C=catalytic amino acid. N=negative design peptide unit (limits or prevents intermolecular polymerization), in this case C-terminal and N-terminal peptide units with increasing positive charge. K=lysine residue replacing hydrophobic residue of the beta-sheet forming peptide. The C-terminal adjacent peptide may have a final flexible soluble tail, e.g., a 4-mer NQRS. Each of the preceding peptide units is linked by a linker of 4 amino acid residues that provides, e.g., a beta turn. [Figure 9] "Non-inverted and inverted" artificial nuclease designs. "Non-inverted" means that the individual beta-strand units of the nuclease share the same primary sequence in the N-to-C-terminal direction, whereas in the "inverted" design, the primary sequence of consecutive beta-strand units is inverted. Boxes 1-7 show the amino acid sequences of the beta peptides in the N-to-C-terminal direction. [Figure 10]Diagram illustrating the construction of pentameric and decamer beta-protein nucleases in the complete modular polypeptide building block of ApGet. The beta-peptide unit (black) is constructed using 5 or 10 beta-peptide units. The units are connected using a series of beta-turns like connecting loops. Each design contains an additional adjacent beta-peptide unit, i.e., a negative design (adjacent arrow), that incorporates a charged residue (lysine, K) to limit intermolecular polymerization. The orientation of the beta-peptide sequence is modified to mimic inversion or non-inversion in an attempt to mimic parallel or antiparallel beta-sheet folding (see again boxes 1-7 in Figure 9 showing the amino acid sequence of the beta-peptides in the N-to-C-terminal direction). The N-terminus of each artificial beta-protein nuclease design is fused with linker 1, just like the ApGet-1.0 protein, onto the C-terminus of the ApGet-i protein, which provides the remainder of the modules required for a complete synthetic genome editing system. ApGet-i can contain several purification / solubility tags to enhance protein expression and solubility. Such a tag is shown providing a maltose binding protein fused to a 6xHis tag at both the N-terminus and C-terminus. The C-terminal His tag is separated from the N-terminus of ApGet by a TEV cleavage site that allows removal of the spacer unit and the entire multi-element tag. [Figure 11]Gel analysis of plasmid DNA cleavage by in vitro transcription and translation (IVTT) expressed ApGet1.0 protein with 10-mer IbD peptide. ApGet-i (expressed with the additional DBD enzyme linker of ApGet1.0 at the C-terminus) and no protein controls were run for each buffer condition to control for activity allowed by the buffer composition or by contaminating protein components of the IVTT mixture. Plasmid DNA was mixed with the same amount of total IVTT protein mixture (by Ab 280 nm) for ApGet1.0 10-mer non-inverted, inverted, ApGeti and water control in reaction buffers with no metal or containing 5 mM Mn2+ or Mg2+. A secondary reaction was also run comparing cleavage of IVTT protein and nuclease inhibitor protein in the presence of Mn2+. [Figure 12] FIG. 1 is a diagram of the plasmid used to test the recruitment of ApGet-i by targeting nucleic acids to the PDS target sequence site in E. coli cells. [Figure 13] Results of testing ApGet-i to reduce transcriptional activity on (a) detector for eYFP expression and (b) detector for LacZα expression in EPI300 E. coli cells. Exemplary Tables 5 and 6 summarize the construct testing. Values presented are the average values of eYFP / OD600 or LacZα / OD600 from triplicate experiments, with standard deviations shown. [Figure 14] Diagram of the "editor" and "detector" plasmids used to test the complete ApGet system in E. coli cells. The ApGet modular polypeptide components had the amino acid sequence of SEQ ID NO: 1, but had the Fok1 nuclease domain instead of the artificial nuclease. [Figure 15]ApGet-FokI editing results obtained using the "editor" and "detector" plasmids exemplified in Figure 14. Editing with the ApGet-Fok1 nuclease construct is evident from the disappearance of the target-containing detector plasmid resulting in reduced bacterial growth on selective medium (see bars 1-3). Results shown are from co-transformation of EPI300 E. coli cells with the ApGet-FokI expressing editor plasmid and a detector plasmid containing the "PDSin" configuration of the target and different lengths of spacers between the target sequences. Detection of editing activity was by measuring OD600 one day after induction of ApGet with anhydrotetracycline (ATC). Control experiments included co-transformation of the ApGet-i expressing editor (ApGet construct, lacking the nuclease domain) or a non-ApGet expressing plasmid with the detector plasmid into EPI300 cells. Bars correspond to expression of ApGet-Fok1 constructs and detector targets as follows: 1. ApGet-Fok12.TE with TE and 5 nt spacer detector and ApGet-Fok13.TE with 8 nt spacer detector and ApGet-Fok14.TE with 14 nt spacer, ApGet5.TE without Fok1 and 5 nt spacer detector, ApGet6.TE without Fok1 and 8 nt spacer detector, ApGet7.vector backbone alone and 5 nt spacer detector 8. vector backbone alone and 8 nt spacer detector 9. vector backbone alone and 14 nt spacer detector. [Figure 16] FIG. 1 is a schematic diagram of the "editor" plasmids used to test the ApGet system with engineered nucleases in bacterial cells. [Figure 17]13 is a schematic diagram of the "detector" plasmid used to test the ApGet system with artificial nuclease in bacterial cells. The detectors were designed in the "PDS-out" configuration (construct f) or the "PDS-in" configuration (construct g) for a pair of PDS target sequence elements. The target and PDS sequences of the PDSin and PDSout detectors are located exactly between the homology arms as in FIG. 14 and are identical to the target and PDS sequences in FIG. 14. Successful targeting is indicated by the observation of nanoluciferase activity in the presence of the appropriate HDR template. The homology-directed repair (HDR) template for nanoluciferase was constructed on the same plasmid expressing ApGet as a separate unit. [Figure 18] Figures 18 and 19 are schematic diagrams showing possible alternative heteroduplex interactions when a pair of genome editing tools of the present invention is used to modify, e.g., cleave, both strands of a target genomic DNA at a selected site. In Figure 18, each polypeptide component is shown as comprising a Fok1 nuclease domain. Two such nuclease domains are shown as dimerizing to provide a functional endonuclease by using two targeting RNAs that target different sequences on opposite strands of the target DNA. Two DNA binding domains are shown as each binding to a PDS in a completely complementary region of the target DNA where the Fok1 endonuclease cleaves. Both PDS sequences are inside the pair of target sequences. This is referred to as using a PDS-in position. In Figure 19, each polypeptide component is shown as comprising a fusion enzyme, which may be, for example, a fusion artificial nickase as described herein. In this case, each DBD is shown as binding to a PDS outside of a single heteroduplex region formed by the targeting sequences of two different targeting RNAs that hybridize to different target strands. Each pair of enzyme domains acts on a different strand within the heteroduplex region. Because the pair of PDS sequences is outside the pair of target sequences, this diagram is referred to as using the PDS-out position. [Figure 19] Same as above. [Figure 20] Figure 20a and b: Results of the test of ApGet editing to restore functional nanoluciferase by triggering homology-directed repair of the disrupted nanoluciferase open reading frame. Different editor plasmids containing the expression cassette for ApGet1.0 (Figure 16) in different configurations and carrying the expression cassette for the nanoluciferase HDR template were co-transformed with a detector plasmid expressing the disrupted open reading frame of nanoluciferase and carrying the target in the "PDSin" and "PDSout" configuration (Figure 17). Colonies of co-transformants of ApGet1.0 with detector were grown in liquid medium with IPTG or with and without IPTG for the same constructs and corresponding antibiotics and the Nanoluc / OD600 signals were analyzed the next day. Error bars are standard deviations of three experiments. [Figure 21] FIG. 1 illustrates the ApGet-i expression editor plasmid used to assess the flexibility of the DBD of SEQ ID NO:4 for binding to mutants of PDS1 (5′-GAGGTC-3′). [Figure 22] Same as above. [Diagram 23] 23a and b show the results for the binding ability of the control plasmid and the tested DBDs to various 6-mer PDSs used together with the detector plasmid of FIG. [Figure 24] FIG. 1 is a schematic diagram of ApGet expression plasmid constructs designed for transient expression of ApGet variants in mammalian cell systems. [Diagram 25] Schematic diagram of the mutated ApGet expression units incorporated into the plasmid construct shown in Figure 24 together with a high copy number bacterial origin of replication and an antibiotic resistance gene. Each expression unit provided a detection tag (FLAG® epitope tag) preceded by a nuclear localization signal linked to the RSBD at the N-terminus of ApGet. [Figure 26]Layout of the ApGet targeting region in exon 4 of the PD-L1 gene. The TE-binding region, designated the PD-L1 target site, is flanked by PDS sites. The orientation of the TE-binding region indicates the forward and reverse complementary strand binding sites for ApGet. [Figure 27] Figure 10: Reduction of PD-L1 expression in cell cultures transfected with ApGet. Quantification of PD-L1 expression in cell samples from respective images (ImageJ). Error bars are standard deviation of signal intensity of individual cells in the sample. [Figure 28] Provided is a table presenting (a) a depiction of the genomic target region and donor template structure discussed in Example 6(ii) for investigating ApGet-mediated repair of a target region in the PD-L1 gene and (b) the editing components used and sample results for PCR confirmation of HDR events. [Figure 29] FIG. 1 shows the genomic reference sequence with the layout of primers used to evaluate targeting of the TSKU gene in HEK293 cells using the ApGet expression construct reported in Example (6)(iii). [Diagram 30] TIDER analysis from the study reported in Example (6)(iii) using Sanger sequencing comparing the levels of substitution from the wild-type sequence between untreated and ApGet transfected samples. [Diagram 31] FIG. 1 shows the configuration of recombinant ApGet protein purified from the E. coli expression system. The configuration details the various purification, solubility and cleavage tags discussed in Example 7. [Diagram 32] FIG. 1 is a diagram of the expression cassette used to express the ApGeti-VP64 construct discussed in Example 8 for transcriptional activation of the ASCL-1 gene in HEK293T cells. [Diagram 33] Results of transcriptional activation of the ASCL-1 gene in HEK293T cells using ApGeti-VP64 constructs expressed in conjunction with four NACs targeting different promoter locations and comparison with dCas9-VP64 fusion constructs. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0042] Detailed Description The targeting nucleic acid for providing the genome editing tool of the present invention will generally be RNA. Targeting element (TE) is complementary to the sequence on target DNA. By "complementary", it is understood that TE hybridizes with target sequence to achieve its targeting purpose, and generally, TE will be completely complementary to target sequence.
[0043] Typically, the TE is 15-25 nucleotides, e.g., 15-20 or 21 nucleotides, preferably 18-20 or 21 nucleotides. However, longer or shorter TEs, e.g., TEs of about 10-35 nucleotides, may be feasible in some circumstances. However, to target a specific sequence in a genome with a single-stranded nucleic acid, e.g., single-stranded RNA, it is necessary that the selected target sequence is unwound and therefore accessible. The inventors hypothesized that this could be aided by providing a small peptide that can bind to a short, proximal predefined dsDNA sequence (PDS), e.g., 3-6 nucleotides, in this example. Without wishing to be bound by theory, it was reasoned that the high affinity binding interaction between the peptide (DNA binding domain or DBD) and the PDS may destabilize Watson-Crick base pairing in the adjacent region of the dsDNA helix, which in turn promotes hybridization of the TE to its complementary target region. When the TE is an RNA sequence, it displaces an R-loop, i.e., a sequence of single-stranded DNA, while its complement forms an RNA / DNA hybrid helix with the invading RNA strand, otherwise forming a region of double-stranded DNA. As mentioned above, this type of DNA-binding domain is considered an essential element of the genome editing tool of the present invention, helping to optimize the efficiency of the required genome modification.
[0044] Providing such DBDs to a given sequence can be accomplished by a variety of methods. Such methods may include the use of computational methods for ligand design. However, a preferred method exemplified herein is the use of a phage display peptide library to select displayed peptides with binding affinity to a selected target dsDNA sequence. The use of such a biopanning method involves the use of approximately 1×10 β-peptides fused to the N-terminus of the phage minor coat protein III by a GGS linker. 9 This is exemplified herein by the selection of DBD candidates using a commercially available M13 phage display library containing a unique 12 amino acid long peptide of the formula (I). See Example 1. By this means, a preferred 15-mer peptide with an incorporated C-terminal GGS was selected (see SEQ ID NO: 4 above), which binds to the 6-mer dsDNA sequence 5'GAGGTC3', representing a possible choice as the PDS. This represents only one DBD-PDS combination that may be used. It will be recognized that many alternative DBD peptides may be found that bind to other desired short dsDNA sequences.
[0045] As mentioned above, the DBD may be, for example, a polypeptide sequence of 70-75 amino acid residues or less, such as 30-mer or less, 25-mer or less, 20-mer or less, preferably 15-mer or less. Given the desirability of minimizing the size of the total polynucleotide components for the genome editing tools of the present invention, however, it is supported to provide a DBD of 15 amino acids or less, for example, about 12-15 amino acids (possibly including a GGS element at the C-terminus). Furthermore, it is recognized that the DBD may not necessarily require 6bp recognition. It could be less than 6bp, for example, 5'NNGG3' or 5'GG3'.
[0046] As mentioned above, the selected DBD may have binding affinity to more than one predefined sequence, for example, more than one 6-mer dsDNA sequence. This may be favored to provide flexibility for use at different DNA sites. Thus, the exemplified DBD of SEQ ID NO: 4 has been shown to be effective with the same targeting nucleic acid when some 6-mer dsDNA sequences other than 5'GAGGTC3' are presented as PDS; see Example 5 and Figure 23b. For example, it may be particularly favored to pair with PDS of sequence 5'TTGGTA3', 5'AAAAAA3' or 5'AAAGTC3'.
[0047] The DBD may be selected, for example, to target portions of the genome that are inaccessible to a CRISPR-Cas system, e.g., a Cas9 CRISPR system, due to PAM restriction. The DBD may be specifically designed, for example, to bind to AT-rich regions of the genome rather than regions that provide a Cas9 PAM in a CGG region, i.e., 5'-NGG-3'.
[0048] The targeting nucleic acid component must be bound by a module of the modular polypeptide component. Preferably, this will be a terminal module. This interaction can be known by various means by linking the nucleic acid with a recognition domain in the polypeptide structure. For example, the targeting nucleic acid component may present a non-nucleic acid label, which is a member of a specific binding pair, as its recognition element for interaction with the modular polypeptide component. This binds to its binding partner presented by the modular polypeptide component. By way of example, the specific binding pair may be biotin-streptavidin or the nucleic acid may be bound to an antigen epitope, in which case the recognition domain of the modular polypeptide may comprise an ScFV for binding to the epitope. However, more preferably, when the targeting nucleic acid component is an RNA, as mentioned above, the recognition element provided will be an RNA motif (alternatively called an RNA scaffold) that binds to a module of the polypeptide component, referred to herein as an RNA scaffold binding domain (RSBD). Many such naturally occurring RNA scaffold-RSBD interactions are known and are often referred to in the literature as RNA aptamer-binding domain interactions for linking RNA sequences to protein domains. Examples of such binding complexes include (i) the 19 nt box B RNA motif and its cognate 22 amino acid RNA binding domain of lambda phage antiterminator N protein (lambda N22 peptide; see SEQ ID NO:2 above), (ii) the MS2 phage operator stem-loop and its MS2 coat protein (MCP) binding domain sequence, (iii) the PP7 phage operator stem-loop and its PP7 phage coat protein (PCP) binding domain sequence, (iv) the phage Com RNA scaffold sequence and its phage Com binding polypeptide, (v) the telomerase RNA Ku binding motif and its Ku protein or RNA binding section, and (vi) the telomerase RNA Sm7 binding motif and its Sm7 protein or RNA binding section.However, non-natural RNA aptamers or scaffolds may be preferred as recognition elements for targeting RNAs that bind to the synthetic binding domains of the modular polypeptide components. Thus, any of the naturally occurring RNA scaffold-RSBD pairs described above may be replaced with variants in which either or both members of the pair are mutated sequences, provided that the required binding affinity is maintained. Phage display peptide libraries may again be used to achieve alternative panning of peptide-RNA scaffold binding pairs or peptide libraries for the identification of peptide binding activity suitable for RNA scaffolds, e.g., RNA structures containing one or more hairpins. The length of the RNA scaffold should preferably be between about 15 and 25 nucleotides, e.g., 19-20 nucleotides as exemplified by the lambda box B sequence. The RNA scaffold-RSBD pair has multiple contact points with a view to providing high affinity interactions and reducing off-target non-specific interactions of the scaffold sequence with any host genome sequences.
[0049] The targeting nucleic acid or its RNA scaffold element and / or connector sequence may choose to incorporate one or more modified nucleotides and / or one or more modified internucleotide linkages. Modified nucleotides for this purpose may include, for example, 2'-O-methyl analogs, 2'-fluoro analogs or 2'deoxy analogs. One may envisage the use of locked nucleic acid (LNA) monomers in which a 2' hydroxyl group is linked to the 4' carbon atom of the sugar ring forming a 2'C,4'-C-oxymethylene bond, thereby providing a bicyclic sugar moiety. Modified bases such as 2-aminopurine, 5-bromouridine, 5-methylcytidine and 5-methoxyuridine may be used. One or more internucleotide linkages may be, for example, phosphorothioate linkages. Such methods for modifying RNA oligonucleotides to aid in stability and required activity in a cellular environment are well known, and any such methods may be utilized in designing targeting nucleic acids for use in accordance with the present invention. It will be appreciated that such RNA oligonucleotide modifications may be envisioned where chemical synthesis of targeting nucleic acids can be performed for RNP delivery of nucleoprotein complexes.
[0050] As a specific example of a modified wild-type phage aptamer that retains binding affinity for its cognate binding protein and can be used in the targeting nucleic acid of the present invention, reference may be made to the MS2 stem-loop mutants taught as RNA stem-loop motifs in WO 2022 / 011232 (Horizon Discovery Limited and Dharmacon, Inc; published 13 January 2022), such as the MS2 motif mutant in which A is changed to 2'-deoxy-2-aminopurine or 2'-ribose-2-aminopurine at position 10 from the 5' end of the stem-loop (F-5 mutant substitution a shown in Figure 2B of the same published WO application).
[0051] In some cases, it may be desirable to provide two or more RNA scaffolds (same or different) on the targeting RNA, for example, two different RNA scaffolds in tandem separated by a connector sequence, to enhance the efficiency of the system or to facilitate interaction with more than one protein via the RSBD. In this way, a single targeting RNA may interact with more than one modular polypeptide component, which may be the same or different, for the nucleoprotein complex of the present invention. Such targeting RNAs with more than one RNA scaffold may also be used to recruit another effector component, for example, a base editor, that is joined to the appropriate RSBD.
[0052] For the purposes of the initial proof of concept studies reported herein, as previously described, the box B RNA motif is expressed in a manner similar to that of its cognate binding peptide, lambda N22.
[0053] [ka] The lambdaN peptide was chosen as a targeting RNA scaffold with a modular polypeptide component that provides a targeting RNA scaffold for targeting RNA molecules. (See Baron-Benhamou et al. (2004) Methods Mol. Biol. 257, 135-154, 'Using the lambdaN peptide to tether proteins to RNAs')
[0054] Whatever the selected recognition element of the targeting nucleic acid, a short connector sequence, for example, 2, 3, 4, 5 or 6 nucleotides, is generally, but not necessarily, provided between the targeting element and the recognition element to facilitate accurate binding to both its complementary sequence and the binding domain, respectively. The RNA scaffold can be placed at the 5' or 3' end of the nucleic acid building block. Thus, in the examples given below, the RNA nucleic acid building block was fused to the 5' end of the following sequence, in which the short connector AATTT is fused to the BoxB RNA scaffold sequence shown in bold:
[0055] [ka]
[0056] It is recognized that the connectors in this sequence may be replaced by alternative sequences, provided that the alternative sequences allow correct function of both the attached targeting element (TE) and RNA scaffold (RS). Moreover, the TE or RS may be the 5' terminal sequence.
[0057] The TE in the targeting nucleic acid for the gene editing tool of the invention can be designed to bind to either strand of dsDNA. See Figure 6.
[0058] It is recognized that the complete modular polypeptide building block must also be designed such that each module has the ability to perform its desired purpose, and thus may require one or two inter-module linkers. Thus, a linker--one or both of the different linker peptides shown as linker 1 and 2 in any of Figures 1, 2 and 6--can be provided between the selected DBD and one or both of the effector building blocks and the recognition domain to link the targeting nucleic acid. Linkers suitable for this purpose that do not impart secondary structure or undesirable domain interactions are well known; see, for example, Chen et al. (2013) Adv. Drug Deliv. Rev. 65, 1357-1369, 'Fusion Protein Linkers: Property, Design and Functionality'. In general, flexible linkers composed of, for example, a stretch of glycine and serine residues or XTEN peptide linkers are preferred (Komor et al, "Programmable Editing of a Target Base in Genomic DNA without Double-Stranded DNA Cleavage." Nature 533 (7603) p.420-424). Suitable linkers may be 2 to 15 or more amino acids long, for example (Gly)n or (GGGGS)n. See also linker 1a used in the study in Example 8. In some cases, flexible linkers much longer than 15 amino acids may be preferred, for example up to 35 amino acids -(GGGGS)7, as exemplified by the linker in ApGet of SEQ ID NO:1 linking the RSBD and DBD. See also the shortened version of this linker (GGGGs)6 used in Example 8. The length of any such linkers may preferably be minimized, with a view to desireing the full length of the modular polypeptide components to be as low as possible, e.g., in the case of artificial nucleases described herein, e.g., preferably 220-250 amino acids or less.However, longer modular polypeptide components, eg, 300 amino acids in length or longer, may be required in some cases, for example if different effector components are used.
[0059] The effector components set forth above may be any of a diverse range of moieties for use in modifying DNA (altering the structure or regulation) including: (i) endonucleases for generating double stranded breaks, e.g., a Fok1 nuclease domain; (ii) nickases; (iii) transcriptional activators such as VP64; (iv) transcriptional repressors; (v) epigenetic modulator enzymes; (vii) recombinases; (viii) transposases; (ix) integrases; and (x) nucleobase modifying enzyme constructs comprising nickases, e.g., the artificial nickases described herein in conjunction with base editors.
[0060] As mentioned above, it may be preferred that the effector component does not have DNA binding ability, but it is not excluded that the effector component has some DNA binding ability, provided that it does not prevent the site-specific modification directed by the targeting element and DBD.For example, it is known to fuse Cas9 protein with transposase or HIV integrase for site-specific insertion of exogenous nucleic acid into genome.Transposase or integrase may also be used as the effector component in the modular polypeptide component for genome editing described now, and the site-specific insertion of co-supplied exogenous nucleic acid is directed by hybridization of the targeting element to its complementary sequence, which is promoted by binding of DBD.
[0061] However, as mentioned above, the provision of an artificial nickase, as now further described below, is particularly preferred. The effector component may, for example, be a fusion construct in which a nickase, possibly including an artificial nuclease as described herein, is provided together with a base editor or a reverse transcriptase.
[0062] Artificial nuclease comprising multimers of linked self-assembling peptides that form a beta-sheet structure, (i) the self-assembling peptides of said multimers each exhibit an alternating pattern of hydrophobic and hydrophilic residues, and all or at least a portion of the hydrophilic residues of each such peptide represent a catalytic triad of amino acids or more than three catalytic amino acids, such that said multimers are capable of cleaving dsDNA at both strands or at a single strand, i.e., exhibit nickase activity; (ii) the self-assembling peptides are joined by a linker to limit the number of multimers to, for example, 20 or less, for example, 4 to 15 or 16, preferably 10 or less, more preferably 5 to 10; (iii) Artificial nucleases are now provided in which the N- and C-terminal flanking sequences are provided linked to multimers to prevent or limit aggregation between multimers of the nuclease protein, which have increased positive charge compared to the peptide units of the multimer due to the inclusion of one or more positively charged residues, e.g., by inclusion of one or more lysine or arginine residues. Other means of capping the multimer at the N- and / or C-termini may be used to aid in performance, e.g., by aiding solubility and / or stability.
[0063] By beta-sheet structure is understood, preferably, as an amyloid-like beta-sheet structure. Self-assembling peptides can be selected and linked, for example, with a view to obtaining antiparallel beta-sheets. The linkers between each such peptide can have the same or different lengths. Thus, the linker length can be 2-25 amino acids, for example, 4-, 5-, 6-, 7-, 8-, 9-, 10-, 11- or 12-mer linkers can be selected. Short linkers, for example 4-mers, suitable for providing beta-turns or beta-turn-like loops can be designed based on knowledge of the amino acid composition of such turns in natural proteins, as further discussed and exemplified below.
[0064] Suitable self-assembling peptides (or, as in single-folded structures, called folded peptides) form structures that can be identified by Thioflavin T (Tht) binding. A fluorescence emission assay of Thioflavin T binding is commonly used to detect amyloid fibrils. Upon binding to amyloid fibrils, ThT emits a strong fluorescent signal at approximately 482-485 nm when excited at 450 nm (Xue et al. (2017) Royal . Soc. Open Sci. 4: 160696, 'Thioflavin T as an amyloid dye: fibril quantification, optimal concentration and effect on aggregation').
[0065] The C-terminal flanking sequence may also preferably provide a flexible tail sequence at its C-terminus to aid in solubility, such as the 4-mer NQGS [SEQ ID NO: 10] or a longer sequence.
[0066] For production as a separate enzyme entity or as part of a fusion construct, such polypeptides comprising artificial nucleases may desirably be produced with an N-terminal tag to aid purification and / or solubility, for example a His tag to aid purification, for example a hexaHis tag, or a tag comprising one or more His tags to aid solubility and / or a sequence such as maltose binding protein (MBP). Preferably, such a tag will be a cleavable tag, for example selected from SUMO or TEV tags. Thus, a suitable tag may comprise 6x His linked to MBP, together with an additional C-terminal 6x His tag separated by a short spacer from the TEV protease cleavage site; see FIG. 10.
[0067] The artificial nickases disclosed herein are preferred for inclusion in the synthetic genome editing tools of the invention as effector components, however, it will be appreciated that the invention extends more generally to polypeptides that comprise or consist of the artificial nucleases presently taught.
[0068] It is recognized that the multimeric self-assembling peptides may be identical, as in the artificial nuclease of SEQ ID NO: 6 discussed above, but one may choose to use two or more non-identical peptides of the same length, e.g., 7-mer. Preferably, the alternating hydrophobic residues of the self-assembling peptide may be leucine and / or isoleucine, most preferably isoleucine, beginning with the N-terminal residue. However, the hydrophobic amino acid residue may also be selected from any hydrophobic amino acid residue, including, e.g., valine. The length may be, e.g., 6-mer to 15-mer, preferably 7-mer to 11-mer, e.g., 7-mer or 9-mer, with 7-mers, preferably 7-mers with alternating isoleucine residues beginning with the N-terminal residue, being preferred.
[0069] It will be appreciated that the peptides of the multimer may be linked consecutively in the N-terminal to C-terminal direction or the alternating peptides may be reversed in whole or in part, and for convenience will preferably be identical. Thus, a multimer formed of a 7-mer self-assembling peptide is: (i) Non-inverted multimer: (H-C1-H-C2-H-C3-H-(X) n -H-C1-H-C2-H-C3-H) n or (ii) Inverted multimer: (H-C1-H-C2-H-C3-H-(X) n -H-C3-H-C2-H-C1-H) n It may be expressed as: H=hydrophobic residue, preferably leucine or isoleucine, most preferably isoleucine; C1, C2 and C3 are different catalytic amino acids, X represents an amino acid linker that may vary, for example, any of 4, 5, 6, 7, 8, 9, 10, 11 or 12 mers, preferably 4 mers designed to provide a beta turn or beta turn-like loop, such as NDGG, DSSG, SSGS, NNGN, GSDG, GNSG, DNGG, DSDG, SSSG, GSEG, GDSG (SEQ ID NOs: 11-21) and other linkers consisting of combinations of amino acids selected from glycine (G), polar amino acids serine (S), asparagine (N), glutamine (Q), and threonine (T) and charged amino acids glutamic acid (E) and aspartic acid (D). Such linkers may be used with any suitable self-assembling peptide to control multimer size, as well as with a view to aiding in the stabilization and solubility of the desired beta structure design. The possible options may be readily discerned by one skilled in the art of protein structure analysis and design considering the examples provided.
[0070] Preferably, all self-assembling peptides are linked in the N-to-C-terminal direction. In this way, the assembly of consecutive monomer units can be provided in an antiparallel beta-sheet structure, as discussed further below. See Figures 8 and 9.
[0071] C1, C2 and C3 may correspond to a trio of different amino acids of a naturally occurring nuclease known to be capable of nicking a single strand of dsDNA, e.g., the known catalytic triad of such a nuclease. Preferably, the selected catalytic amino acids may correspond to the catalytic amino acids of the RuvC nuclease domain present in endonucleases, e.g., Cas9 nuclease, namely, glutamic acid (E), aspartic acid (D) and histidine (H).
[0072] The RuvC nuclease domain present in Cas9 nuclease was reported to have four catalytic amino acids, His983, Asp986, Asp10 and Glu762. It was revealed that mutation of any of these catalytic amino acids results in loss of function (Nishimasu et al. (2014) Cell 156, 935-949). However, in the context of the above multimers, it was revealed that it is possible to use identical peptide units with alternating hydrophobic residues for simplification, e.g., only a single His, Asp and Glu, so that the peptide units maintain the tendency to form stable β-sheet like supramolecular structures and all the necessary catalytic amino acid residues are in close proximity for a functioning nickase.
[0073] Thus, by way of example, a preferred 7 amino acid peptide for the above linkage is IEIDIHI. It is contemplated that this peptide can also be made of any different length between 5 and 11 amino acids. In the artificial nucleases of the invention, such peptides may optionally be fused to a linker sequence providing the aforementioned aspartic acid (D residue) in the multimeric structure. The D residue may be incorporated into the beta-turn design with a view to potentially increasing the solubility of the complete nuclease.
[0074] As mentioned above, multimers of the same self-assembling peptide are flanked by N- and C-terminal sequences with increased positive charge to reduce aggregation tendency. This may be by substitution of at least one hydrophobic residue in an otherwise identical peptide to the self-assembling peptide for the multimer with a positively charged residue, for example, preferably by substitution of the hydrophobic residue with lysine. For example, if the multimer is formed of 7-mer units of alternating hydrophobic and catalytic residues starting with leucine or isoleucine, preferably, for example, 7-mer units of IEIDIHI, such C- and N-terminal units may have the same sequence except for a positively charged amino acid, for example, lysine (K) residue substitution for isoleucine at the 3rd and / or 5th position from the N-terminus. Thus, the N-terminal flanking sequence may be IEKDIHI (SEQ ID NO: 22) or IEIDKHI (SEQ ID NO: 23) fused to a 4-mer linking sequence for binding to the first 7-mer peptide unit of the multimer of the same peptide unit. The C-terminal flanking sequence may be selected from the same sequence with an additional C-terminal tail sequence. Thus, the N-terminal flanking sequence can be IEKDIHI and the C-terminal flanking sequence can be IEIDKHI, or vice versa. One or both lysine residues can be alternating positively charged residues, e.g., arginine. In this way, there is repulsion of nuclease assembly, for example, when such flanking units are linked to a multimer of 10 identical peptides. As mentioned above, the C-terminal tail sequence can be selected to aid in solubility, e.g., NQGS.
[0075] An example of such a functional nickase formed entirely from linked synthetic peptides is as follows:
[0076] [ka] Ten identical self-assembling 7-mer peptides (peptide IbD with a hydrophobic isoleucine residue) are shown in bold; Provided is the amino acid sequence of SEQ ID NO:6 which can alternatively be represented as individual peptide units, all in the N-to-C-terminal direction.
[0077] All these peptides are linked from the N- to C-terminus direction by a 4-mer linker intended to provide a beta turn so that successive peptide monomers are arranged in an antiparallel beta sheet conformation as follows:
[0078] [ka] It will be appreciated that the orientation is opposite.
[0079] It will be appreciated that one or more of the beta turns may be varied and / or the C-terminal NQGS sequence and still maintain the desired nuclease activity.
[0080] Substitution of artificial nuclease of SEQ ID NO:1 It is recognized that the artificial nuclease provided in SEQ ID NO: 1 can be replaced by any alternative effector component desired for modification of the target dsDNA site, with the maintenance or modification of the linker to the DBD. By way of example, such a functional synthetic genome editing tool now taught has the artificial nuclease of SEQ ID NO: 1 replaced by a Fok1 nuclease domain (referred to herein as the ApGet-Fok1 construct). It is recognized that any ApGet-i containing a nucleic acid recognition module linked to a DBD can be further linked to a Fok1 nuclease domain. Such modular polypeptide constructs can be used, for example, with a pair of targeting nucleic acids with different TEs to provide a functional Fok1 nuclease capable of cleaving dsDNA at the target site. One arrangement for this is illustrated in FIG. 18. See also the illustration below, which uses the "editor" plasmid illustrated in FIG. 14 for expression of the ApGet-Fok1 construct that generates a double-stranded break in the detector plasmid. It will be appreciated that the DBD and targeting elements can be altered to generate double-stranded breaks at different dsDNA target sites, however, with a view to minimizing the size for expression of the complete modular polypeptide components of the gene editing tools now taught, it is preferred to replace the Fok1 nuclease domain with an artificial nickase of the invention.
[0081] As a further illustration, the artificial nuclease provided in SEQ ID NO: 1 can be replaced by a transcriptional activator, e.g., VP64, and the resulting modular polypeptide can be combined with one or more targeting nucleic acids designed to target the promoter location, thereby activating or enhancing the transcription of one or more coding sequences operatively linked to the promoter; see Example 8. Such effector component replacement can involve modification of one or more of the RSBD, DBD, the linker between the RSBD and DBD (linker 2, L2 shown in FIG. 1) and the linker between the effector component and DBD (linker L1, L1 shown in FIG. 1), provided that the necessary functions of the RSBD, DBD and effector component are maintained. Thus, for example, as illustrated by Example 8, the artificial nuclease of SEQ ID NO: 1 can be replaced by a VP64 polypeptide, with shortening of linker 2 (see linker 2b in SEQ ID NO: 70) or modification of linker 1 (see linker 1a in SEQ ID NO: 69).
[0082] Further aspects of the invention It is recognized that the artificial nickase of the present invention has applications beyond its inclusion in the fully synthetic gene editing tool discussed above.Artificial nickase can be used as part of another entity for DNA modification, such as fused with dCas9, possibly also linked to further components for DNA modification, such as base editor or reverse transcriptase.Artificial nickase can be expressed from polynucleotide, such as expression vector, as part of fusion protein, such as modular polypeptide component of DNA modification tool for use with targeting nucleic acid, or as separate protein.
[0083] Polynucleotides capable of expressing the polypeptides of the invention, for example comprising or consisting of the artificial nucleases or nickases of the invention, and host cells transformed with such polynucleotides also form aspects of the invention.
[0084] Also provided are polynucleotides encoding the artificial nuclease or nickase proteins of the invention with such tag sequences, which are N-terminal tag sequences for aiding purification and solubility as discussed above, e.g., a cleavable tag sequence that provides both a His tag and a tag for aiding solubility, e.g., a maltose binding protein (MBP) sequence. Such polynucleotides may be expression cassettes for the production of nuclease or nickase proteins in in vitro transcription and translation systems or expression vectors for the expression of such nuclease proteins in host cells, e.g., bacterial cells such as E. coli cells or yeast cells. Such transformed host cells represent further aspects of the invention.
[0085] A method for producing a polypeptide comprising an artificial nuclease or nickase of the invention, comprising: (i) culturing the above host cells or using an in vitro transcription-translation system in which the polypeptide is produced together with an N-terminal tag sequence to aid in purification and solubility, preferably together with a cleavable tag, such as a cleavable His-MBP-His fusion tag; (ii) isolating the polypeptide using an N-terminal tag, and said N-terminal tag is cleavable; (iii) cleaving the tag from the polypeptide; and (iv) separating the polypeptide from the tag. Further provided is a method comprising:
[0086] As mentioned above, it is particularly preferred that the artificial nuclease of the present invention, preferably the artificial nickase of the present invention, is provided as part of a modular polypeptide of the artificial gene editing system of the present invention. In this case, the artificial nuclease or nickase is generally joined to the DBD by a linker sequence so that the DBD can bind to the selected PDS and the nuclease can function at the required target site. For the expression of such a modular polypeptide, an expression cassette, e.g., a vector, is provided that codes for the entire modular polypeptide. In some cases, an N-terminal tag, e.g., a cleavable His tag, can be provided (see FIG. 16).
[0087] A preferred method for recombinant production of ApGet-i or complete ApGet modular polypeptides comprising the artificial nucleases of the present invention is now taught herein. Thus, such modular polypeptide components, or any modular polypeptide components, for use in the nucleoprotein complexes of the present invention for target sequence modification may be initially expressed in a host cell, such as, for example, E. coli, with multiple component tag sequences to aid in solubility, purification and / or detection, including protease cleavage sites that can remove undesired tag sequences prior to use. Such multiple component tag sequences may be linked to an NLS, whereby upon protease cleavage, the NLS is retained at the N- or C-terminus, possibly with additional sequences that do not interfere with the required target modification, such as terminal sequences including detection sequences. Such a multiple component sequence tag may desirably provide all of (i) at least one His tag, e.g., at least one hexaHis tag, (ii) a polypeptide sequence that aids in solubility, e.g., MBP or small ubiquitin-like modifier (SUMO), (iii) a protease cleavage site, and (iv) a detection sequence, e.g., an epitope tag. As an example of such a preferred multiple component tag, it has been found to be beneficial to express ApGet or ApGet-i in E. coli with an N-terminal extension that provides, in the N-to-C-terminal direction, all of: (i) a hexaHis tag, (ii) a solubility tag of MBP or SUMO, (iii) a linker that provides a spacer, (iv) a TEV protease cleavage site, and (v) an epitope tag, such as a FLAG® epitope grafted to an NLS. Upon protease cleavage, the final modular polypeptide for targeted modification retains both the NLS and the epitope tag together with a short N-terminal leader sequence, e.g., a 4-mer N-terminal leader sequence such as GWGS (see Figure 31). The linker between the solubility tag and the TEV protease cleavage site may desirably be an asparagine-rich linker rather than a linker that increases the requirement for glycine and serine.It has been found that the expression of ApGet further benefits from the incorporation of a Strep-tag® at the C-terminus to aid in separation from prematurely terminated products. However, it is recognized that a variety of other N- and C-terminal tags may be used when initially expressing the modular polypeptide components of the invention to aid in solubility and / or purification and / or detection and / or membrane penetration. For example, as described above, the modular polypeptide components may in some instances be expressed including cell-penetrating peptide sequences.
[0088] Uses of the Synthetic Genetic Engineering System of the Invention As mentioned above, the nucleoprotein complexes of the invention can be delivered to the host cell as active complexes, for example, by electroporation, or together with one or both of the polypeptide component expressed by the polynucleotide and the targeting nucleic acid component.
[0089] Delivery of ApGet components into cells The required nucleic acid components can be delivered to cells by any of a variety of suitable methods without limitation.Many such methods are known.Thus, the synthetic RNA molecule can be directly introduced into the cells of interest by electroporation, nucleofection, transfection, by nanoparticles, by virus-mediated RNA delivery, by non-virus-mediated delivery, by extracellular vesicles (e.g., exosomes and microvesicles), by eukaryotic cell transfer (e.g., by recombinant yeast) and other methods that can package nucleic acid molecules and provide delivery to target viable cells.Other methods for the introduction of RNA molecules include non-integrative transient transfer of DNA polynucleotides, in which the relevant sequence is transcribed into cells. This includes, without limitation, by using DNA-only vehicles (e.g., plasmids, minicircles, minivectors, ministrings, protelomerase-generated DNA molecules (e.g., Doggybones, artificial chromosomes (e.g., HACs), cosmids) or using vehicles such as nanoparticles, extracellular vesicles (e.g., exosomes and microvesicles), by eukaryotic cell transfer (e.g., by recombinant yeast), transient viral transfer with AAV, non-integrating viral particles (e.g., lentivirus and retrovirus based systems), cell penetrating peptides and other techniques that can mediate the introduction of DNA into cells without direct integration into the genomic landscape. Another method for the introduction of RNA components includes the use of integrative gene transfer techniques for stable introduction of the machinery for RNA transcription into the genome of the target cell. This can be controlled by constitutive or promoter-inducible systems that attenuate RNA expression. Such techniques for stable gene transfer include integrative viral particles (e.g., lentivirus, adenovirus and retrovirus based systems), transposase-mediated transfer (e.g., Sleeping Beauty). Examples of suitable techniques include, but are not limited to, siRNA, siRNA, siRNAs ...
[0090] The delivery of protein or peptide-acting components such as artificial nickases of the present system can be carried out by the same technology, but in some circumstances, there is an advantage to mediate the delivery by different methods. Such applicable methods are listed below, but are not limited to them. First, the direct introduction of protein molecules into the cells of interest can be by electroporation, nucleofection, transfection, by nanoparticles, by virus-mediated packaged delivery, by extracellular vesicles (e.g., exosomes and microvesicles), by eukaryotic cell transfer (e.g., by recombinant yeast) and other methods that can package macromolecules for delivery to target living cells without integration into the genome landscape. Other methods for the introduction of protein molecules include non-integrative transient transfer of DNA polynucleotides that contain sequences related to transcription and translation to provide the required intracellular protein molecules. Again, this includes, without limitation, the possible use of DNA-only vehicles (e.g., plasmids, minicircles, minivectors, ministrings, protelomerase-generated DNA molecules (e.g., Doggy Bones), artificial chromosomes (e.g., HACs), cosmids) or DNA delivery vehicles such as nanoparticles, extracellular vesicles (e.g., exosomes and microvesicles), by eukaryotic cell transfer (e.g., by recombinant yeast), transient viral transfer by AAV, non-integrating viral particles (e.g., lentivirus and retrovirus-based systems) and other techniques that can mediate the introduction of DNA into cells without direct integration into the genome landscape. Another method for the introduction of protein component(s) includes the use of integrative gene transfer techniques for stable introduction of the machinery for transcription and translation into the genome of the target cell. Many such methods are well known in the genome modification field. Control can again be by constitutive or inducible promoter systems. The design may allow the system to be removed once utility has been fulfilled (e.g., introducing a Cre-Lox recombination system).Such techniques for stable gene transfer include, but are not limited to, integrating viral particles (e.g., lentivirus, adenovirus and retrovirus based systems), transposase-mediated transfer (e.g., Sleeping Beauty and PiggyBac), and other techniques that facilitate integration of target DNA into the cells of interest.
[0091] Expression system As mentioned above, it may be desirable to express one or more of any protein and / or RNA components from the encoding polynucleotide. This can be achieved in a variety of ways.
[0092] For example, one or more intermediate vectors may be used for introduction into a prokaryotic or eukaryotic cell and suitable for replication and / or transcription to express the required component(s). The expression vector may be used for administration to a selected host cell, for example, a plant cell, an animal cell, for example, an avian, mammalian or human cell, a fungal cell, a bacterial cell, or a protozoan cell. Thus, the present invention provides a nucleic acid encoding any of the above-mentioned RNA components or proteins of the present invention. Preferably, the nucleic acid is isolated and / or purified.
[0093] Examples of expression constructs useful in the application of the nucleoprotein complex of the present invention include vectors, such as plasmids or viral vectors, into which the nucleic acid sequence of the present invention is inserted in a forward or reverse orientation. In a preferred embodiment, the construct further comprises a regulatory sequence, including a promoter, operably linked to the sequence. Many suitable vectors and promoters are known to those of skill in the art and are commercially available.
[0094] Examples of expression vectors useful in applications of the nucleoprotein complexes of the present invention include chromosomal, non-chromosomal and synthetic DNA sequences, bacterial plasmids, phage DNA, baculovirus, yeast plasmids, vectors derived from combinations of plasmids and phage DNA, viral DNA such as vaccinia, adenovirus, fowlpox virus, and pseudorabies.
[0095] The vector may include sequences suitable for amplifying expression. In addition, the expression vector preferably contains one or more selectable marker genes that provide a phenotypic trait for selection of transformed host cells, such as dihydrofolate reductase or neomycin resistance in eukaryotic cell culture, or tetracycline or ampicillin resistance in E. coli.
[0096] The vector(s) for use in the application of the nucleoprotein complex of the invention, for example, expressing the modular polypeptide component(s) and / or one or more targeting RNAs in a cell, or in the production of the protein components of the invention, can be designed with appropriate control sequences for such expression or production in a selected host cell. Examples of suitable expression hosts include bacterial cells (e.g., E. coli, Streptomyces, Salmonella typhimurium), fungal cells (yeast), insect cells (e.g., Drosophila and Spodoptera frugiperda (Sf9)), animal cells (e.g., CHO, COS, and HEK293), adenovirus, and plant cells. The selection of a suitable host for the production of the protein of the invention, e.g., an artificial nuclease, is within the capabilities of one of ordinary skill in the art.
[0097] Applicable Thus, in a further aspect, a nucleic acid or combination of nucleic acids is provided for the provision of one or more nucleoprotein complexes of the present invention in a host cell.However, direct delivery of the active complexes (protein components plus targeting RNA) of the present invention to a host cell is not excluded.Such a host cell may be a prokaryotic or eukaryotic cell.The host cell may be any eukaryotic cell, including, for example, bacterial cells, yeast cells, insect cells, and fungal cells, as well as plant cells and human and non-human animal cells, such as, for example, hamster (for example, CHO cells), monkey (for example, Vero cells), rat, mouse, bird or chicken cells.It is understood that suitable host cells include ex vivo cells, such as ex vivo stem cells, induced pluripotent stem cells (iPSCs) and ex vivo T cells, including engineered CAR-T cells.
[0098] The nucleic acid may preferably be a vector, for example a viral vector, such as a recombinant adeno-associated virus (rAAV) vector, lentivirus, retrovirus, adenovirus, or Sendai virus vector, capable of expressing both at least one modular polypeptide component of the present invention, including the artificial nickase of the present invention, and at least one targeting nucleic acid component. To target more than one DNA site, more than one targeting nucleic acid component may be provided to the host cell. In this case, each targeting nucleic acid component may interact with a single modular polypeptide component, i.e., more than one targeting nucleic acid, for example, a pair, is operatively combined with a single DBD that meets a single PDS. However, each targeting nucleic acid component of a pair may interact with a different modular polypeptide component, each having a DBD that binds to a different PDS. Regardless of which of these delivery modes is selected, the DNA target site for the targeting nucleic acid component may be outside the pair of PDS sites (referred to as a "PDS-in" configuration) or the DNA target site for the targeting nucleic acid component may be inside the pair of PDS sites (referred to as a "PDS-out" configuration). Such configurations are illustrated in Figures 18 and 19, where each targeting nucleic acid targets a different DNA strand of dsDNA.
[0099] It is recognized that the pair of nucleoprotein complexes of the present invention for use as described above may comprise, as an effector component, a Fok1 nuclease domain joined to the DBD by a suitable linker sequence, i.e., represent an ApGet-i construct (herein referred to as ApGet-Fok1 modular polypeptide) joined to the catalytic domain of Fok1 endonuclease by a linker. Such a pair of complexes may act in a PDS-in or PDS-out configuration to generate double-strand breaks in dsDNA, e.g., genomic DNA for gene knockout. See FIG. 18, which illustrates a pair of ApGet-Fok1 modular polypeptides functioning in a PDS-in configuration for this purpose.
[0100] For vector delivery of the pair of nucleoprotein complexes of the present invention that provide artificial nickases to cells to cleave both strands of dsDNA, one or more vectors, preferably a single vector, are particularly preferred. Such delivery may also achieve gene knockout by non-homologous end joining. However, it is recognized that such nucleoprotein complex delivery may involve providing a sequence template for homologous recombination. This sequence template may be provided by a complete separate expression construct or by expression from a separate expression cassette on the same vector. Figure 16 illustrates a single vector construct that can express all of the modular polypeptides that comprise the artificial nuclease, the targeting nucleic acid components for interacting with the modular polypeptide, and the sequence for homologous repair of the enzyme coding sequence by cleavage of the pair of target sites.
[0101] As a further aspect of the present invention, a method is provided for modifying one or more target nucleic acid sequences in a host cell using one or more nucleoprotein complexes of the present invention or one or more nucleic acids for their provision. However, such a method as claimed does not extend to a method for modifying human germline identity or as such a method of medical treatment performed on the human or animal body. As mentioned above, desirably, a single polypeptide component for the provision of one or more nucleoprotein complexes can be provided in the cell by expression from a vector. The same vector can preferably also express the required one or more targeting nucleic acids, preferably one or more targeting RNAs.
[0102] As yet another aspect of the present invention, there is provided a combination of (i) at least one polypeptide component for the genome modification tool of the present invention, or a polynucleotide capable of expressing it, and (ii) one or more targeting nucleic acids that can interact with or link to said polypeptide component(s), or one or more polynucleotides that can express it, for use in the above-mentioned method or for use in therapeutic treatment. Such a combination may be provided in the form of a kit for one or more specific applications, for example, a single polypeptide component or a polynucleotide encoding it may be provided together with one or more targeting nucleic acids, or one or more polynucleotides designed to target one or more desired nucleic acid positions. However, as mentioned above, both the polypeptide component and one or more targeting nucleic acids, preferably targeting RNA, may be preferably expressed from a single vector.
[0103] By way of example, it will be appreciated that the present invention may be useful for editing cells for use in bioproduction, vaccine production, research tools and reagents, the creation of genetically modified animals and plants, and cell therapy. Components of the present invention may be utilized for gene therapy applications.
[0104] The nucleoprotein complex of the present invention may be utilized, for example, to provide cells for use in therapy. One such application is to generate therapeutic cells, such as T cells engineered to express chimeric antigen receptors (CAR-T) or T cell receptors (TCR). CAR-T / TCR cells may be derived from primary T cells or differentiated from stem cells. Suitable stem cells include, but are not limited to, mammalian stem cells, such as human stem cells, including, but not limited to, hematopoietic, neural, embryonic, induced pluripotent stem cells (iPSC), mesenchymal, mesodermal, hepatic, pancreatic, muscle and retinal stem cells. Other stem cells include, but are not limited to, mammalian stem cells, such as mouse stem cells, e.g., mouse embryonic stem cells.
[0105] The nucleoprotein complexes of the present invention may be used to knock out, modify or increase the expression of a single gene or multiple genes in various types of cells or cell lines, including but not limited to cells of mammalian origin. The nucleoprotein complexes of the present invention may be adaptable for multiple gene modifications, including genetically modifying multiple genes or multiple targets within the same gene, as known in the art. The techniques may be used for many applications, including but not limited to knocking out genes to prevent graft-versus-host disease by making non-host cells non-immunogenic to the host or to prevent host-versus-graft disease by making non-host cells resistant to attack by the host. These approaches are also relevant to creating allogeneic (off-the-shelf) or autologous (patient-specific) cell-based therapeutics. Target genes may include but are not limited to T cell receptors, major histocompatibility antigen (MHC class I and class II) genes, including B2M, and genes involved in the innate immune response.
[0106] As stated above and by way of summary, the DNA modification system now taught comprises: 1. Introduction of Ribonucleoprotein Particles 2. mRNA 3. Plasmids expressing the targeting RNA components and the modular polypeptide components 4. Virus-like particles containing any of the above items 1, 2 or 3. 5. Lipid nanoparticles containing any of the above 1, 2 or 3. 6. Exosomes containing any of the above items 1, 2 or 3. 7. Liposomes containing any of the above mentioned in 1, 2 or 3 8. Viral vectors expressing targeting RNA components and modular polypeptide components, e.g., rAAV vectors It will be appreciated that the polypeptide may be delivered into a cell in a variety of ways well known to those of skill in the art for cell transformation, including by immunohistochemistry.
[0107] As mentioned above, it is envisioned that the synthetic genome editing tools now taught, including artificial nucleases formed from linked synthetic peptide units, may advantageously enable genome editing for gene knockout, or genetic modification by homologous recombination, using only a single vector for delivery of the required expression cassettes, e.g., a vector capable of expressing all of the ApGet modular polypeptides and at least one targeting RNA, and possibly also a template sequence for homologous recombination between pairs of cleavage sites.
[0108] The following non-limiting examples are provided to illustrate the invention and further describe the derivation and testing of a nickase in the modular polypeptide of SEQ ID NO:1, preferably as a proof of concept of an artificial nickase that may be used as an element of the synthetic genome editing system of the invention. EXAMPLES
[0109] [Example 1] Selection of DNA-binding domains As mentioned above, a commercially available M13 phage display library was used (New England Biolabs, catalog number E8111L), which contains approximately 1 × 10 9 Two different 6 nt dsDNA baits were used with biotin tags immobilized on streptavidin-coated 96-well plates, as specified in Table 1 below.
[0110] [Table 1]
[0111] An ELISA assay was used to assess binding of peptides from the phage display library to the PDS1 and PDS2 dsDNA baits using horseradish peroxidase (HRP)-labeled anti-M13 phage antibody. After three rounds of biopanning and sequence analysis of bound peptides in the library, two 12-mer peptide sequences were selected, chemically synthesized, and modified as presented in Table 2 below.
[0112] Further details of the biopanning and selection of DBDs are presented below.
[0113] Biopanning and selection of DBDs 1. Blocking step of streptavidin coated 96 well plates (Fisher Scientific, UK) with blocking buffer (0.1 M NaHCO3 (pH 8.6), 5 mg / ml BSA, filter sterilized and stored at 4°C). 2. Discard the blocking solution and then wash with 1x TBST for 6 rounds. 3. Precomplexing step (this step can be performed at the same time as the blocking step described above), the amplified phage library was incubated with each of the dsDNA baits in separate reactions. Each bait contains an annealed double-stranded DNA oligo (dsDNA) containing one of the PDS sites. In the examples provided herein, the forward strand of each of the dsDNA oligos described above was 3'-end labeled with TEG-biotin (Sigma / Merck). The single-stranded DNA oligonucleotide sequences used to generate the dsDNA PDS-containing baits in this step are presented in Table 1 above. The length of the PDS should be kept as short as possible (to allow reliable oligonucleotide synthesis and stable base pairing to form dsDNA). Conditions for the precomplexing step: 5 micL phage library (10 11 pfu) are incubated with 50 nM dsDNA bait (final concentration) in 200 μl TBST (total volume of binding reaction). 4. Immobilization Step The pre-complexed library was immobilized on a streptavidin-coated 96-well plate (Fisher Scientific, UK). 5. Discard unbound phages and wash 10 times with TBST. 6. Elution step: 0.2 M glycine-HCl (pH 2.2) for 10-20 min, followed by neutralization (15 μM Tris-HCl, pH 9.1 to 100 μM 0.2 M glycine-HCl, pH 2.2). 7. Amplification and titering steps (for non-amplified and amplified eluates) were performed according to the manufacturer's instructions (NEB) before each selection (panning) round. In total, 2-3 rounds were performed. 8. Sequence analysis step Subsequent analysis of clonal PCR and Sanger sequences (20 clones for each bait in each test round) was performed on the "blue" (with functional b-gal) clones grown on LB agar plates from the titering analysis (NEB).
[0114] Additional modifications to the above selection procedure included: 1. After the introduction of a separate counter-selection round (before round 2 or before round 1), the unbound phage library (after step 1) was harvested, amplified and tested by titering before the actual biopanning round. 2. Introduction of a pre-screening round after the counter-selection round. This was done in experiments with baits that contained a PDS followed by an "extension region" sequence, thus allowing the biopanning procedure to be carried out at 37° C. The pre-screening step was carried out with oligos that contained "extension regions only", thus allowing the phage clones that bind to these extension regions to be removed.
[0115] Peptide 2 in the phage library had high binding ability to PDS1 dsDNA bait. Peptide 4 was selected based on its binding to PDS2 dsDNA bait. These peptides are joined to the phage coat protein in the library by a GGS linker. Therefore, peptide 1 and peptide 3 with this linker at the C-terminus were also provided for further testing.
[0116] [Table 2]
[0117] It was reasoned that phage clones carrying the selected DNA-binding peptides would alter the DNA duplex in the region displaying the cognate PDS, which could result in recognition by DNA nucleases such as T7E1, commonly used to detect insertions or deletions (indels), resulting from the activity of genome editing enzymes such as Cas9, Cas12a, etc. The results of such cleavage assays on destabilized DNA duplexes using extended dsDNA baits and T7E1 and Surveyor enzymes, presented in Table 3 below, support this hypothesis (see Figures 3a-c), i.e., binding of clones displaying the selected fusion peptides to dsDNA destabilizes the structure of the DNA duplex to promote accessibility to RNA hybridization. The resulting transient unwinding and formation of a heteroduplex structure can be envisioned, as shown in Figure 2.
[0118] [Table 3]
[0119] The modified peptides presented in Table 2 were used to test the sequences for use as DNA binding domains without M13 phage ligation. Modifications included amidation and acetylation at the N- and C-termini, respectively, to mirror natural peptide modifications that improve peptide stability and prevent degradation. In the case of peptides 1 and 3, an additional C-terminal linker (GGS) was provided as previously described. This also allowed assessment of binding properties to the cognate PDS by the presence of a linker that allows ligation to different non-phage proteins.
[0120] When presented with extended dsDNA containing the cognate PDS, the results of cleavage assays using the T7E1 enzyme for binding and non-duplex structure formation clearly pointed to the peptide 1 sequence as the best candidate to be the DNA-binding domain when PDS1 was the default cognate PDS selected (see Figure 4).
[0121] These results were further supported by performing a fluorescent DNA binding assay. FRET experiments were performed to assess the binding and unwinding activity of peptide 1, which has the ability to promote RNA hybridization upon presentation of PDS1 in an extended dsDNA target, as shown diagrammatically in Figure 5. It was hypothesized that upon binding, peptide 1 would cause some degree of transient unwinding of the dsDNA to promote binding of the targeting RNA. The appropriate targeting RNA was tagged with a 15 bp complementary sequence using the fluorescent donor, 6-carboxyfluorescein (FAM). The target DNA was tagged with the fluorescent acceptor Cy5 at the 3' end or Texas Red at the 5' end. RNA-DNA hybridization was revealed to be promoted by binding of peptide 1 to its cognate PDS1 sequence, reflected by a FRET signal-Cy5 or Texas Red fluorescent signal upon FAM excitation (see Figure 5).
[0122] Taken together, the results from the cleavage and FRET experiments are consistent with peptide 1 binding to its respective PDS in dsDNA and causing partial and transient unwinding of dsDNA to facilitate hybridization of complementary RNA. Thus, peptide 1 amino acid sequence (SEQ ID NO: 4) incorporating a C-terminal GGS linker is an example of a supported polypeptide sequence for use as a DNA binding domain in genome editing tools of the present invention, such as the ApGet system. It is recognized that similar studies can find short polypeptide sequences that are suitable for binding to other desired predefined dsDNA target sequences of 6 or less than 6 nucleotides. Site-directed mutagenesis can be utilized to further improve the binding ability to any desired PDS.
[0123] It is recognized that peptide 1 is generally provided with at least a second linker for incorporation into the complete required modular polypeptide component, which further includes a targeting nucleic acid recognition domain and an effector component. In addition, it may be desirable to extend the incorporated GGS linker. However, this is achieved by maintaining the desired binding properties to the PDS, which can be confirmed by appropriate further testing. Thus, peptide 1 as DBD can be provided with a further sequence of G and S residues, for example the linker of SEQ ID NO:5 used in the AgPet of SEQ ID NO:1.
[0124] [ka] A C-terminal extension by ALK may be used to link an effector component, for example an artificial nuclease.
[0125] [Example 2] Provision of artificial nickase The first step in the design of artificial nickases was the selection of ordered polypeptides capable of forming supramolecular self-assembled nanostructures. For this purpose, beta-sheet forming peptides of 7 or 9 amino acids in length with an alternating pattern of hydrophobic leucine and hydrophilic lysine residues were selected. Such patterns are known to form self-assembled nanostructures and have a tendency towards amyloid structure formation. This pattern was further modified by substitution of amino acids at hydrophilic positions based on known catalytic amino acids of the RuvC1 nuclease domain present in naturally occurring nucleases, e.g., Cas9.
[0126] It is contemplated that multiple such beta-sheet forming peptides can be joined by linkers or loops to control the number shown below for a 7-mer, where H is a hydrophobic amino acid, C is a catalytic amino acid, and X is a connecting linker of length n amino acids. Each of the peptide units can be identical, where each C is a different amino acid of a catalytic trio, a recognized catalytic triad, or a trio of different amino acid types that confer catalytic function of a known naturally occurring nuclease, such as the RuvC1 nuclease domain. Alternating linked peptide units may be linked N-terminus to C-terminus or back again. (H-C1-H-C2-H-C3-H-(X) n -H-C1-H-C2-H-C3-H) n or (H-C1-H-C2-H-C3-H-(X) n -H-C3-H-C2-H-C1-H) n
[0127] Table 4 below presents the exact peptide sequences of 7 or 9 amino acids that were initially used to develop efficient artificial nickases. The starting sequence (bA) as described above was an inactive 7-mer of alternating leucine and lysine residues. Naturally occurring nuclease domains are listed to provide at hydrophilic positions of the amino acids of the catalytic trio. For example, peptides bB, bD, iBD, bE, bF all had all lysines replaced by three amino acids [glutamic acid (E), aspartic acid (D), and histidine (H)] that match the catalytic amino acids of the RuvC1 nuclease domain and are separated by hydrophobic residues.
[0128] As mentioned above, the RuvC1 nuclease domain has been reported to have four catalytic amino acids, His983, Asp986, Asp10 and Glu762. Mutation of any of these catalytic amino acids was found to result in loss of function (Nishimasu et al. (2014) Cell 156, 935-949). However, in the context of the multimers mentioned above, for simplicity, it was reasoned that it may be sufficient to use identical peptide units with alternating hydrophobic residues and, for example, only a single His, Asp and Glu, thereby maintaining the tendency of the hydrophobic residues to form a stable β-sheet-like supramolecular structure while keeping all the necessary catalytic amino acid residues in close proximity to the functioning nuclease. However, the 9-mer peptide bE was provided with two Asp residues plus glutamic acid and histidine.
[0129] In some cases, leucine was also replaced by isoleucine. Thus, bD is a 7-mer starting with leucine and alternating with E, D and H in that order. IbD has the same amino acid pattern except for the replacement of leucine by isoleucine (I) as a hydrophobic residue. It may be that isoleucine residues can help the formation of β-sheet structure more efficiently than leucine residues because the former is more hydrophobic than the latter.
[0130] [Table 4]
[0131] Peptide IbD and peptide bD are previously represented above as SEQ ID NO: 7 and SEQ ID NO: 8, respectively. The remaining peptides are SEQ ID NOs: 47-59.
[0132] Such peptides were tested for their ability to self-assemble into β-sheet supramolecular structures and nickase activity. The ability to give β-sheet supramolecular structures was determined by thioflavin T fluorescence binding (Xue et al. (2017) Royal Soc. Open Sci. 4: 160696). Nickase activity was examined by molecular beacon assays combined with gel plasmid cleavage assays.
[0133] Thioflavin T assay Thioflavin T (ThT) was mixed with the peptides in reaction buffer. The peptides were diluted (from a stock solution of 4.4 mM in DMSO) to 200 μM using reaction buffer (the same buffer used for the first molecular beacon assay presented below). ThT was added to the peptides at 25 μM in reaction buffer using pulse vortexing and centrifugation. Plates containing 200 μl of peptide / Tht solution per cell were incubated at 37° C. The emission signal at 485 nm was measured every 5 min using an excitation wavelength of 450 nm. An increase in the emission signal was interpreted as an indication that ThT had bound to amyloid-like fibrils. For peptide bD, a high emission signal was measured, indicating immediate rapid self-assembly of the peptide to form such fibrils (see FIG. 7).
[0134] Nickase activity survey First, each of the peptides was screened for DNA binding and cleavage using a molecular beacon assay. For the purposes of these studies, the following DNA oligonucleotides containing 33 nucleotides were used:
[0135] [ka] The first 6 nucleotides of each end complement each other to form a 6 bp double-stranded DNA stem. This results in a 23 nucleotide single-stranded DNA loop region. The hairpin structure formed contains a fluorescent molecule (6-FAM) as the 5' modification and a fluorescent quencher (BHQ1) as the 3' modification. In the hairpin conformation, the fluorescent signal of the molecular beacon is reduced. When the peptide binds or causes cleavage, the hairpin structure disassembles, releasing the fluorophore from the quencher.
[0136] Reaction conditions were chosen to reflect conditions that might be expected to enhance the likelihood of DNA cleavage by each peptide. As mentioned above, peptide design was inspired by the catalytic active sites of known nucleases. In many cases, the commonly preferred divalent metal for such sites is Mg. 2+ However, Mg 2+ The generation of ionic reaction intermediates requires precise orientation geometry and charge requirements and can be substrate-specific and highly selective. 2+ The Mn ion requires fewer coordination sites to promote catalysis. 2+ Ions have been used to screen mutant nucleases in vitro because such ions tend to exhibit loose substrate specificity and Mg 2+ This is because the defective enzyme can be rescued compared to the Mn ion. 2+ The ions were used in this case for initial activity screening of peptides.
[0137] Each reaction was carried out in a reaction buffer of 25 mM Tris-HCl (pH 7.5), 130 mM NaCl, 27 μM KCl, and 5 mM MnCl 2. Molecular beacons were provided at a concentration of 400 nM.
[0138] After 15 min of incubation, the emission signal of the fluorophore was measured for each peptide mixed with molecular beacons at a concentration ratio of 500:1.
[0139] Peptides bL and bN gave relatively high luminescence signals under such conditions. Therefore, these peptides were used to investigate the mechanism of Mn 2+ Variations in assay conditions with respect to concentration, metal ion dependency and peptide concentration were further assessed.
[0140] metal concentration The concentration of manganese chloride used in the initial screen was in excess. This was to enhance the possibility of forming a reactive intermediate to drive the reaction to completion. Excess amounts of metal ions are often used in in vitro reactions. However, too much metal can have an inhibitory effect on the reaction. By using a range of manganese chloride concentrations from 0 to 5 mM, the optimal manganese chloride concentration for the bN peptide was found to be 0.25 mM.
[0141] metal dependence Reactions with different metal ions showed differences in the signals emitted by molecular beacons using bL or bN peptides. Degradation of the molecular beacon was more extensive in the presence of zinc chloride ions compared to manganese, calcium, and magnesium chloride for the bL peptide. In contrast, the bN peptide showed a high increase in signal, independent of the metal in the reaction. A high signal was observed without any metal ions. However, the signal was significantly higher in the presence of Mn 2+ The effect of cleavage on DNA cleavage was greater in the presence of ions, suggesting that potentially different mechanisms of DNA cleavage can be generated by using different peptide designs containing different active site regions.
[0142] Effect of peptide concentration Molecular beacons at 400 nM are 2+ or Zn 2+ The bN peptide was added at increasing concentrations from 0 to 800 μM in the presence of Mn ions. 2+and Zn 2+ However, at a concentration of 400 μM, Mn 2+ The signal from the bN peptide with ions is Mg 2+ The overall increase in signal measured for the degradation of the molecular beacon was observed with increasing concentrations of peptide bN. 2+ The signal of the bN peptide in the presence of ions did not saturate after 400 μM, and Zn 2+ In the presence of ions the signal was found to still increase up to 800 μM bN peptide.
[0143] Gel plasmid cleavage assay The signal emitted by the molecular beacon arises from the cleavage of DNA by the peptide, and the resulting oligo has a low T m Dissociation of the double-stranded region of the hairpin structure could occur due to the range (melting temperature). However, a signal could also be generated by the peptide binding to DNA and distorting the shape of the hairpin structure. This could change the proximity of the fluorophore and quencher, which could also result in signal emission. Therefore, the molecular beacon assay cannot distinguish between binding and cleavage.
[0144] For this reason, further gel plasmid cleavage assays were used to look for cleavage. Such gel cleavage assays can distinguish between intact (no cleavage), nicked (ssDNA cleavage) and linearized (dsDNA cleavage) plasmids because the three plasmid species migrate differently on an agarose gel. Intact plasmids have a supercoiled structure and therefore can migrate easily through an agarose gel, and ssDNA cleavage on the plasmid relaxes this supercoiled structure. This results in a relaxed circle / nick structure that migrates more slowly through the pore. dsDNA cleavage of the plasmid results in a linearized plasmid. The linear structure migrates faster than the relaxed circle structure, but not as fast as the supercoiled structure.
[0145] To optimize the assay, plasmid DNA (pBR322) was 2+ Increasing concentrations of peptides bL or bN were mixed in the presence of ions. This was found to result in an increased percentage of nicked plasmid after 12 hours. The percentage of nicked species increases with increasing peptide concentration for both bL and bN. The percentage of nicked plasmid decreases with peptide concentrations greater than 600 μM for the bN peptide, while the amount of nicked plasmid continues to increase at such higher peptide concentrations for the bL peptide. An increasing proportion of plasmid DNA is also observed in the wells of the gel. This DNA is able to bind to the peptide, causing a complex that cannot migrate into the gel.
[0146] Peptides bL, bN and bD all demonstrated some degree of cleavage, however peptide bD was selected for further study due to its higher propensity to give rise to highly stable amyloid beta sheet structures as assessed by the ThT assay.
[0147] Protein Design The percentage of cleavage caused by peptide bD was found to increase with peptide concentration, suggesting that assembly of the peptide into a supramolecular structure enhances its functionality as a cleavage enzyme. However, uncontrolled peptide assembly can lead to large aggregates. Therefore, constructs were designed in which several identical peptides are linked by short amino acid linkers or loops that can be expected to create beta-turn-like structures. This not only connects the monomer units, but also allows for close proximity of the monomer units, which promotes the peptide to create the desired beta-sheet structure.
[0148] The conceptual design demonstration considered the following principles: 1. Beta-strand mimetic peptides of 7 or 9 amino acids in length (hereafter referred to as beta-peptides) in which an alternating pattern of hydrophobic and hydrophilic residues is used. The hydrophobic leucine residues are interchangeable with isoleucine. 2. Provision of a four amino acid linker to connect the individual beta-strand mimetic monomers designed to impart a beta turn and promote a beta-sheet fold between the connecting peptide units. 3. Provision of flanking units of negative design, i.e., to prevent intermolecular binding and aggregation between individual beta sheet protein molecules.
[0149] In the final proof of concept construct, the leucine residues of peptide bD were replaced with isoleucine to give the 7-mer peptide ibD, since ibD was found to produce no observable aggregation even at the highest tested concentration. Since isoleucine is more hydrophobic than leucine, it is possible that the former could support the formation of β-sheet structure more efficiently than the latter. The bD peptide did not show the greatest signal from in vitro peptide functionality studies. However, this peptide showed the least tendency to aggregate and precipitate, and therefore, the bD peptide was expected to be a more stable peptide as a monomer unit of larger multimeric constructs such as designed decamers. To enhance its potential to fold into stable secondary beta structures, the leucine was replaced with isoleucine. This has previously been demonstrated to enhance enzymatic activity of amyloidogenic peptides (Rufo et al. Short peptides self-assemble to produce catalytic amyloids. Nature Chem. (2014) 6, 303-309) and was applied to potentially generate more stable folds for multimeric constructs.
[0150] Betaturn Design To create beta turns that can connect the beta peptide units to each other, the appropriate length and amino acid composition of each beta turn must be designed. Both the design of the beta turn in the natural protein and the type of amino acid side chains required to orient the beta chain sequence (beta peptide units) into hydrogen bonds to create a beta sheet structure are considered. Studies have shown that the optimal beta turn length in natural proteins is often 2, 4 or 5 residues long. In natural proteins, beta turns consisting of 2 or 5 amino acids often fold to favor a particular enantiomerism, and beta turns with a length of 4 amino acids do not fold to favor any particular enantiomerism. Therefore, in the proof of concept, an amino acid length of 4 was selected to provide a linker between the selected self-assembling peptides.
[0151] The role of the beta turn is to facilitate this self-assembly by bringing the peptide units closer together. Therefore, there is no need to design rigid beta turns with amino acid residues that may enforce a particular turn conformation and limit the potential structural folding of the beta peptide. On the contrary, the design of the beta turn is much more flexible. In natural proteins, polar amino acid residues (serine, asparagine, glutamine and threonine) are often present in beta turns. The hydrophilic side chains of these amino acids can easily hydrogen bond to stabilize potential beta turn conformations. Furthermore, the side chains of these amino acids are uncharged and therefore less likely to contribute to structural rearrangements or interact with active site residues. To stabilize the predominantly hydrophobic beta peptide units, charged amino acids (glutamic acid and aspartic acid) are also incorporated during the beta turn design, potentially increasing the solubility of the complete protein. The choice of the fourth amino acid is further restricted to preferentially glycine. The small and flexible side chain of this amino acid prevents steric repulsions and allows the beta-beta connection.
[0152] Negative design of adjacent units The ability of beta peptides to easily self-assemble into beta-sheet structures is advantageous for creating beta-sheet protein structures, but the self-assembly of beta peptides is uncontrollable, potentially creating large fibrillar structures. By introducing beta turns, the number of beta peptides per molecule can be controlled. However, this does not prevent the intermolecular assembly of these beta-sheet protein molecules. To prevent this, a "negative design" is incorporated into the beta-sheet protein design, as seen in natural β-sheet proteins (see Richardson and Richardson (2002 PNAS 99, 2754-2759)). For this, a lysine amino acid residue, a single amino acid with a charged side chain, was introduced into each adjacent beta-peptide unit, thereby providing a charged residue at the intermolecular interface to inhibit the aggregation of beta-sheet protein molecules.
[0153] In a proof of concept construct, the second and third hydrophobic residues (positions 3 and 5) were used in place of the lysine amino acid residues in the respective N- and C-terminal beta-peptide units, see FIG.
[0154] Beta-sheet peptide orientation FTIR experiments suggested that the beta-peptide units adopted a predominantly antiparallel beta-sheet structure, but the same experiments also suggested that the self-assembly was dynamic and evolved over time with the inclusion of balanced beta-sheets and some unstructured loop regions. It was recognized that the artificial constructs could have two different types of secondary structure formation, where monomers bind to each other in the same N-to-C-terminal direction or in opposite directions. To create a non-inverted beta-sheet-like assembly, each peptide unit was connected in the N-to-C-terminal direction using a beta-turn. This results in the catalytic residue at position 2 of the first peptide unit potentially hydrogen bonding to the catalytic residue at position 6 of the second peptide unit (Figure 9). Thus, for example, consecutive ibD peptides could be linked by an N4 linker providing a beta-turn as follows:
[0155] [ka] See again Figures 8 and 9.
[0156] To create an inverted beta-sheet-like structure, the sequence of every other peptide unit is inverted, resulting in the catalytic residue at position 2 (N-to-C-terminal direction) of peptide unit 1 potentially hydrogen bonding to the catalytic residue at position 2 of peptide unit 2 (Figure 9).
[0157] [Example 3] Fusion of an artificial nickase to the ApGet modular polypeptide Taking the above design aspects, the beta peptides were first assembled into pentamers consisting of 5 beta peptides and decamers consisting of 10 beta peptides. The pentamers and decamers (Figure 8) were flanked at their N- and C-termini by beta peptide units incorporating a negative design to reduce aggregation as explained above. Additionally, inverted and non-inverted designs were created (see Figure 9). Each peptide unit represents a potential beta strand and is separated by a beta turn of 4 amino acids. The N-terminus was fused directly to the C-terminus of the DBD shown in SEQ ID NO:1 through a linker (Linker 1). A flexible and soluble region was provided at the end of the C-terminal flanking region consisting of asparagine, glutamine, glycine and serine.
[0158] Further testing of nickase activity as part of a longer fusion construct providing both the RSBD and DBD, each protein construct was fused to the C-terminal region of the construct ApGet-i (SEQ ID NO: 1 minus the final linker and nickase) by providing a linker sequence. The linker selected was:
[0159] [ka] It was. This final construct is called ApGet1.0 nuclease protein.
[0160] Such modular polypeptides were expressed using two different systems using different mechanisms of action or in vitro characterization studies: (i) in vitro transcription translation (IVTT) and 2) bacterial recombinant protein production (E. coli protein production). To enhance solubility and to facilitate downstream protein purification steps, a series of protein tags were attached to the ApGet1.0 sequence (see FIG. 10). The first protein design contained an N-terminal hexahistidine tag (6His-tag) preceding a protease specific cleavage site (TEV cleavage site). The cleavage site can be cleaved by TEV protease, allowing removal of the 6His-tag after purification of the ApGet1.0 protein. The second protein design was considered to potentially enhance the solubility of the ApGet1.0 construct by fusing the maltose binding protein (MBP) to the N-terminus of ApGet1.0. The MBP tag was provided together with an N-terminal 6His-tag. The C-terminus was fused to a second 6His tag followed by a spacer region consisting of glycines and serine followed by a TEV protease cleavage site. The spacer reduces steric hindrance of MBP upon folding of ApGet1.0 and allows proper access to the TEV protease for successful cleavage of the His-MBP-His tag from ApGet1.0.
[0161] In vitro transcription and translation (IVTT) An IVTT expression system was used to enhance expression as discussed above and initially screen the ApGet1.0 protein for cleavage activity. Cell-free in vitro transcription and translation systems translate proteins without the constraints of a cellular environment. Such systems are often used to express insoluble and toxic proteins that are not stable in cells and would otherwise aggregate or cause cell death, reducing protein yield. The PURExpress® IVTT system (New England BioLabs) was used, which allows proteins to be expressed without proteases, RNases and DNases. This was desirable because the presence of proteases can reduce total protein yield and RNases / DNases can interfere with in vitro cleavage assays.
[0162] The MBP-fused ApGet1.0 was cleaved prior to any cleavage assay to remove the large MBP protein, which could sterically hinder the action of the ApGet1.0 protein. The cleaved protein was semi-purified prior to the cleavage assay by performing a crude (batch) purification of the protein mixture with Ni-NTA resin. The resin captures the His-MBP-His fusion tag and the 6His tag, which extracts the His-TEV protease from the protein solution.
[0163] First non-specific cleavage assay using ApGet1.0 protein Initial cleavage assays were performed with ApGet1.0 decamer protein expressed in IVTT and incubated with plasmid DNA substrates. To serve as a negative control, assays with ApGet1.0 decamer were performed alongside IVTT-expressed ApGet-i (expressed with the additional inclusion of the DBD enzyme linker of ApGet1.0 at the C-terminus) to control for miscellaneous activity from the IVTT mixture and a no protein control to control for different buffer compositions. Without an associated targeting nucleic acid, any cleavage of DNA substrates by ApGet1.0 protein was expected to be non-specific. The design principles used for the artificial nucleases predict that divalent metal cations are essential for activity, so no metal, Mn 2+ or Mg 2+ Buffers containing Mn were tested. 2+ and Mg 2+ was used at 5 mM. The remaining buffer components reflected physiological conditions: 25 mM Tris-HCl pH 8.0, 150 mM NaCl.
[0164] As an example, the ApGet1.0 protein of SEQ ID NO:1 is 2+ or Mn 2+ It was found to nick plasmid DNA (representing ssDNA cleavage) in the presence of ions. As mentioned above, this is ApGET, which uses the artificial nuclease of SEQ ID NO:6.
[0165] [ka]
[0166] In this sequence, all 7-mer monomer units of IbD (IEIDIHI) are linked by a 4-mer linker in the N- to C-terminal direction without inversion to provide a beta-turn. The adjacent N- and C-terminal peptide units are shown in bold immediately above, with the second and third hydrophobic residue positions, respectively, replaced by a positively charged lysine residue. As mentioned above, the C-terminal unit is also provided in the C-terminal sequence of NQGS to aid in solubility.
[0167] In the absence of divalent metals, no cleavage was detected. The percentage of nicked species was 100% for Mg 2+ Compared to Mn 2+ The no protein control and ApGet-i were unable to cleave the DNA substrate in either the presence or absence of metal, indicating a specific cleavage activity associated with the ibD peptide decamer (see FIG. 11).
[0168] Recombinant protein expression in E. coli Expression of the ApGet1.0 10-mer protein using the IVTT system is useful as it can be screened for initial cleavage activity. However, increased protein yield, purity and the ability to quantify the protein can help clarify the efficiency of the enzyme and understand its mechanism of action. E. coli recombinant techniques can produce larger yields of protein and purification can be aided using a series of columns for chromatography.
[0169] Expression and purification For increased high-level expression, the DNA sequence encoding the MBP-tagged ApGet1.0 protein discussed above was placed into a high-copy plasmid, pET24a, which upon transformation ensures that each cell contains many copies of the ApGet1.0 protein, thereby increasing protein yield. This was further enhanced by using a T7 promoter. This plasmid was used to transform DE3 lysogenic bacterial cells containing the T7 RNA polymerase gene, and expression of the RNA polymerase was induced by IPTG.
[0170] The host cells chosen were BL21(DE3)RIPL codon-plus E. coli cells (Agilent Technologies). The MBP-fused ApGet1.0 protein was overexpressed for 12–18 h at an induction temperature of 18°C. The ApGet-i protein (again containing a C-terminal linker) was similarly expressed at 37°C for 4 h.
[0171] To purify the ApGet1.0 protein, cell lysates containing overexpressed ApGet1.0 protein were passed through a Ni-NTA affinity-based chromatography column using the 6His-tag to purify the MBP-fused ApGet1.0 protein from the lysate. A second cation exchange column was used to further purify the MBP-ApGet1.0 protein, followed by a cleavage reaction with TEV protease, which cleaves the target protein from the entire His-MBP-His protein tag. Final separation of the His-MBP-His tag and TEV protease from the target protein was performed by batch purification using Ni-NTA resin. The target protein remains in solution, while the tag and protease bind to the resin.
[0172] [Example 4] Proof of concept testing in bacterial cells We first devised a test platform to investigate the on-target mobilization of ApGet polypeptides in E. coli cells. E. coli was co-transformed with a first plasmid, named "editor", encoding the ApGet arrangement to be tested, and a second plasmid, the "detector", encoding a factor that generates a differential signal by gene editing mediated by the ApGet system. The two plasmids have replication origins of distinct compatibility and contain different antibiotic resistance genes that allow the co-transformation and stable maintenance of both plasmids in the E. coli population.
[0173] Delivery of the ApGet-i system into bacterial cells and demonstration of transcriptional inhibition The editor plasmid is a multicopy plasmid that encodes the ApGet polypeptide under the control of the native SpCas9 promoter and the nucleic acid component (NAC) under the control of the J23119 promoter, a commonly used strong synthetic promoter that drives gRNA expression in E. coli (see, for example, in Qi et al. Repurposing CRISPR as an RNA-guided platform for sequence specific control of gene expression. Cell (2013) 152, 1173-1183). The detector plasmid contains a Lac promoter-driven eYFP or LacZa protein coding sequence, the 5' upstream region of which contains four identical targets complementary to the targeting elements placed in the ApGet system. See Figure 12.
[0174] The target element sequence (TE) in the editor plasmid is:
[0175] [ka] It is.
[0176] The targeting element in NAC is fused to a BoxB RNA scaffold sequence using the connector AATTT, which binds to the RSBD of the ApGet polypeptide.
[0177] The targets in the two detector plasmids are as follows (SEQ ID NO:62): Target sequences in PC1 and PC2 from 5' to 3' (sequences in grey boxes are PDS, sequences in white boxes are targets, separated by other sequences in between):
[0178] [ka]
[0179] In preliminary experiments, an "ApGet-i" system was created in which the protein components encoded by the editor plasmids were identical to the concatenated RSBD and DBD components of the ApGet polypeptide of SEQ ID NO:1 minus the artificial nickase. Nevertheless, it was expected that successful recruitment of the ApGet modular polypeptide-NAC complex to the PDS-containing target would inhibit the transcription of eYFP, reducing the fluorescent signal or inhibiting the expression of LacZa. This turned out to be the case; see Figures 13a and b. After induction, a decrease in eYFP fluorescence or in the activity of the LacZa gene (in hydrolysis of o-nitrophenyl-beta-D-galactosidase) was observed only when either the detector plasmid was delivered together with the ApGet-i expression plasmid, indicating the transcription inhibitory activity of ApGet-i.
[0180] Table 5 below summarizes constructs tested with detection of expression of eYFP. Table 6 below summarizes constructs tested with detection of expression of LacZa. Construct 1 = plasmid for expression of ApGet-i. Construct 2 = eYFP detector plasmid. Construct 3 = LacZa detector plasmid. Construct 4 = plasmid with a backbone similar to construct 1 but without the ApGet-i expression cassette.
[0181] [Table 5]
[0182] [Table 6]
[0183] Testing with the Fok1 nuclease domain Next, testing of the ApGet concept was extended by fusing the catalytic unit of FokI nuclease to the DBD domain of ApGet-i through a linker (linker of SEQ ID NO:5 shown above in SEQ ID NO:1) to create a fully functional RNA-dependent genome editing enzyme. The nuclease linker was selected in part to optimize the nucleotide sequence for the synthesis of synthetic DNA fragments. In this case, the editor plasmid of the ApGet-FokI system consisted of a nucleic acid component (NAC) expressed under the control of the constitutive J23119 promoter and the ApGet-FokI protein under the control of the inducible Tet-On promoter (Tet repressor (TetR) regulates the expression of the Tet promoter (TetP) by binding to the Tet operator (TetO) sequence). The backbone of the editor plasmid was derived from pET28a and contains a ColE1 origin of replication and a kanamycin resistance expression unit (see Figure 14). The detector plasmid contained a targeting cassette with two identical target sequences (recognized by a single ApGet NAC component expressed from the editor plasmid used in the experiments) on opposing DNA strands, e.g., with a PDS sequence located within its respective target sequence (we term this the "inward" or "PDS-in" configuration). The target element sequence and the target sequence were as follows: Target element sequence:
[0184] [ka] Target sequence (grey box indicates PDS, white box indicates target sequence): "PDSout" arrangement:
[0185] [ka] "PDSin" configuration with 5nt spacer:
[0186] [ka] "PDSin" configuration with 8nt spacer:
[0187] [ka] "PDSin" configuration with 14nt spacer:
[0188] [ka] (SEQ ID NOs: 64 to 67)
[0189] The detector plasmid also contained elements from pCC1 that contained the ori2 and oriV origins of replication as well as an antibiotic resistance gene, distinct from the antibiotic resistance gene on the editor plasmid, in this particular example a chloramphenicol resistance gene (again, see Figure 14).
[0190] Successful recruitment of ApGet-FokI would be expected to result in a site-specific double-strand break and loss of the detector plasmid, which should reduce bacterial growth since this plasmid contains a resistance gene to the antibiotic present in the culture medium (Figure 15). To facilitate detection of the editing activity of ApGet, these tests were performed by co-transformation of the editor and detector plasmids into EPI300 cells (EpiCentre), which were engineered to encode the oriV activator trfa under the tight regulation of the L-arabinose-inducible ParaBAD promoter. Growth of EPI300 cells co-transformed with the editor and detector plasmids in a growth medium lacking L-arabinose would result in very low copy numbers of the detector plasmid (because the multicopy ori-V origin of replication was not activated). This was expected to render the cells more sensitive to chloramphenicol when the detector plasmid was lost as a result of the DNA editing activity of ApGet-FokI. Detection of editing activity was by measuring OD600 one day after induction of ApGet with anhydrotetracycline (ATC). Control experiments included co-transformation of ApGet-i expressing editor (ApGet construct, lacking the nuclease domain) or non-ApGet expressing plasmid and detector plasmid into EPI300 cells. Successful testing is illustrated by the decrease in OD600 observed with expression of ApGet-FoK1 in cells also containing a detector plasmid with the "PDS-in" format for PDS-target sequence pairs and spacers between these pairs of 5, 8 or 14 nt (see bars 1-3 in Figure 15).
[0191] Testing with artificial nickase modules Following successful testing of the ApGet platform fused to FokI nuclease, we extended testing to ApGet fused to an artificial nuclease (AN) containing the preferred selected artificial nickase (SEQ ID NO: 6) as previously described.
[0192] In summary, the ApGet system was delivered to E. coli bacterial cells with the goal of correcting a disrupted gene. The cells were co-transformed with a detector plasmid carrying the disrupted nanoluciferase gene and an editor plasmid expressing the ApGet system, in which the polypeptide component of the system (e.g., a modular polypeptide having the amino acid sequence of SEQ ID NO:1) is under the control of a T7 promoter and Lac operator, and the RNA component of the system is expressed under the control of a synthetic promoter (J23119). The repair template is also embedded in the editor plasmid along with roughly 200 bp right and left homology arms.
[0193] The sequences from 5' to 3' of the donor or repair template elements were as follows: (the sequence of the left homologous arm is underlined and the sequence of the right homologous arm is highlighted in grey):
[0194] [ka]
[0195] Various ApGet polypeptides with different lengths and orientations of nuclease peptide units were tested (see FIG. 16). The expression cassette was modified compared to the expression cassette for expression of ApGet polypeptide production for purification and testing as discussed above. Expression trials showed that adding the soluble tag MBP enhanced protein expression in E. coli. However, such a tag was not necessary to obtain editing activity. However, the cleavable 6His tag was retained. Specifically, we tested ApGet1.0, which has a nuclease domain consisting of up to 10 peptide units (decamers) in the "non-inverted" and "inverted" protein translation orientations. To maximize the detection sensitivity of ApGet editing activity, the detector plasmid provides a "gain-of-function" detection system, in which the generation of a nick in the double-stranded DNA sequence of the detector element can promote homology-directed repair (HDR) to restore detector activity in the presence of a donor template. To further increase the detection activity, a detector was designed in the form of a nanoluciferase gene expressed under the control of a T7 promoter, disrupted by a targeting cassette, and recognized by ApGet1.0. The targeting cassette is composed of two identical target sequences tandemly aligned in a "PDS-out" or "PDS-in" configuration (see constructs f and g, respectively, in Figure 17). The PDS-in and PDS-out detectors were located exactly between the homology arms, using the same targeting sequences previously used and described above.
[0196] We reasoned that the targeting cassette, consisting of two target sequences located in close proximity to each other, would increase the clustered nicking activity of ANs in ApGet1.0 and thus increase the chances of double-strand breaks, which in turn might trigger more efficient HDR, resulting in the expression of functional nanoluciferase.As mentioned above, the homology-directed repair (HDR) template for nanoluciferase was constructed on the same plasmid expressing ApGet as a separate unit.
[0197] Figures 18 and 19 further illustrate the use of PDS-in and PDS-out configurations of PDS-target sequence pairs targeted by the synthetic genome editing system of the present invention.
[0198] The various configurations of ApGet1.0 that were tested are listed in Table 7 below.
[0199] [Table 7]
[0200] BL21(DE3) cells were co-transformed with each of these configurations and either of two types of detectors in which the nanoluciferase open reading frame (Nanoluc ORF) was disrupted (PDS-in or PDS-out configuration). The Nanoluc signal / OD600 ratio, which reflects the efficiency of HDR to restore functional nanoluciferase, was analyzed for each sample (Figure 19). HDR events were confirmed by Sanger sequencing analysis as well as deep amplicon sequencing of the repaired nanoluciferase coding sequence of the detector units.
[0201] These results clearly demonstrated the discernible activity of the ApGet1.0 system in promoting HDR in E. coli, apparently by nicking the integration site. The ApGet modular polypeptide of SEQ ID NO: 1 with the appropriate targeting nucleic acid is preferred for this purpose.
[0202] [Example 5] Binding ability of selected DBDs to mutant PDS As mentioned above, the array
[0203] [ka] The DBD of was selected based on its binding affinity to a dsDNA sequence represented by the 5' to 3' sequence 5'GAGGTC3' (the predefined sequence (PDS) initially selected for use in targeting). Further studies were undertaken to determine the importance of various nucleotides in this PDS for binding to the same DBD using a nanoluciferase detection assay (Promega). This was based on previous findings described above that expression of ApGet-i in bacteria can repress transcription of sequences containing target sites for the PDS used for NAC plus DBD selection.
[0204] The ApGet-i expression plasmid (editor plasmid) was provided as shown in Figure 21. The nanoluciferase expression cassette was cloned into a second plasmid vector (detector plasmid) carrying the PDS originally used for DBD selection and a variant of that PDS at the target region shown in Figure 22.
[0205] Bacterial cells were transformed with an editor plasmid in which ApGet-i expression was under the control of the TetR promoter. NAC was also expressed by the same plasmid under the control of the constitutive J23119 promoter. Detector and editor plasmids with different resistance genes and compatible origins of replication were provided. The ApGet-i expression plasmid has a CoIE1 high copy origin of replication. The detector plasmid has Ori2 and OriV origins of replication. Transformation of the detector and editor plasmids was carried out in EPI300 cells (Lucigen), which allow copy control of the detector plasmid. The transformed cultures were induced the next day with anhydrotetracycline (ATC) to induce expression of ApGet-i, and nanoluciferase expression was induced with IPTG (isopropyl β-D-1-thiogalactopyranoside). The nanoluciferase signal was monitored on a microplate reader and normalized to take into account the extent of bacterial growth in the culture.
[0206] A plasmid expressing LacZ protein with a similar backbone to the editor plasmid but without the ApGet-i ORF was used as a control (Figure 23a), which allowed us to compare the reduction in the expression level of the luciferase gene under the influence of different PDSs.
[0207] The experimental results are summarized in Figure 23b.
[0208] It has been demonstrated that DBDs directed to different PDSs of DNA are effective in inhibiting luciferase expression. This confirms that the DBD of the present invention is flexible in its ability to bind to DNA. For example, the DBD of SEQ ID NO: 4 can be expected to be effective in targeting not only the 5'GAGGTC3' sequence in dsDNA, but also, for example, 5'TTGGGTC3' and 5'AAAAAA3' with good, if not better, affinity.
[0209] [Example 6] Gene editing research in mammalian cells The editing efficiency of ApGet in mammalian cells was tested in HEK293T and HAP1 cells. For this purpose, optimized configurations of ApGet in plasmid or RNP form were created. For transient expression of ApGet in mammalian cell culture, ApGet and NAC expression units were constructed in a circular vector of minimal required length, containing only a high copy replication origin and an antibiotic resistance gene that allows efficient production of this vector (Figure 24). ApGet and NAC units were expressed under RNA pol II and RNA pol III promoters, respectively: CMV, Cbh / CAG or Ef1A promoters were used for the expression of ApGet and U6 promoter for the expression of each of the two NAC units. The variants of ApGet expression cassette used are more fully illustrated in Figure 25. A nuclear localization signal was provided at the N-terminus after a FLAG® epitope tag.
[0210] (i) Detection of PD-L1 / CD274 editing by ApGet using immunofluorescence microscopy In this example, the target region of ApGet was selected to affect the expression of all known alternatively spliced transcripts with the open reading frame of the PD-L1 gene (also known as CD274) (Q9NZQ7-1, Q9NZQ7-2, Q9NZQ7-3, Uniprot). The layout of the ApGet targeting region in exon 4 of the PD-L1 gene target, including the binding site and PDS site of the ApGet target element, is shown in Figure 26.
[0211] ApGet RNP was delivered into haploid HAP1 cells by electroporation (NEON, Thermofisher). Haploid cells were chosen so that the direct phenotypic consequences of any reading frame alterations or gross deletion mutations could be predicted since these cells only contain a single copy of the target gene. Two hours after transfection, interferon gamma was added to the transfected cell cultures to induce PD-L1 expression. Twenty-four hours after transfection, cells were fixed and PD-L1 was detected using an immunofluorescently labeled antibody.
[0212] Quantification of PD-L1 expressing cells was performed by ImageJ and values were plotted into a graph using MS Excel; see Figure 27.
[0213] This experiment confirmed that PD-L1 protein expression was greatly reduced in response to interferon-gamma induction, consistent with gene editing at the PD-L1 locus targeted by the ApGet system.
[0214] (ii) ApGet activity promotes genomic integration of donor template sequences through homology-directed repair (HDR). HAP1 cells were co-transfected with vectors expressing ApGet (Figure 25, construct 1), NAC to two flanking regions of the target region in the PD-L1 gene in the form of RNA oligos, and a donor template in the form of a single-stranded DNA oligo (ssODN, Alt-R, IDT). The donor template was designed to have an HA tag sequence with a stop codon at the 3' end and flanked by 50 bp homology arms on both sides. The homology arms of the donor template were designed to bind to sequences 30 bp upstream and downstream of the target region to minimize the effect of potential ApGet activity on and adjacent to the target region on the already integrated HDR template (Figure 28a). Cells were harvested 24 hours after transfection, and genomic DNA was extracted and analyzed by HDR-specific PCR, in which one of the primers binds to the integrated HA tag and another primer binds to the genomic sequence outside the homology arms. The layout and results of the experiment described in this example are presented in Figure 28b.
[0215] The experiments provide evidence that ApGet-mediated editing triggered the incorporation of an HA tag into chromosomal DNA by the HDR mechanism.
[0216] (iii) ApGet induces chromosome editing at the target site in human cells HEK293T cells were transfected with an ApGet expression construct (j and k shown in Figure 25) directed at the TSKU gene, containing PDS sequences juxtaposed on either side of the target site; see Figure 29. Genomic DNA was collected 48 hours after transfection and amplified by PCR using primers flanking the target site. Sanger sequencing-derived DNA sequence traces using fluorescent chain-terminating residues were obtained from the PCR products. These were analyzed using TIDER, a cloud-based software package that calculates the extent of gene editing by assessing the proportion of wild-type sequence at each position. Non-transfected samples were used as controls (Brinkman et al, Nucleic Acids Res. (2018)).
[0217] The results shown in Figure 30 indicate a higher substitution rate from the wild-type sequence from sequence position 5' to base pair 150 in the test sample compared to the control sample, strongly supporting the assertion that the ApGet system edits chromosomal DNA at the target site in human HEK-293 cells.
[0218] [Example 7] Recombinant production of ApGet proteins Further studies on ApGet protein production were undertaken using recombinant protein expression and purification techniques. These studies were undertaken with a view to optimizing ApGet-i and ApGet production (hereinafter referred to as "protein of interest" or POI) by expression in Escherichia coli (E. coli). The POI was modified to allow soluble expression and ease of purification by producing fusions with various tags. The starting configurations used for these fusion proteins are detailed in Figure 31 and consisted of the following: His tag (6x histidine) for immobilized metal affinity chromatography (IMAC) based purification, maltose binding protein (MBP) or small ubiquitin-like modifier (SUMO) soluble tags, a spacer followed by a TEV protease cleavage site, a FLAG® epitope tag, a nuclear localization sequence (NLS), a POI sequence and finally a C-terminal Strep Tag® II sequence. This tag may be used in both ApGeti and ApGet production. However, ApGeti can also be purified without the StrepII tag and therefore expressed by simply including linker 1 (SEQ ID NO:5) at the C-terminus.
[0219] Several glycine or serine residues were incorporated between the His tag and the solubility tag. Since the spacer between the solubility tag and the TEV protease site was initially a glycine-serine-rich linker, this was replaced by an asparagine-rich linker to reduce the number of glycine and serine residues in the protein construct.
[0220] TEV protease-mediated cleavage in these constructs generated an NH2-terminal leader consisting of the amino acids Gly-Trp-Gly-Ser (GWGS), thus introducing additional aromatic amino acid residues to improve the sensitivity of absorbance measurements at 280 nm for protein quantification.
[0221] When incorporated as a COOH-terminal tag, the StrepII tag allows the use of affinity chromatography with a specifically engineered version of streptavidin to select for the full-length protein. This allowed the removal of prematurely terminated translation products from the final protein preparation. Through this method, production of the full-length ApGet protein complete with purification, detection and solubility tags was successfully achieved.
[0222] [Example 8] Transcriptional activation using ApGeti linked to VP64 Experiments were undertaken to demonstrate the ability of the ApGeti platform linked to non-nuclease effectors to provide gene regulatory modifications as well as the modularity of the individual components of the system. The artificial nuclease in the ApGet modular polypeptide corresponding to SEQ ID NO: 1 was replaced with the VP64 transcription activator protein to provide a construct named ApGeti-VP64, and the ability of this construct to transcriptionally activate the ASCL1 gene in HEK293T cells was determined. Two variants of ApGeti-VP64 with linker variations and dCas9-VP64 constructs were also tested.
[0223] The ASCL-1 gene was selected as a known target gene for such CRISPR / Cas-based transcriptional activators with low background transcription levels. Four locations in the gene promoter region were selected to be targeted using four different NACs. Each NAC consisted of a targeting element joined to the same RNA scaffold by a short connector (box B lambda phage sequence). As previously mentioned, three different modular ApGet polypeptides were used with the same VP64 effector and the same RSBD (lambda N22 peptide). Linker 1 between the VP64 component and DBD or linker 2 between the RSBD and DBD were varied as shown in Table 8 below.
[0224] [Table 8] The amino acid sequences of the various components of the ApGeti-VP64 modular polypeptides listed above are provided in Table 9.
[0225] [Table 9]
[0226] Each ApGeti-VP64 mutant was expressed together with four separate NAC transcription units from the same plasmid. As shown in Figure 32, all plasmid constructs had an overall structure including a minimal vector backbone (containing a ColE1 origin of replication and a bacterial resistance gene) and an insert containing five separate transcription units: two U6 promoter-driven NAC units, followed by an Ef1-α promoter-driven ApGeti-VP64 expression unit in the reverse direction of transcription, followed by another two separate U6 promoter-driven NAC units. The sequences of the different NAC expression units as well as the ApGeti-VP64 mutant expression units are presented in Table 10 below. Table 11 summarizes the expression units of each plasmid construct.
[0227] NAC numbering (1-4) is according to the position of the NAC in the construct (5' to 3'). NAC1 and NAC2 sequences are given on the reverse complement strand as they are in reverse transcriptional orientation relative to ApGeti-VP64 and NAC3 and NAC4. Each NAC expression unit contains a U6 promoter followed by a terminator sequence including RS (bold), TE (underlined sequence) and 7xt. Each ApGeti-VP64 expression sequence contains an Ef1-α promoter and 5' untranslated region followed by ApGeti-VP64 coding sequence (underlined) including RSBD (italic capital letters), followed by linker 2 (bold), DBD (bold capital letters), linker 1 (capital letters) and the VP64 module (italic letters).
[0228] Table 10 NAC1 (SEQ ID NO:72)
[0229] [ka] NAC2 (SEQ ID NO: 73)
[0230] [ka] NAC3 (SEQ ID NO:74)
[0231] [ka] NAC4 (SEQ ID NO: 75)
[0232] [ka] ApGeti-VP64 (SEQ ID NO: 76)
[0233] [ka] RSBD (Lambda N22) DNA sequence (SEQ ID NO:77)
[0234] [ka] Linker 1 (L1) DNA sequence (SEQ ID NO:78)
[0235] [ka] Modified Linker 1 (L1a) DNA sequence (SEQ ID NO:79)
[0236] [ka] Linker 2 (L2) DNA sequence (SEQ ID NO:80)
[0237] [ka] Modified Linker 2 (L2b) DNA sequence (SEQ ID NO:81)
[0238] [ka] VP64 module DNA sequence (SEQ ID NO:82)
[0239] [ka] DBD DNA sequence (SEQ ID NO:83)
[0240] [ka]
[0241] [Table 10]
[0242] Cell growth and transfection HEK293T cells were transfected with plasmids expressing the ApGet-i protein fused to VP64 plus the four NACs using Lipofectamine 3000 reagent according to the manufacturer's instructions. 2.5 mg of the ApGeti-VP64 expression plasmid was transfected at 0.7–1 × 10 per well in a 6-well plate format (ThermoFisher Scientific). 6 Used in HEK293T cells.
[0243] RNA and cDNA preparation Total RNA was extracted 48 hours after transfection using a commercially available RNA preparation kit (Monarch Total RNA Miniprep Kit, NEB). Harvested and purified RNA was also transcribed using a commercially available kit, Improm-II™ Reverse Transcription System. mRNA expression levels were quantified using predesigned Taqman qPCR assays (IDT). Target Cq values (FAM dye) were normalized to ActinB Cq values (HEX dye). Fold change in target gene (ASCL1) expression was determined by comparison with mock-transfected controls.
[0244] result The results of Taqman qPCR assays assessing transcriptional activation by various ApGet-VP64 polypeptides are shown in Figure 33. For each sample, the ASC1-specific signal was normalized to the ActB signal. The relative fold change in the transcription level of the ASCL-1 gene normalized to the level of mock-transfected cells (vector backbone without ApGet-VP64) is presented in the graph.
[0245] All of the ApGeti-VP64 polypeptides tested demonstrated transcriptional activation.
Claims
1. A nucleoprotein complex for use in modifying a target nucleic acid sequence, comprising: (A) a targeting nucleic acid; and (B) a modular polypeptide component; the targeting nucleic acid (i) a targeting nucleic acid element (TE) that is complementary to a region of the target; (ii) a recognition element that specifically interacts with a nucleic acid recognition module of the modular polypeptide component; and (iii) a connecting sequence linking the targeting element and the recognition element; The modular polypeptide components are linked as separate functional modules: (a) the nucleic acid recognition module lacking enzymatic activity; (b) a DNA binding domain (DBD) for assisting in melting and / or unwinding the double helix of the target, wherein the DBD recognizes a predetermined sequence in the target, and wherein module (a) and the DBD are not found in conjunction in nature; and (c) an effector component for use in modifying a target, whereby site-specific modification occurs as directed by said TE and DBD; Including, Module (b) is coupled to module (a) and module (c), respectively; Nucleoprotein complexes.
2. A nucleoprotein complex as described in claim 1, wherein module (b) is linked to module (a) by a first linker, and module (b) is linked to module (c) by a second linker.
3. The nucleoprotein complex of claim 1 , wherein the targeting nucleic acid is RNA.
4. 2. The nucleoprotein complex of claim 1, wherein the recognition element of the targeting nucleic acid is an RNA scaffold (RS) that binds to an RNA scaffold binding domain (RSBD) that provides the module (a) of the modular polypeptide component.
5. The building blocks include (A) a complete nucleic acid building block (NAC) and (B) a modular polypeptide building block, The NAC: (i) a targeting element (TE) that is complementary to a region of the target; (ii) an RNA scaffold (RS) that specifically binds to the RNA scaffold binding domain (RSBD) of the modular polypeptide component; and (iii) a connecting sequence linking the TE and RS Including, The modular polypeptide components are linked as separate functional modules: (a) the RSBD; (b) a DNA-binding domain linked to (a), which recognizes a predetermined sequence in a target, wherein module (a) and the DBD are not found linked in nature; (c) an effector component linked to (b) for use in modifying a target, whereby the site-specific modification directed by the TE and the DBD occurs; and (d) optionally or if necessary, a first linker and / or a second linker linking the DBD to the effector component and RSBD, respectively; The nucleoprotein complex of claim 4, comprising:
6. 5. The nucleoprotein complex of claim 4, wherein the targeting nucleic acid provides two or more RNA scaffolds, which may be the same or different.
7. A nucleoprotein complex as described in claim 5, wherein the targeting nucleic acid or the NAC provides two or more RNA scaffolds, which may be the same or different.
8. wherein the DBD: (i) a polypeptide sequence of 70-75 amino acid residues or less, optionally wherein the DBD is a 30-mer or less, a 25-mer or less, a 20-mer or less, or a 15-mer or less; and / or (ii) The nucleoprotein complex of claim 1, wherein the DBD binds to a predetermined sequence of 3 to 6 nucleotides.
9. The nucleoprotein complex of claim 1 , wherein the DBD is capable of binding to more than one predetermined sequence.
10. 2. The nucleoprotein complex of claim 1, wherein the modular polypeptide components further comprise a nuclear localization signal (NLS) and / or an organelle localization signal, the NLS and / or the organelle localization signal being provided at the N- or C-terminus, optionally connected to additional sequences that aid in detection.
11. 2. The nucleoprotein complex of claim 1, wherein the effector component is selected from (i) an endonuclease or Fok1 nuclease domain for generating a double-strand break, (ii) a nickase, (iii) a transcriptional activator, (iv) a transcriptional repressor, (v) an epigenetic modulator enzyme, (vii) a recombinase, (viii) a transposase, (ix) an integrase, and (x) a nucleobase-modifying enzyme construct.
12. A nucleic acid or a combination of nucleic acids for the provision of one or more nucleoprotein complexes according to any one of claims 1 to 11 in a host cell.
13. 12. A method for modifying one or more target nucleic acid sequences using one or more nucleoprotein complexes or one or more nucleic acids for providing same in a host cell according to any one of claims 1 to 11, provided that the method is not a method for modifying human germline identity or a method of treatment performed on the human or animal body.
14. For use in a method for therapeutic treatment, comprising: (i) 12. A combination of at least one modular polypeptide component for a nucleoprotein complex according to any one of claims 1 to 11, or a polynucleotide encoding the same, and (ii) one or more targeting nucleic acids as defined in claim 1, or one or more polynucleotides encoding the same, capable of linking to said polypeptide component(s).