Methods and compositions

The use of a retargetable Group II intron with a selectable marker and meganuclease cleavage site in nucleic acid editing addresses off-target issues, ensuring precise and efficient editing in eukaryotic cells, particularly mammalian cells, by minimizing off-target effects and enabling effective selection.

WO2026074269A1PCT designated stage Publication Date: 2026-04-09FORGE GENETICS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing nucleic acid editing methods, such as CRISPR-Cas9, suffer from high off-target effects, difficulty in sequence design, and inefficient cleavage, leading to populations of cells with undesired edits and potential danger due to unedited cells persisting.

Method used

A method utilizing a retargetable Group II intron with a selectable marker and a nucleic acid sequence for nuclease cleavage, ensuring precise insertion and editing by designing the cleavage site to minimize off-target effects, using meganucleases for high specificity, and employing strategies to ensure the selectable marker is only expressed after correct integration.

Benefits of technology

Achieves a high percentage of correctly edited cells with minimal off-target edits, allowing for precise editing and efficient selection of desired edits in eukaryotic cells, particularly mammalian cells, reducing the risk of dangerous cell propagation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025052127_09042026_PF_FP_ABST
    Figure GB2025052127_09042026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to novel means for the ex vivo editing of mammalian cells. The invention provides methods and accompanying nucleic acids and compositions for putting the invention into practice.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Methods and Compositions

[0002] Field

[0003] The invention is in the field of nucleic acid editing.

[0004] Background

[0005] Various methods, including CRISPR Cas9 based technologies, have been used to edit the nucleic acid of cells, including mammalian cells.

[0006] These methods typically involve the generation of a double-stranded break in the target cell genome followed by repair of the break via non-homologous end joining, resulting in insertions and deletions (indels) to disrupt gene expression and are not generally suitable for precise modifications, such as an in-frame deletion or insertion.

[0007] These techniques such as CRISPR, have high rates of off-target effects, i.e. creating edits in the genome at undetermined, and not easily screenable locations. Tight regulatory control is required over the expression of the CRISPR components, and it can be difficult to find appropriate sequences to use as guide RNAs, making the CRISPR based methods complicated. Furthermore, CRISPR-Cas editing can trigger a p53- mediated DNA damage response, leading to cell cycle arrest and apoptosis which is the innate way that cells prevent propagation of cells with deleterious mutations. Off- target effects in p53 disrupts this check-point, allowing these potentially dangerous and oncogenic cells to propagate within the population.

[0008] Cleavage of the double-stranded DNA by the Cas endonuclease can also not be 100% effective, resulting in the propagation of unedited but viable cells, which can persist into the therapeutic population.

[0009] The result is a complicated method, and a population of cells that are often intended to be therapeutic, but which harbour amounts of off-target edited, and potentially dangerous cells.

[0010] The invention described below seeks to address all of the problems mentioned above. Summary of the invention

[0011] Whilst prior art populations of edited cells comprise high numbers of cells that carry undesired, and often dangerous off-target edits, the inventors of the present invention have devised a new method for the production of eukaryotic cell populations such as mammalian cell populations which comprise a high percentage of correctly edited cells vs the percentage of unedited cells or off-target edited cells. In some preferred instances every cell in the population comprises the desired, on-target edit, and no off- target edits.

[0012] The method of the invention is particularly suited for use with eukaryotic cells, such as mammalian cells, including human cells, and more particularly for the ex vivo or in vitro editing of eukaryotic cells for example the ex vivo or in vitro editing of mammalian cells such as human cells. In some embodiments of the invention described herein the various components required for putting the invention into practise are suitable for putting the invention into practice in a eukaryote, preferably a mammal, preferably a human and are not collectively suitable for putting the invention into practice in a prokaryote.

[0013] The method utilises a retargetable Group II intron which delivers a selectable marker operably linked to a promoter (herein terms a selectable marker promoter due to its association with the selectable marker) and a nucleic acid sequence that is a cleavage site for a nuclease. Group II introns are self-catalytic RIMA molecules, or ribozymes, capable of inserting into a target nucleic acid sequence. The Group II intron can be designed to specifically insert into a given target and the skilled person knows how to design regions of a Group II intron so that insertion into the desired target nucleic acid occurs. Self-insertion of the Group II intron into a target cell genome or other nucleic acid present in the target cell results in insertion of the selectable marker and the nucleic acid sequence that is a cleavage site for a nuclease into the target genome at a target site.

[0014] Group II introns are known, and include LI.LtrB (from Lactococcus lactis). The skilled person is aware of other Group II introns such as Ecl5 (from Escherichia coli), Rmlntl (from Sinorhizobium meliloti'), and TeI3c (from Thermosynechococcus elongatus). Preferably, the group II intron is LI.LtrB. Group II intron can be used provided it can insert into a target nucleic acid. The Group II intron may be naturally occurring (other than for the editing required to target the intron to the desired sequence), or may be a synthetic Group II intron. The nucleic acid cassette that comprises a sequence that encodes the Group II intron, the selectable marker and the nucleic acid sequence that is a cleavage site for a nuclease is herein termed an "insertion cassette" since it inserts into the target cell genome via the Group II intron. The nucleic acid cassette may be linked to other nucleic acid components.

[0015] Any one or more of the sequences described herein, or component parts of larger sequences, may be codon optimised for optimal function in a target cell. For example in some embodiments the nucleic acids described herein have been codon optimised for use in mammalian cells, for example human cells. For example, group II intron, selectable marker, sequence that is a cleavage site for a nuclease and / or IEP may be codon optimised for a eukaryotic cell, optionally a mammalian cell, optionally a human cell.

[0016] The sequence that is a cleavage site for a nuclease should not be a sequence that is found in the native target cell genome, or should be a site that is not found frequently in the target cell genome. By target cell genome we include the meaning of any nucleic acid within the target cell, for example chromosomal nucleic acid, mitochondrial nucleic acid, episomal nucleic acid, plasmid nucleic acid, artificial chromosome, etc. The skilled person is well able to screen genomic sequence information (including any nucleic acid sequences typically found in a given target cell) to identify suitable nuclease sites for use in any given situation, for example by searching screening publicly available sequence databases. Target cell genome and target nucleic acid may be used interchangeably. The target cell genome or target nucleic acid may be a target genetic locus of a chromosome of a cell.

[0017] As described herein, preferably the nucleic acid sequence that is a cleavage site for a nuclease is a cleavage site for a meganuclease. Meganucleases, also called homing endonucleases, have large recognition sites of approximately 14-25bp, meaning that they occur very infrequently in nature. In some embodiments then the sequence that is a cleavage site for a nuclease is a site that is at least 14 bp long, for example at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48 or 50 or more bp long.

[0018] Exemplary nucleases and the corresponding recognition sites are show in Figure 1 and include for example I-Scel, I-Ppol, I-Dmol, I-Njal, I-Dirl, I-Bmol, BneMS4ORFIP, F- Cphl, F-ECOT3I, F-ECOT5I, F-EcoT5II, F-EcoT5IV, F-PhiU5I, F-Scel, F-Scell, F-TevI, F- TevII, F-TevIII, F-TevIV, H-Drel, H-Drel, I-AabMI, I-AchMI, I-Anil, I-ApeKI, I-BanI, I- BasI, I-Bth0305I, I-Bthll, I-BthORFAP, I-Ceul, I-Chul, I-Cmoel, I-Cpal, I-CpaII, I- CpaMI, I-Crel, I-CreII, I-CsmI, I-Cvul, I-Ddil, I-GpeMI, I-Gpil, I-Gzel, I-GzeII, I- HjeMI, I-Hmul, I-HmuII, I-Llal, I-Ltrl, I-LtrWI, I-MpeMI, I-Msol, I-NanI, I-Nfil, I-Nitl, I-OmiII, I-Onul, I-PakI, I-PanMI, I-PfoP3I, I-PnoMI, I-PogTE7I, I-Porl, I-Scal, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-SecIII, I-SmaMI, I-SpomI, I-SscMI, I- Ssp6803I, I-TevI, I-TevII. I-TevIII. I-TslI. I-TsIWI, I-TspO61I, I-Twol, I-Vdil41I, - Aval, PI-BciPI, PI-HvoWI, PI-MgaI„ PI-MleSI, PI-MtuI, Pl-PabI, Pl-Pabll, Pl-Pful, PI- PfuII, Pl-Pkol, Pl-PkoII, PI-PspI, PI-PspI, Pl-Scal, Pl-Scel, PI-Tful, PI-TfuII, PI-Thyl, PI-Tlil, PI-Tlill, PI-Tmal, PI-TmaKI, Pl-Zba.

[0019] In some embodiments the nucleic acid sequence that is a cleavage site for a nuclease comprises at least 2, 3, 4, 5 or more sequences that are cleavage sites for nucleases, for example cleavage sites for meganucleases. The multiple sequences may be all the same sequence, i.e. multiple recognition sequences for the same endonuclease for example the same meganuclease, or two or more of the multiple sequences may be different, for example representing recognition sequences for two or more different nucleases, for example two or more different meganucleases.

[0020] Other suitable nucleic acid sequences that are cleavage sites for a nuclease include a nucleic acid recognition sequence for a corresponding TALEN, or a nucleic acid recognition sequence for a corresponding Zinc finger nuclease. However preferably the nucleic acid sequence that is a cleavage site for a nuclease is a site for a meganuclease.

[0021] It is also possible to use CRISPR technology with the invention, for example in those instances the sequence that is a recognition site for a nuclease is a sequence towards which a guide RNA is designed, targeting a corresponding nuclease such as a Cas protein. However, CRISPR / Cas technology is considered to not cleave with as high a fidelity as a meganuclease, and there are difficulties in designing appropriate guide sequences. It is preferred that one or more meganucleases are used, and that the insertion nucleic acid comprises one or more recognition sites for the one or more meganucleases.

[0022] In some embodiments the nuclease is not a Cas protein. In some embodiments the sequence that is a recognition site for a nuclease is not a sequence towards which a guide RNA has been devised. The sequence that is a cleavage site for a nuclease can be positioned anywhere within the Group II intron other than within the sequence that encodes the selectable marker, since expression of the selectable marker must be independent of whether the sequence that is a cleavage site for a nuclease has been cleaved or not. The sequence that is a cleavage site for a nuclease should also not be too close to the selectable marker sequence (and any corresponding promoter sequence). In the event that the site is cleaved and repaired by NHEJ rather than homology directed repair, it is important that the selectable marker is still expressed. As described below, expression from the selectable marker following break repair is indicative of a cell which has not undergone the desired editing event, and a cell in which the selectable marker is lost is indicative of the desired editing event. If the break is made too close to the promoter / coding sequence of the selectable marker and repaired by NHEJ, some resection of the break ends may occur, affecting the promoter / selectable marker sequence such that is incapable of being expressed.

[0023] For this reason the sequence that is a cleavage site for a nuclease should be position such that cleavage occurs at least Int, 2, 3, 4, 5, 6, 7, 8, 9, 10, 13, 14, 16, 20, 22, 24, 26, 28, 30, 32, 24, 26, 28, 30, 32, 34, 36, 38, 40 or more nucleotides away from the start of the promoter sequence or the start of the open reading frame or the end of the open reading frame of the selectable marker. In some instances it is preferred if the sequence that is a cleavage site for a nuclease should be position such that cleavage occurs at least Int, 2, 3, 4, 5, 6, 7, 8, 9, 10, 13, 14, 16, 20, 22, 24, 26, 28, 30, 32, 24, 26, 28, 30, 32, 34, 36, 38, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more nucleotides away from the start of the promoter sequence or the start of the open reading frame or end of the open reading frame of the selectable marker.

[0024] Any of the sequences that are provided as part of the Group II intron can be inserted anywhere in the Group II intron provided that it doesn't interfere with the RIMA structure such that it prevents insertion into the host nucleic acid, or significantly reduces retrotransposition efficiency. The skilled person knows or is easily able to determine suitable positions for the selectable marker and sequence that is a cleavage site for a nuclease in the Group II intron without affecting the Group II intron core function of insertion into the target nucleic acid. For example, in the case of the L. lactis LtrB intron, insertion of sequences in the loop region of domain IV is a preferred location. See for example Plante and Cousineau 2006 RA 12; 1980-1992. Preferably the size of the inserted sequences, i.e. the selectable marker sequence, selectable marker promoter, and the sequence that is a recognition site for a nuclease are kept as small as possible to avoid impacts on retrotransposition efficiency. In some embodiments the total size of the selectable marker, selectable marker promoter and the sequence that is a recognition site for a nuclease is Ikb or less than Ikb.

[0025] The insertion cassette can be provided to a cell in the form of DNA. Transcription of the DNA sequence that encodes the Group II intron from a DNA dependent RNA polymerase promoter, such as a T7 promoter, produces the corresponding RNA, which is self-catalytic and capable of insertion into the target cell genome, where second strand synthesis occurs via reverse transcription, with subsequence replacement of the inserted RNA strand with a second DNA strand. Accordingly in some embodiments the insertion cassette as described herein in operably linked to a promoter such as a DNA dependent RNA polymerase promoter, such as a T7 promoter, that allows transcription across the insertion cassette to generate an insertion cassette RNA transcript.

[0026] In some embodiments the insertion cassette promoter is a cell-type specific promoter, for example a eukaryotic cell-type specific promoter for example a mammalian celltype specific promoter. In some embodiments, the insertion cassette promoter is an Hl promoter.

[0027] The selectable marker is typically present in the insertion cassette operably linked to its own promoter capable of driving expression of the selectable marker once integrated into the target genome. There are some instances where insertion of the insertion cassette into the target genome occurs in such a way so as to place the selectable marker under the control of an endogenous promoter. However, typically the selectable marker will be present in the insertion cassette under the control a promoter that is also located in the insertion cassette. The skilled person will recognise that this means that when the insertion cassette is provided to the cell in the form of a DNA vector, for example a plasmid, expression of the selectable marker prior to insertion of the insertion cassette in to the target genome is possible. In this case the selectable marker is expressed and available for selection prior to the desired insertion event occurring. The selectable marker promoter can be any promoter capable of driving expression once in a given target cell. For example in some instances where the cell to be edited is a mammalian cell, the selectable marker promoter is a promoter that is functional in a mammalian cell. It is important that in all instances where the insertion cassette has inserted into the target cell nucleic acid, that the selectable marker is expressed, and is expressed to a detectable level. In some instances then the selectable marker promoter is a strong promoter. In some instances the selectable marker promoter is a constitutive promoter. In some instances the selectable marker promoter is a strong constitutive promoter.

[0028] Ideally the selectable marker would only be expressed once the insertion cassette is inserted into the target genome (whether at the correct target site or not) or other target nucleic acid such as a plasmid, i.e. where the insertion cassette is a DNA nucleic acid and is part of for example a plasmid, it is possible that expression from the selectable marker promoter of the selectable marker from the plasmid can occur, giving rise to cells that express the selectable marker prior to any insertion events. There are various strategies to mitigate this, all of which are compatible with the present invention.

[0029] For example in a first strategy the insertion cassette may be present on a non- replicative, non-selectable plasmid. In these cases, introduction of the plasmid into a cell may cause a transient expression of the selectable marker from the plasmid, but since the plasmid is not selected for, over a short period of time the plasmid would be lost from the cell. At this time point, any cells that continue to express the selectable marker should be doing so from a genome integrated insertion cassette (that as stated above comprises the Group II intron and nuclease site), or other target-nucleic acid integrated insertion cassette for example where the target nucleic acid is present on a plasmid rather than the genome.

[0030] In some strategies the selectable marker may be termed a retrotransposition-activated marker (RAM) since the marker is only expressed upon integration of the nucleic acid into the target sequence. For example, in a second strategy, the open reading frame that encodes the selectable marker comprises a removable intervening sequence arranged such that when the removable sequence is present expression from the selectable marker promoter does not produce a functional protein, and wherein removal of the removable intervening sequence allows expression of a functional protein or RNA. In some instances the removable intervening sequence comprises a Group I intron that disrupts expression of the selectable marker from the plasmid version of the insertion cassette. Group I introns are self-catalytic and once the insertion cassette (including the Group II intron and the selectable marker which in this embodiment comprises a Group I intron) is transcribed into RNA, or is otherwise in an RNA form, the Group I intron is able to self-excise, leaving an in-frame selectable marker capable of proper translation. When using a RAM, the marker must be antisense to the direction of the Group II intron. The Group I is in a sense orientation relative to the Group II intron, and therefore antisense relative to the marker. When the marker is transcribed from the plasmid, the Group I is transcribed in an antisense direction, and it cannot excise from the mRNA in this form. Only when it forms a part of the mRNA from expression of the Group II intron is it in its sense orientation, and it can then excise. In this case the marker is now intact but is itself in an antisense orientation, so only once it integrates can it then be transcribed from its own promoter without a Group I intron interrupting it.

[0031] A third strategy involves ensuring that the insertion cassette comprises the promoter and associated selectable marker in the antisense orientation - i.e. expression of the insertion cassette into an RNA transcript comprises within it the selectable marker in an antisense direction. This means that the selectable marker sequence present in the insertion cassette RNA transcript is not capable of translation and not capable of producing a functional selectable marker. Only when the insertion cassette is inserted into the target cell genome (or other target nucleic acid), and reverse transcribed, is the sense strand produced, and only at that point can transcription and translation of a functional selectable marker occur. This approach is suitable for use where the insertion cassette is provided to the target cell as part of a piece of RNA.

[0032] This "antisense" strategy can also be used in the context of delivering the insertion cassette in an RNA form complexed with proteins that are required for insertion of the insertion cassette into the cell genome (or other target nucleic acid), i.e. a ribonuclear complex, i.e. the invention provides an insertion cassette as set out herein that is an RNA and wherein the selectable marker is provided in antisense form so that expression of the selectable marker only occurs following integration into the genome (or other target nucleic acid). These latter two embodiments have advantages over the approach that utilises the Group I intron. Group I introns do not always splice out with 100% efficiency. This means that in some instances an insertion cassette, which comprises the Group II intron and the selectable marker sequence interrupted by the Group I intron can insert into the target genome (or other target nucleic acid), but be undetectable due to failure of the Group I intron to splice and produce the detectable selectable marker.

[0033] To allow the Group II intron to insert into the target genome (or other target nucleic acid), various protein activities have to be present to allow the Group II intron to invade the target sequence, integrate, and reverse transcribe. These are all activities that native Group II introns encode within themselves as part of a protein complex called an Intron Encoded Protein (IEP). Group II introns typically encode an IEP that also allows the Group II intron to self-excise and retarget a different portion of the genome. It is crucial that the Group II intron used in the invention, and any accessory proteins such as lEPs do not allow the self-excision of the Group II intron. Accordingly, in all embodiments, the Group II intron is not capable of self-excision and is not exposed to a protein complement that allows for self-excision. Self-excision would remove the selectable marker and the nuclease site, giving an outward phenotype of a desired cell in which a desired gene editing event has taken place (see further below), but would in actual fact be wild-type at the target locus, as the Group II intron inserts and then excises. In the invention, the only way that the Group II intron should be lost from the genome (or other target nucleic acid) is via desired homology directed repair, as set out elsewhere herein.

[0034] Accordingly, in addition to introducing to the target cell the insertion cassette, the additional functions required for insertion, but not excision, must also be present in the cell. However, these functions must be present only transiently, for example provided on a non-replicative plasmid or under the control of an inducible promoter, so that the ability for the Group II intron to insert is time limited, to prevent multiple insertion events occurring. The skilled person is able to select suitable means for the transient expression of these protein functions.

[0035] These functions can be provided to the cell in any means. The functions can be provided on a separate nucleic acid to the insertion cassette. The functions can be provided within the insertion cassette i.e. the insertion cassette also comprises a portion that encodes the one or more necessary accessory functions. The accessory functions can be provided on the same nucleic acid molecule as the insertion cassette, but not as an integral part of the insertion cassette. For example, the insertion cassette can be provided as part of a plasmid, and the plasmid may also comprise one or more portions that express one or more accessory proteins with accessory functions. Since these two functions are not linked, insertion of the insertion cassette into the target nucleic acid does not also result in insertion of the coding regions for the accessory protein functions.

[0036] The accessory functions can be provided by an intron-encoded protein (IEP). An example of a suitable IEP is LtrA. LtrA comprises a nucleotide sequence substantially as set out in SEQ ID NO: 156; or a sequence that is: a)at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NO: 156; b)about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% identical to SEQ ID NO: 156; and / or c)75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 156.

[0037] The accessory functions can be provided by an accessory protein that is encoded on the same nucleic acid as the intron, but that is not encoded by the intron. Accordingly, in some embodiments, the intron is part of a nucleic acid that further comprises a nucleotide sequence encoding an accessory protein, wherein the accessory protein is not encoded by the intron. In some embodiments, the accessory protein is LtrA. In some embodiments, the nucleotide sequence encoding the accessory protein is operably linked to a promoter. In some embodiments, the nucleotide sequence encoding the accessory protein is operably linked to a promoter that is a different promoter to a promoter that is operably linked to or comprised by the intron. In some embodiments, the nucleotide sequence encoding the accessory protein is operably linked to a promoter selected from the group comprising or consisting of: CMV promoter, an EFla promoter, an SV40 promoter, a UBC promoter, a PGK promoter, a CAG promoter, an SFFV promoter, an MSCV promoter.

[0038] In some embodiments, expression of the accessory protein causes: i) an RNA transcript from the insertion cassette promoter to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand; or ii) where the insertion cassette is an RNA molecule, the insertion cassette to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand.

[0039] The insertion cassette and the accessory proteins, such as the IEP, can be provided to the cell on a single nucleic acid molecule such as a single double stranded DNA molecule, single-stranded DNA molecule, or as an RNA molecule.

[0040] In some particular embodiments, the insertion cassette may be provided to the cell as an RNA complexed with the accessory proteins such as complexed with an IEP protein. The association between the accessory proteins such as the IEP protein and the RNA may be achieved by the use of an aptamer sequence that can be included in the RNA and which has binding affinity for one or more proteins that have the necessary function, such as the IEP protein. Regardless of the form in which the insertion cassette is provided to the cell, it is generally advantageous if the nucleic acid comprising the insertion cassette is not capable of replicating in the cell. Accordingly, in some embodiments, the nucleic acid comprising the insertion cassette does not comprise an origin of replication. In some embodiments, the nucleic acid comprising the insertion cassette does not comprise an origin of replication that is capable of driving replication of the nucleic acid in the target cell. In some embodiments, the nucleic acid comprising the insertion cassette comprises an origin of replication that is not capable of driving replication of the nucleic acid in the target cell. As will be understood, an origin of replication that is "capable of driving replication" can recruit protein complexes that replicate a nucleic acid, such as DNA, and initiate nucleic acid replication such that a second copy of the nucleic acid is generated. For example, where the nucleic acid comprising the insertion cassette is a plasmid, an origin of replication that is "capable of driving replication" can initiate rolling circle replication of the plasmid, leading to the generation of a second copy of the plasmid. In embodiments where the target cell is a eukaryotic cell - such as a mammalian cell - the nucleic acid comprising the insertion cassette comprises a bacterial origin of replication. In some embodiments, the bacterial origin of replication is a ColEl origin. In some embodiments, the bacterial origin of replication is a BBR1 origin. In embodiments where the target cell is a eukaryotic cell - such as a mammalian cell - the nucleic acid comprising the insertion cassette comprises a bacterial origin of replication and does not comprise a mammalian origin of replication.

[0041] In some embodiments, the nucleic acid comprising the insertion cassette further comprises a selectable marker that can be used during nucleic acid amplification and cloning, but that is not suitable for use in selecting the target cell or for maintaining the nucleic acid in the target cell. This selectable marker is not encoded by the intron. For example, where the target cell is a eukaryotic cell - such as a mammalian cell - the nucleic acid comprising the insertion cassette may comprise a bacterial selectable marker that is suitable for use in selecting a bacterial cell. In some embodiments, the bacterial selectable marker is a bacterial antibiotic resistance marker. In some embodiments, the bacterial antibiotic resistance marker is selected from the group comprising or consisting of: an ampicillin resistance cassette, a chloramphenicol resistance cassette, a kanamycin resistance cassette, a streptomycin resistance cassette, a tetracycline resistance cassette, a spectinomycin resistance cassette, an erythromycin resistance cassette, a gentamycin resistance cassette, a rifampicin resistance cassette, a trimethoprim resistance cassette, a carbenicillin resistance cassette, a neomycin resistance cassette, a G418 / geneticin resistance cassette, a hygromycin resistance cassette, a paromomycin resistance cassette, a tobramycin resistance cassette, a puromycin resistance cassette, a blasticidin resistance cassette, a zeocin resistance cassette, a penicillin resistance cassette, and a lincomycin resistance cassette. In some embodiments, the bacterial selectable marker is a bacterial auxotrophic marker. In some embodiments, the bacterial auxotrophic marker is selected from the group comprising or consisting of: hisA, hisB, hisD (histidine biosynthesis), leuB (leucine biosynthesis), ilvA (isoleucine / valine biosynthesis), metE (methionine biosynthesis), argA, argE (arginine biosynthesis), thrC (threonine biosynthesis), pyrF (UMP biosynthesis (uracil pathway)), purA (adenine biosynthesis), thyA (thymidine biosynthesis), galK (galactose utilization), and araA (arabinose utilization). In some instances the selectable marker is a marker that is capable of selection in both mammalian and bacterial cells, but is under the control of a promoter that operable in a bacterial cell but is not operable in a eukaryotic cell such as a mammalian cell.

[0042] Expression of the accessory protein functions, such as expression of an IEP causes: i) an RIMA transcript from the insertion cassette promoter to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand; or ii) where the insertion cassette is an RNA molecule, the insertion cassette to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand.

[0043] The insertion cassette may also be provided to the cell in the form of an editing complex that comprises an RNA molecule complexed with one or more accessory proteins that comprise the functions to allow insertion of the RNA into the target nucleic acid, for example complexed with an IEP as set out herein. The RNA molecule may be a transcript from a DNA version of the editing cassette as set out herein. The association between the RNA and the accessory proteins can be achieved by any suitable means, for example in some instances the editing cassette comprises a portion that when transcribed produces a portion of the RNA that is an RNA aptamer that has specificity to the one or more accessory proteins. The invention therefore also provides this editing complex.

[0044] Once the target cell or cell population has been exposed to the insertion cassette, a first screening step selects for cells into which the insertion cassette, i.e. the nucleic acid that encodes the Group II intron, selectable marker and the nucleic acid sequence that is a cleavage site for a nuclease have inserted into the target nucleic acid, by determining the presence of the selectable marker. Group-II introns are relatively promiscuous, and at this point there will be some cells that have been selected on the basis of expression of the selectable marker in which the Group-II intron and associated markers have inserted at an undesired location. Since the method is an ex vivo or in vitro method this is not a concern since the off-target cells are not yet within a target subject.

[0045] Once cells that express the selectable marker have been identified, they are collected for further processing.

[0046] The selectable marker can be any protein or RIMA capable of allowing the selection of a population of cells that express the selectable maker. For example any protein or RNA that confers a physically detectable phenotype on a target cell that expresses said marker. The RNA or protein may directly affect physical properties of the cell, for example as set out herein a preferred selectable marker is a fluorescent protein. However, selectable markers may also be indirect selectable markers. For example a selectable marker may be an RNA aptamer that binds to a target cell protein, subsequently affecting physical properties of the cell.

[0047] Preferably the physically detectable phenotype is a phenotype that can be detected by a cell sorter, for example by a FACS machine. Preferably the selectable maker is a protein or fragment thereof that has the ability to produce light, for example to produce light of a particular wavelength after interrogation with a laser. In other embodiments the selectable marker affects the size and / or morphology of the cells in such a way so as to allow the cells to be distinguished on the basis of those that do and those that do not express the selectable marker.

[0048] Preferably the selectable maker is a protein that is a fluorescent protein or a bioluminescent protein, or a protein or antigen towards which a fluorescently labelled antibody or binding domain thereof, or aptamer, can bind, so as to allow detection of the fluorescence or bioluminescence and capture of the fluorescent or bioluminescent cells via fluorescence activated cell sorting (FACS). The antibody or fragment thereof can be any fragment that is capable of binding to the selectable marker protein or antigen, and that can be labelled with a fluorescent entity. For example the antibody or binding fragment thereof can be an ScFv, Fab etc. Such antibody fragments are well known in the art, as are suitable fluorescent moieties that can be conjugated to said antibody or fragment thereof. The fluorescent protein may be any fluorescent protein, for example may be a blue fluorescent protein, green fluorescent protein, yellow fluorescent protein, and red fluorescent protein. The blue fluorescent protein may be selected from the group comprising or consisting of: TagBFP, mTagBFP2, EBFP, EBFP2, Azutire, mKamalal, cyan fluorescent protein, ECFP, Cerulian, CyPet, mTurquoise, mTurquoise2, and AmCyanl. The green fluorescent protein may be selected from the group comprising or consisting of: miniGFPl, GFP, EGFP, mNeonGreen, and mGFP. The yellow fluorescent protein may be selected from the group comprising or consisting of: YFP, Citrine, Venue, and YPet. The red fluorescent protein may be selected from the group comprising or consisting of: mCherry, mCardinal, mScarlet, DsRed, mOrange, mRaspberry, mFruit, mK, TagRFP, mKate, mRuby, FusionRed, and DsRed-Express. There are many more commercially available fluorescent proteins that are compatible with the present invention.

[0049] Preferably the sequence that encodes the protein is as small as possible. In preferred embodiments "mini" versions of fluorescent proteins, such as miniGFPl as set out in Liang et al 2022 Front Bioeng Biotechnol 10: 1039317.

[0050] The bioluminescent protein may be selected from the group comprising or consisting of: luciferase, luxA, and luxB.

[0051] Where the selectable marker is a protein or antigen towards which an antibody can bind, for example a fluorescently labelled antibody - the protein or antigen can be any protein or antigen, provided it is not already present within the cell. The protein or antigen chosen to be used as a selectable marker and a corresponding antibody or fragment thereof should show no or minimal background fluorescence in a cell which does not comprise the insertion cassette. The skilled person knows how to screen sequence information of a given target cell for the presence of a particular protein or antigen sequence, and is also able to conduct routine experiments to determine an appropriate protein / antigemantibody pair for use in a given target cell.

[0052] The fluorescent moiety that is conjugated to an antibody may be any fluorescent moiety, for example preferably a fluorescent moiety compatible with FACS. There are many such moieties commercially available.

[0053] Preferably the selectable marker is not an antibiotic resistance gene. Selection of cells into which the insertion cassette has inserted can therefore be performed using FACS or a flow cytometer directly on cells that express a fluorescent protein or a bioluminescence protein, or on cells which have been exposed to a fluorescently labelled antibody with binding affinity to the selectable marker, where the selectable marker is a protein or antigen to which a fluorescently labelled antibody can bind.

[0054] In one embodiment, the insertion cassette is comprised by a nucleic acid further comprising: a) a nucleotide sequence encoding an accessory protein, wherein the nucleotide sequence encoding the accessory protein is not comprised by the insertion cassette as provided herein; b) an origin of replication that is not capable of driving replication of the nucleic acid in the target cell as provided herein; and c) a bacterial selectable marker as provided herein.

[0055] In one aspect, the invention provides a nucleic acid comprising: a) an insertion cassette as provided herein; b) a nucleotide sequence encoding an accessory protein, wherein the nucleotide sequence encoding the accessory protein is not comprised by the insertion cassette as provided herein; c) an origin of replication that is not capable of driving replication of the nucleic acid in the target cell as provided herein; and d) a bacterial selectable marker as provided herein.

[0056] The selected cells (i.e. those that have been selected on the basis of expression of the selectable marker) are then exposed to two components simultaneously: 1) an editing cassette that comprises regions of nucleic acid that are homologous to the target region, and can also comprise an intervening sequence that is to be inserted into the target region; and 2) the nuclease that corresponds to the nucleic acid sequence that is a cleavage site for a nuclease. The regions of homology to the target region are homologous to regions either side of the inserted insertion cassette, i.e. are homologous to regions either side of the entire Group II intron, selectable marker and nucleic acid sequence that is a cleavage site for a nuclease.

[0057] The skilled person is able to design the presence or absence or sequence of the intervening sequence that is located between the two regions that are homologous to the target region so as to elicit the desired editing event. For example the targeting event might be a deletion, an insertion, an inframe deletion, an inframe insertion, a mutation or any other editing event.

[0058] The editing cassette can be provided on a separate nucleic acid molecule to the nucleic acid that encodes the nuclease, or can be provided on the same nucleic acid.

[0059] The editing cassette can be provided on a separate DNA molecule to the nucleic acid that encodes the nuclease, or can be provided on the same nucleic acid.

[0060] The editing cassette may be provided as a single stranded DNA molecule, or may be provided as a double-stranded DNA molecule. The nuclease may be provided in the form of a nucleic acid from which the nuclease is transcribed and / or translated in the target cell; or may be provided directly as a protein molecule.

[0061] In some embodiments the editing cassette and nuclease is provided as a nucleic acid protein complex. For example in some embodiments the editing cassette is provided as a single stranded DNA molecule which also comprises an aptamer sequence that is capable of binding to the relevant nuclease, such that the single stranded editing cassette and nuclease are provided as a single entity. This embodiment is considered to be advantageous since it targets the editing cassette to the site of the nuclease cleavage, driving the desired homology directed repair over NHEJ, discussed below.

[0062] In some embodiments, the nucleic acid comprising the editing cassette does not comprise an origin of replication. In some embodiments, the nucleic acid comprising the editing cassette does not comprise an origin of replication that is capable of driving replication of the nucleic acid in the target cell. In some embodiments, the nucleic acid comprising the editing cassette comprises an origin of replication that is not capable of driving replication of the nucleic acid in the target cell. In some embodiments where the target cell is a eukaryotic cell - such as a mammalian cell - the nucleic acid comprising the editing cassette comprises a bacterial origin of replication. In some embodiments, the bacterial origin of replication is selected from the group comprising or consisting of: a ColEl origin, a BBR1 origin, a pl5A origin, a pSClOl origin, an R6K origin, an RK2 / RP4 origin, and an RSF1010 origin. In some embodiments, the bacterial origin of replication is a ColEl origin. In some embodiments, the bacterial origin of replication is a BBR1 origin. In some embodiments where the target cell is a eukaryotic cell - such as a mammalian cell - the nucleic acid comprising the editing cassette comprises a bacterial origin of replication and does not comprise a mammalian origin of replication. In some embodiments, the nucleic acid comprising the editing cassette further comprises a selectable marker that can be used during nucleic acid amplification and cloning, but that is not capable of being selected in the target cell. For example, where the target cell is a eukaryotic cell - such as a mammalian cell - the nucleic acid comprising the editing cassette may comprise a bacterial selectable marker that can be used when the nucleic acid is being amplified in a bacterial cell. In some embodiments, the bacterial antibiotic resistance marker is selected from the group comprising or consisting of: an ampicillin resistance cassette, a chloramphenicol resistance cassette, a kanamycin resistance cassette, a streptomycin resistance cassette, a tetracycline resistance cassette, a spectinomycin resistance cassette, an erythromycin resistance cassette, a gentamycin resistance cassette, a rifampicin resistance cassette, a trimethoprim resistance cassette, a carbenicillin resistance cassette, a neomycin resistance cassette, a G418 / geneticin resistance cassette, a hygromycin resistance cassette, a paromomycin resistance cassette, a tobramycin resistance cassette, a puromycin resistance cassette, a blasticidin resistance cassette, a zeocin resistance cassette, a penicillin resistance cassette, and a lincomycin resistance cassette. In some embodiments, the bacterial selectable marker is a bacterial auxotrophic marker. In some embodiments, the bacterial auxotrophic marker is selected from the group comprising or consisting of: hisA, hisB, hisD (histidine biosynthesis), leuB (leucine biosynthesis), ilvA (isoleucine / valine biosynthesis), metE (methionine biosynthesis), argA, argE (arginine biosynthesis), thrC (threonine biosynthesis), pyrF (UMP biosynthesis (uracil pathway)), purA (adenine biosynthesis), thyA (thymidine biosynthesis), galK (galactose utilization), and araA (arabinose utilization). In some embodiments the selectable marker is under the control of a promoter that is operable in bacterial cells but which is not operable in eukaryotic cells such as mammalian cells.

[0063] Once expressed in the selected cells, the nuclease will cleave the nucleic acid sequence that is a cleavage site for the nuclease, wherever it is present in the genome (or other target nucleic acid) resulting in a double-strand break.

[0064] Double-strand breaks tend to have three fates in mammalian cells: 1) the break is not repaired, triggering cell death by apoptosis; 2) the break is repaired by non- homologous end joining (NHEJ), and the two fragments are simply ligated back together though in some instances the ends can be resected prior to ligation; or 3) the break is repaired by homology directed repair (HDR) using the homologous regions on the editing cassette. These three different fates, in a cell that has a correctly targeted insertion cassette or in a cell that has an off-target insertion cassette, can be readily distinguished as follows:

[0065] In a cell with a correctly targeted insertion cassette, the cell will comprise the selectable marker and the nucleic acid sequence that is a cleavage site for the nuclease, flanked by regions that are homologous to regions in the editing cassette, in which case, cleavage of the nuclease site by the nuclease will result in a double strand break that: i) is repaired by HDR, using the editing cassette for repair. In this instance the selectable marker is lost via recombination and the editing cassette sequence replaces the native sequence, resulting in the desired editing event. These cells are viable since the break has been repaired, but do not express the selectable marker. Where the selectable marker is a fluorescent protein or a bioluminescent protein, these cells do not fluoresce or luminesce, and can be selected for this feature using FACS. These are the desired cells with the desired edit(s) and can be used for a range of purposes, including administration to a subject for therapeutic purposes. ii) is repaired by NHEJ. In this case, the cell is viable since the break has been fixed. However, the selectable marker is still present at the target locus. These cells then will still express the selectable marker and be selected as such and discarded as being undesirable. Where the selectable marker is a fluorescent protein or a bioluminescent protein, these cells fluoresce or luminesce, and can be selected for this feature using FACS. These are undesired cells which do not have the desired edit(s) and can be discarded. iii) is not repaired. These cells will not be viable and will be lost from the population.

[0066] In a cell with an incorrectly targeted insertion cassette i.e. the Group II intron has inserted at an off-target site, the cell will comprise the selectable marker and the nucleic acid sequence that is a cleavage site for the nuclease, but in this instance the flanking regions that are homologous to regions in the editing cassette are not present (since the regions of homology will have been designed to match the desired target locus rather than the undesired off-target locus). In this case, cleavage of the nuclease site by the nuclease will result in a double strand break that: i) cannot be repaired by HDR, since the regions of the editing cassette do not match the regions flanking the insertion cassette of the off-target locus. Accordingly the selectable marker is not lost by recombination. ii) is repaired by NHEJ. As above, the cell is viable since the break has been fixed. However, the selectable marker is still present at the target locus. These cells then will still express the selectable marker and be selected as such and discarded as being undesirable. Where the selectable marker is a fluorescent protein or a bioluminescent protein, these cells fluoresce or luminesce, and can be selected for this feature using FACS. These are undesired cells which do not have the desired edit(s) and can be discarded. iii) is not repaired. These cells will not be viable and will be lost from the population.

[0067] It is clear then that the desired population of cells that comprise the correct editing event, will be viable (since the break has been repaired), and will not express the selectable marker (e.g. the fluorescent protein or bioluminescent protein) since the break will have been fixed using HDR. Viable cells that are fluorescent or luminescent are undesired and can be discarded.

[0068] In this way it is straightforward to remove cells in which the insertion cassette inserted at the incorrect locus, and to remove cells in which the editing event did not occur.

[0069] Accordingly the method of the invention can also include the step of selecting cells which are viable but which do not express the selectable marker. Cells which are viable and that do express the selectable marker are unwanted and should be discarded.

[0070] It is also clear that to achieve the desired edit, the double strand break should be repaired by HDR and not by NHEJ. In some embodiments the various nucleic acids, cassettes and methods described herein are optimised and formulated in a way so as to increase the chance of a double-stranded break being repaired by HDR rather than NHEJ. For example as set out above, the editing cassette can be supplied as a single stranded DNA molecule complexed with the appropriate nuclease protein for example via an aptamer sequence that can be included in the editing cassette. In some instances it is considered that nucleic acid sequences that are cleavage sites for a nuclease that generate larger overhangs after cleavage drives HDR over NHEJ. In some instances the GC content of the overhangs following cleavage is considered to be important. In some instances a meganuclease that cuts a short distance from the recognition site, for example lObp downstream or upstream of the recognition site is preferred as it allows for more flexibility in the design of the overhangs.

[0071] The invention provides a method of editing the nucleic acid for example the genome of a target cell for example a target cell that is a mammalian cell such as a human cell, and also provides each of the various nucleic acid component parts for putting the invention into practice.

[0072] It will be clear to the skilled person that a key feature of the present invention is the ability to specifically target and edit nucleotides of a host cell genome, for example a mammalian genome. However it is also possible to edit target nucleic acid sequences that are present on extra-chromosomal genetic elements, such as plasmids or minichromosomes. Accordingly the target sequence may be present in the genome for example on a chromosome, or may be present on a plasmid, cosmid, minichromosome or other episomally maintained element.

[0073] For example the invention provides a nucleic acid that is an insertion cassette as set out above. The insertion cassette comprises a first portion that encodes a Group II intron and also a sequence that is or encodes a selectable marker and a nucleic acid sequence that is a cleavage site for a nuclease. Preferences for each of these elements are as described elsewhere herein.

[0074] The relative location of the sequence that is or encodes a selectable marker and the nucleic acid sequence that is a cleavage site for a nuclease can be present in the insertion cassette in any order. For example the sequence that is or encodes a selectable marker may be 5' to the nucleic acid sequence that is a cleavage site for a nuclease, or the sequence that is or encodes a selectable marker may be 3' to the nucleic acid sequence that is a cleavage site for a nuclease. Whatever the relative positioning, the sequences are positioned within the Group II intron such that upon insertion of the Group II intron into the target cell genome / nucleic acid (whether at the intended target site or at an off-target site), the sequence that is or encodes a selectable marker and the nucleic acid sequence that is a cleavage site for a nuclease are both inserted into the target cell genome / nucleic acid.

[0075] As set out elsewhere, and as will be apparent to the skilled person, the Group II intron is designed to have a portion of sequence that is able to hybridise with a portion of the target sequence within the target cell genome. The skilled person knows how to design Group II introns to target particular sites. The invention also provides an editing cassette as set out herein. The editing cassette that comprises regions of nucleic acid that are homologous to the target region, and can also comprise an intervening sequence that is to be inserted into the target region. Preferences for features of the editing cassette are as set out elsewhere herein.

[0076] The editing cassette nucleic acid comprises at least one portion that is a region of homology to the target nucleic acid. In some instances the region of homology comprises an intervening sequence that is to be inserted into the target region, in which case the editing cassette can be said to have two regions of homology. The regions of homology may be contiguous with a target sequence, such that a first and second homologous sequence are homologous to two regions of target nucleic acid that are contiguous with one another. In other instances the two regions of homology are homologous to regions which are not contiguous with one another. In this latter case, following homology directed repair, the section of the target nucleic acid that is in between the two regions of homology will be lost, i.e. generating a deletion. The regions of homology have to have sufficiently high homology so as to allow the editing cassette to be used to repair the double stranded break using homology directed repair.

[0077] The design of editing cassettes for use in editing a cell by the repair of a double strand break by homology directed repair is well known by the skilled person and the skilled person will be able to design such cassettes that result in any desired type of editing event, for example an in-frame deletion, other deletion, gene knockout, gene knockin, promoter replacement, single point mutation etc.

[0078] The editing cassette may be provided as a circular or linear nucleic acid. The editing cassette may be provided as a single stranded nucleic acid or a double stranded nucleic acid. The editing cassette may be a DNA or an RIMA. In some embodiments the editing cassette is provided as part of a larger circular double stranded DNA molecule such as a plasmid. In other instances the editing cassette is provided as a single stranded DNA molecule.

[0079] The editing cassette may comprise a further portion that is a nucleic acid that is an aptamer with binding specificity for a nuclease, for example for the nuclease that corresponds to the nuclease cleavage site that is present in the insertion cassette. In some instances the editing cassette is a single stranded DNA molecule that comprises a portion that is an aptamer with binding specificity to a nuclease, for example preferably to a meganuclease as set out herein. The invention also provides a nuclease expressing cassette. The nuclease expressing cassette comprises a sequence that encodes a nuclease operably linked to a promoter. Preferences for the nuclease are as set out elsewhere herein and is preferably a meganuclease. The nuclease expressing cassette may encode more than one different nuclease, i.e. for use in instances where the insertion cassette comprises sequences that are cleavage sites for more than one different nuclease. The nuclease expressing cassette may be provided as a circular or linear nucleic acid. The nuclease expressing cassette may be provided as a single stranded nucleic acid or a double stranded nucleic acid. The editing cassette may be a DNA or an RIMA. In some embodiments the nuclease expressing is provided as part of a larger circular double stranded DNA molecule such as a plasmid. In other instances the nuclease expressing is provided as a single stranded DNA molecule.

[0080] The promoter of the nuclease expressing cassette may be any promoter, provided it is operable in a mammalian cell since the nuclease expressing cassette will be inserted into a mammalian cell for expression. In some instances the promoter is: a) an inducible promoter; b) a constitutive promoter; c) a eukaryotic promoter; d) a mammalian promoter; e) a human promoter; and / or f) a promoter that is not operable in a prokaryote.

[0081] The nuclease that is expressed from the nuclease expressing cassette may be a fusion protein comprising the nuclease fused to a nuclear localisation signal (NLS) to target the nuclease to the nucleus. In some embodiments, the nuclear localisation signal is selected from the group comprising or consisting of: SV40 NLS, c— Myc NLS, p53 NLS, nucleoplasmin NLS, yeast GAL4 NLS, Xenopus nucleoplasmin NLS, adenovirus E1A NLS, polyoma virus large T antigen NLS. In some embodiments, the nuclear localisation signal (NLS) is a mammalian nuclear localisation signal (NLS) or is a NLS from a non-mammalian cell but which is operable in a mammalian cell. In some embodiments, the mammalian nuclear localisation signal (NLS) is selected from the group comprising or consisting of: SV40 NLS, c-Myc NLS, p53 NLS, nucleoplasmin NLS. In some embodiments, the nuclear localisation signal has an amino acid sequence selected from the group comprising or consisting of: SEQ ID NO: 160 (SV40 Large T antigen NLS (PKKKRKV)), SEQ ID NO: 161 (c-Myc NLS (PAAKRVKLD)), SEQ ID NO: 162 (p53 NLS (PQPKKKPL)), SEQ ID NO: 163 (nucleoplasmin NLS (KRPAATKKAGQAKKKK)). In some embodiments, the NLS from a non-mammalian cell but which is operable in a mammalian cell is selected from the group comprising or consisting of: yeast GAL4 NLS, Xenopus nucleoplasmin NLS, adenovirus E1A NLS, polyoma virus large T antigen NLS. In some embodiments, the nuclear localisation signal or (NLS) from a non-mammalian cell but which is operable in a mammalian cell has an amino acid sequence selected from the group comprising or consisting of: SEQ ID NO: 164 (yeast GAL4 NLS (MKLKRAARRRRRKR)), SEQ ID NO: 165 (adenovirus E1A NLS (KRPRP)), SEQ ID NO: 166 (polyoma virus large T antigen NLS (VSRKRPRP)).

[0082] In some instances the editing cassette and the nuclease expressing cassette are part of the same nucleic acid molecule. For example the invention provides an editing and nuclease expressing cassette that comprises the editing cassette of the invention and the nuclease expressing cassette of the invention, both present on the nucleic acid molecule, for example both present in the same double stranded DNA molecule, for example part of the same plasmid. The editing cassette and the nuclease expressing cassette do not need to be contiguous with one another and may be located in different regions of the nucleic acid, for example different regions for a plasmid. The editing and nuclease expressing cassette may comprise more than one different nuclease expressing cassette.

[0083] The invention also provides an editing complex. The editing complex comprises both the editing cassette and the nuclease protein. In this way the editing cassette, which has the template for double strand break repair, is targeted to the site of the double strand break, as it accompanies the nuclease to the site for cleavage. In this way repair by HDR can be promoted over repair by the undesired NHEJ. The editing complex may comprise an editing cassette that is a single stranded DNA molecule complexed with a nuclease as set out herein, for example complexed with a meganuclease as set out herein. The association between the editing cassette nucleic acid and the nuclease protein can be achieved by any suitable means. In some instances the components are complexed by the use of an aptamer. An aptamer is a nucleic acid that has binding specificity to a particular target in a similar way that an antibody has binding specificity to an antigen. The editing cassette may comprise a further nucleic acid portion that is an aptamer that has binding specificity to a particular nuclease for example to a particular meganuclease. In some instances, for example where the insertion cassette comprises more than one sequence that is a cleavage site for more than one different nuclease, the editing cassette that is part of the editing complex can comprise more than one aptamer, with each aptamer having binding specificity to a different nuclease for example to a different meganuclease. The nuclease protein that is part of the editing complex may be fused to an NLS. For example in some embodiments the nuclease is a meganuclease fused to an NLS.

[0084] The design of aptamers to have specificity to particular targets is well known in the art and poses no difficulties to the skilled person.

[0085] Any of the nucleic acid cassettes or nucleic acids described herein may be provided in any format. For example in some instances any of the nucleic acid cassettes or nucleic acids of the invention are provided as part of a plasmid, a phagemid, a cosmid, a viral genome or fragment thereof, or an artificial chromosome.

[0086] The invention also provides various mammalian nucleic editing systems comprising various combinations of the nucleic acid cassettes described herein. For example in some embodiments the system comprises any two or more of:

[0087] The insertion cassette according to the invention;

[0088] The insertion complex according to the invention;

[0089] The editing cassette according to the invention;

[0090] The nuclease expressing cassette according to the invention;

[0091] The editing complex according to the invention; and / or

[0092] The editing and nuclease expressing cassette according to the invention.

[0093] For example in some embodiments the mammalian nucleic acid editing system comprises:

[0094] The insertion cassette according to the invention; and i) the editing cassette according to the invention and the nuclease expressing cassette according to the invention; ii) the editing and nuclease expressing cassette according to the invention; or iii) the editing complex according to the invention.

[0095] In some embodiments the mammalian nucleic acid editing system of the invention comprises:

[0096] The insertion complex of according to the invention; and i) the editing cassette according to the invention and the nuclease expressing cassette according to the invention; ii) the editing and nuclease expressing cassette according to the invention; or iii) the editing complex according to the invention.

[0097] In some embodiments the mammalian nucleic acid editing system of the invention comprises:

[0098] A nucleic acid comprising an insertion cassette of the invention and an nucleotide sequence encoding an accessory protein of the invention, wherein the nucleic acid encoding the accessory protein is not comprised by the insertion cassette; and i) the editing cassette according to the invention and the nuclease expressing cassette according to the invention; ii) the editing and nuclease expressing cassette according to the invention; or iii) the editing complex according to the invention.

[0099] The invention also provides various kits for putting the editing methods of the invention into practice. For example the kits of the invention may comprise various combinations of the nucleic acid cassettes described herein. For example in some embodiments the kit comprises any two or more of:

[0100] The insertion cassette according to the invention;

[0101] The insertion complex according to the invention;

[0102] The editing cassette according to the invention;

[0103] The nuclease expressing cassette according to the invention;

[0104] The editing complex according to the invention; and / or

[0105] The editing and nuclease expressing cassette according to the invention.

[0106] For example in some embodiments the kit comprises:

[0107] The insertion cassette according to the invention; and i) the editing cassette according to the invention and the nuclease expressing cassette according to the invention; ii) the editing and nuclease expressing cassette according to the invention; or iii) the editing complex according to the invention.

[0108] In some embodiments the kit of the invention comprises:

[0109] The insertion complex of according to the invention; and i) the editing cassette according to the invention and the nuclease expressing cassette according to the invention; ii) the editing and nuclease expressing cassette according to the invention; or iii) the editing complex according to the invention.

[0110] In some embodiments the kit of the invention comprises:

[0111] A nucleic acid comprising an insertion cassette of the invention and an nucleotide sequence encoding an accessory protein of the invention, wherein the nucleic acid encoding the accessory protein is not comprised by the insertion cassette; and i) the editing cassette according to the invention and the nuclease expressing cassette according to the invention; ii) the editing and nuclease expressing cassette according to the invention; or iii) the editing complex according to the invention.

[0112] The skilled person will appreciate that the mammalian nucleic acid editing system and the kit should be suitable for putting the editing methods of the invention into practice in a mammalian cell.

[0113] The invention also comprises various cells that comprise one or more of the nucleic acids, such as cassettes, of the invention.

[0114] Accordingly the invention provides a cell comprising any one or more of :

[0115] The insertion cassette according to the invention;

[0116] The insertion complex according to the invention;

[0117] The editing cassette according to the invention;

[0118] The nuclease expressing cassette according to the invention;

[0119] The editing complex according to the invention; and / or

[0120] The editing and nuclease expressing cassette according to the invention.

[0121] The cell can be any cell. For example in some instances the cell is a prokaryotic cell such as a bacterial cell which acts as a cloning and replication vector for the nucleic acids of the invention.

[0122] In other instances the cell is a eukaryotic cell, preferably a mammalian cell, preferably a human cell. As set out above, the invention provides methods for editing the nucleic acid of a cell, and in particular provides methods for editing the nucleic acid of a mammalian cell such as a human cell. A method of the invention is described above.

[0123] The invention also provides a method of editing a target nucleic acid present in a target cell, wherein the method comprises:

[0124] (a) providing a target cell or a plurality of target cells;

[0125] (b) introducing into the target cell or plurality of target cells one or more insertion cassettes according to the invention; and

[0126] (c) incubating the target cell or plurality of target cells under conditions that permit integration of the insertion cassette into the target nucleic acid.

[0127] Preferences for the target cells, and the insertion cassette are as described elsewhere herein. For example the target cell may be a human cell. The insertion cassette may be present on a DNA plasmid, or the insertion cassette may be a single stranded DNA molecule that may be complexed with accessory proteins.

[0128] The method of the invention may include a further step, step (d), comprising detecting the presence of the selectable maker, wherein detection of the selectable marker indicates that the insertion nucleic acid is present in the target cell, and selecting those cells in which the presence of the selectable marker is detected for further processing. Cells in which the selectable marker have not been detected can be discarded.

[0129] Preferences for the selectable marker are set out elsewhere herein. For example preferably the selectable marker is a fluorescent protein and detecting the selectable marker and selection of cells which express the selectable marker is accomplished using FACS.

[0130] As set out elsewhere, depending on what the selectable marker is, other means such as the use of flow cytometry, microscopy, spectrophotometry and / or visual inspection may be used.

[0131] As set out above in the explanation as to how the method works, cells which express the selectable marker are cells in which the insertion cassette has inserted into the nucleic acid. This insertion may be in to the desired target site, or in an off-target site. The method of the invention therefore can comprise a further step, step (e) which utilises the nuclease site present in the insertion cassette along with the editing cassette to select for those cells which both have the insertion cassette at the correct target region, and which undergo HDR so as to result in a correctly edited cell.

[0132] In step (e), the selected cells are exposed to both i) the corresponding nuclease which can cleave the nuclease site present in the insertion cassette which is now inserted somewhere in the target nucleic acid, and ii) the editing cassette which is used for subsequent repair of the cleaved nuclease site by HDR.

[0133] Step (e) therefore comprises introducing into the cells selected for further processing in step (d): i) the nuclease expressing cassette of the invention and the editing cassette of the invention; ii) the editing and nuclease expressing cassette of the invention; or iii) the editing complex of the invention.

[0134] The editing complex of the invention (which as set out above is a single stranded NDA that has regions of homology to the target sequence and which is complexed to the corresponding nuclease protein, for example complexed via an aptamer / protein interaction) is considered to have particular benefits since the editing cassette, which is the template for HDR, is present in the vicinity of the nuclease as it cuts the target nucleic acid, driving repair by HDR rather than NHEJ.

[0135] Where the nuclease is to be expressed within the target cell from a nuclease expressing cassette, the cells are incubated under conditions so as to allow expression of the nuclease. Where the nuclease is provided as a protein as part of the editing complex, it is not necessary to express the nuclease from within the target cell(s).

[0136] The cells are incubated under conditions that permit HDR of the double stranded break produced by the nuclease, using the editing cassette as the template for repair.

[0137] As set out above, the various cassettes used in the method may be provided on separate nucleic acid molecules, or some cassettes may be provided as part of the same nucleic acid molecule. The cassettes may be double stranded or single stranded; DNA or RNA; and / or circular or linear. The method may further comprise a step (f) of selecting for the cells which are correctly edited. In accordance with the invention, the correctly edited cells will be viable (since the double strand break has been repaired) and will not express the selectable marker since following HDR the Group II intron and selectable marker will be lost from the cell. In instances then where the selectable marker is a fluorescent protein, the step (f) comprises selecting viable cells that do not fluoresce. Such a selection step can be performed readily using FACS or other means.

[0138] It will be clear to the skilled person that due to the various steps of selecting the desired cells, the method of the invention is well suited to being an ex vivo or in vitro method rather than a method that can be performed directly on a subject. Accordingly preferably the method of editing a target nucleic acid present in a target cell is an ex vivo method or an in vitro method, and the target cell or cells are a cell or cells that have been obtained from a subject, for example a subject in need of gene editing.

[0139] The invention therefore provides an ex vivo or in vitro method of editing one or a plurality of cells wherein said method comprises the method of editing the nucleic acid of a target cell of the invention.

[0140] In a specific embodiment the invention provides an ex vivo or in vitro method of editing one or more mammalian cells, optionally one or more human cells, said method comprising:

[0141] (a) providing a mammalian cell or a plurality of mammalian cells that have been obtained from a mammal;

[0142] (b) introducing into the mammalian cell or plurality of mammalian cells one or more insertion cassettes according to the invention; and

[0143] (c) incubating the mammalian cell in conditions that permit integration of the insertion cassette into the target nucleic acid, wherein the selectable marker is a fluorescent protein and the nucleic acid sequence that is a cleavage site for a nuclease is a cleavage site for a meganuclease.

[0144] The method may also comprise step (d) detecting mammalian cells which express the fluorescent protein for further processing, optionally wherein expression of the fluorescent protein is performed using a FACS.

[0145] The method may also comprise step (e) introducing into the cells selected for further processing in step (d): i) the nuclease expressing cassette of the invention and the editing cassette of the invention; ii) the editing and nuclease expressing cassette of the invention; or iii) the editing complex of the invention, wherein the nuclease is a meganuclease capable of cleaving the sequence that is a cleavage site for a nuclease, optionally wherein the meganuclease is fused to a nuclear localisation signal and step (f) incubating the mammalian cell or mammalian plurality of cells under conditions that a) permit expression of the meganuclease in the target cell or plurality of target cells; and b) permit homologous recombination between a region of the editing cassette that has homology to the target nucleic acid.

[0146] The method may also comprise the further step (g) of selecting cells following incubation that are viable and that do not fluoresce, optionally wherein said selecting is performed using FACS.

[0147] In another specific embodiment the invention provides an ex vivo or in vitro method of editing one or more mammalian cells, optionally one or more human cells, said method comprising:

[0148] (a) providing a mammalian cell or a plurality of mammalian cells that have been obtained from a mammal;

[0149] (b) introducing into the mammalian cell or plurality of mammalian cells one or insertion complexes as set out herein; and

[0150] (c) incubating the mammalian cell in conditions that permit integration of the insertion cassette into the target nucleic acid, wherein the selectable marker is an antigen towards which a fluorescent (or otherwise labelled antibody) can bind and the nucleic acid sequence that is a cleavage site for a nuclease is a cleavage site for a meganuclease.

[0151] The method may also comprise step (d) detecting mammalian cells which express the antigen towards which a fluorescently labelled antibody can bind, for example wherein expression of the selectable marker that is an antigen is performed using a fluorescently labelled antibody and FACS. The method may also comprise step (e) introducing into the cells selected for further processing in step (d) the editing complex of the invention, wherein the nuclease is a meganuclease capable of cleaving the sequence that is a cleavage site for a nuclease, optionally wherein the meganuclease is fused to a nuclear localisation signal and step (f) incubating the mammalian cell or mammalian plurality of cells under conditions that permit homologous recombination between a region of the editing cassette that has homology to the target nucleic acid.

[0152] The method may also comprise the further step (g) of selecting cells following incubation that are viable and that do not express the antigen towards which a fluorescently labelled antibody can bind, optionally wherein said selecting is performed using a fluorescently labelled antibody and FACS.

[0153] The invention also provides a cell, or a population of cells that have been edited according to any of the methods of editing nucleic acid set out herein.

[0154] Accordingly the invention provides an edited cell that has been edited according to any of the methods of the invention.

[0155] Following editing and selection for cells which have lost the marker, any off-target or unedited cells should have been lost from the population, leaving a population of highly pure correctly edited cells.

[0156] Accordingly the invention also provides a population of edited cells that have been edited according to any of the methods of the invention, wherein the population of edited cells comprise no cells with an off-target gene edit event.

[0157] In some instances the population of edited cells is such that at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99% or 100% of the cells comprise a desired editing event and no off-target editing events.

[0158] The invention also provides a pharmaceutical composition comprising an edited cell or a population of edited cells according to the invention. The pharmaceutical composition may comprise one or more agents that make the composition suitable for administration to a mammal for example to a human. For example may comprise one or more pharmaceutically acceptable excipients or adjuvants.

[0159] The invention also provides therapeutic uses of the edited cells or populations of edited cells of the invention. For example in some instances the cells may have been edited to remove a faulty gene, or may have been edited to express or overexpress a therapeutic protein or RNA.

[0160] Accordingly the invention provides a therapeutic method of treatment comprising administering an edited cell according to the invention or a population of edited cells according to the invention or a pharmaceutical composition according to the invention to a subject in need thereof.

[0161] The invention also provides an edited cell according to the invention or a population of edited cells according to the invention or a pharmaceutical composition according to the invention for use in a method of treatment wherein the method comprising administering the edited cell or population of edited cells or a pharmaceutical composition to a subject in need thereof.

[0162] The invention also provides the use of an edited cell according to the invention or a population of edited cells according to the invention or a pharmaceutical composition according to the invention for use in a method of manufacture of a medicament.

[0163] The invention also provides the following numbered embodiment paragraphs:

[0164] 1. An insertion cassette that is a nucleic acid comprising at least: i) a sequence that encodes a group II intron; ii) a sequence that encodes a selectable marker operably linked to a selectable marker promoter; and iii) a nucleic acid sequence that is a cleavage site for a nuclease.

[0165] 2. The insertion cassette of embodiment 1 wherein the sequence of the group II intron is designed to have a region that is capable of hybridising to a region of a target nucleic acid.

[0166] 3. The insertion cassette of any of the preceding embodiments wherein the sequence that encodes a selectable marker, selectable marker promoter and the nucleic acid sequence that is a cleavage site for a nuclease are arranged within the insertion cassette such that upon insertion of the Group II intron a target nucleic acid, optionally into a target cell genome, all of the sequence that encodes a selectable marker, selectable marker promoter and the nucleic acid sequence that is a cleavage site for a nuclease are inserted into the target cell genome / nucleic acid.

[0167] 4. The insertion cassette of any of the preceding embodiments wherein the selectable marker is a marker that confers a physically detectable phenotype on a target cell that expresses said marker.

[0168] 5. The insertion cassette of embodiment 4 wherein the physically detectable phenotype is a phenotype detectable by a cell sorter, optionally by a FACS machine.

[0169] 6. The insertion cassette of embodiment 4 or 5 wherein the physically detectable phenotype is selected from: the ability to produce light, optionally produce light of a particular wavelength after interrogation with a laser; size; and / or morphology.

[0170] 7. The insertion cassette of any of embodiments 4-6 wherein the physically detectable phenotype is the ability to produce light and the selectable marker is: a) a fluorescent protein; or b) a bioluminescent protein.

[0171] 8. The insertion cassette of any of the preceding embodiments wherein the selectable marker is: a) a fluorescent protein; or b) a bioluminescent protein; c) not an antibiotic resistance gene.

[0172] 9. The insertion cassette of embodiment 7 or 8, wherein the fluorescent protein is selected from the group comprising or consisting of: blue fluorescent protein, green fluorescent protein, yellow fluorescent protein, and red fluorescent protein.

[0173] 10. The insertion cassette of embodiment 9, wherein: a)the blue fluorescent protein is selected from the group comprising or consisting of: TagBFP, mTagBFP2, EBFP, EBFP2, Azutire, mKamalal, cyan fluorescent protein, ECFP, Cerulian, CyPet, mTurquoise, mTurquoise2, and AmCyanl; b)the green fluorescent protein is selected from the group comprising or consisting of: miniGFPl, GFP, EGFP, mNeonGreen, and mGFP; c)the yellow fluorescent protein is selected from the group comprising or consisting of: YFP, Citrine, Venue, and YPet; and / or d)the red fluorescent protein is selected from the group comprising or consisting of: mCherry, mCardinal, mScarlet, DsRed, mOrange, mRaspberry, mFruit, mK, TagRFP, mKate, mRuby, FusionRed, and DsRed-Express.

[0174] 11. The insertion cassette of embodiment 7 or 8 wherein the bioluminescent protein is selected from the group comprising or consisting of: luciferase, luxA, and luxB.

[0175] 12. The insertion cassette of any of embodiments 1-3 wherein the selectable marker is a protein or antigen to which a fluorescently labelled antibody or fragment thereof can bind, for example a fluorescently labelled ScFV or Fab fragment.

[0176] 12. The insertion cassette of any of the preceding embodiments wherein the sequence that encodes a selectable maker is only expressed once the transposition of the Group II intron and the selectable maker into a target nucleic acid has occurred.

[0177] 13. The insertion cassette of any of the preceding embodiments wherein the sequence that encodes a selectable marker comprises a removable intervening sequence arranged such that when the removable sequence is present expression from the selectable marker promoter does not produce a functional protein, and wherein removal of the removable intervening sequence allows expression of a functional protein or RNA, optionally wherein the selectable marker is arranged as a RAM.

[0178] 14. The insertion cassette of embodiment 13 wherein the removable intervening sequence is a group I intron, optionally wherein the group I intron is flanked by exons that enable self-splicing of the group I intron.

[0179] 15. The insertion cassette of any of the preceding embodiments wherein the selectable marker and / or the selectable marker promoter are present in the cassette in the antisense orientation, such that: a) direct transcription from the selectable marker promoter from the insertion cassette is not possible; and / or b) translation of a transcript produced from the selectable marker promoter from the insertion cassette does not produce a functional selectable marker protein or RNA. 16. The insertion cassette of any of the preceding embodiments wherein the nucleic acid is DNA.

[0180] 17 The insertion cassette of any of the preceding embodiments wherein the nucleic acid is RIMA.

[0181] 18. The insertion cassette of any of the preceding embodiments wherein the nucleic acid is double-stranded.

[0182] 19. The insertion cassette of any of the preceding embodiments wherein the nucleic acid is single-stranded.

[0183] 20. The insertion cassette of any of any of the preceding embodiments wherein the nucleic acid is linear.

[0184] 21. The insertion cassette of any of the preceding embodiments wherein the nucleic acid is a single-stranded RNA molecule.

[0185] 22. The insertion cassette of embodiment 21 wherein the selectable marker and / or selectable marker promoter is present in an antisense form, such that translation from the single-stranded RNA does not produce a functional selectable marker protein or RNA.

[0186] 23. The insertion cassette of embodiment 22 wherein once the insertion cassette is inserted into the target nucleic acid and converted to double stranded DNA, the sense strand of the selectable marker is present, allowing for production of a functional selectable marker protein or RNA.

[0187] 24. The insertion cassette of any of the preceding embodiments wherein the nucleic acid sequence that is a cleavage site for a nuclease is: i) is a meganuclease recognition site; ii) a recognition site for a corresponding TALEN; iii) a recognition site for a corresponding ZFN.

[0188] 25. The insertion cassette of any of the preceding embodiments wherein the nucleic acid sequence that is a cleavage site for a nuclease is a nucleic acid sequence towards which a guide RNA is generated to target an endonuclease, optionally a Cas protein, to said sequence.

[0189] 26. The insertion cassette of any of embodiments 1-24 wherein the nucleic acid sequence that is a cleavage site for a nuclease is not a nucleic acid sequence towards which a guide RNA is generated to target an endonuclease, optionally a Cas protein, to said sequence.

[0190] 27. The insertion cassette of any of the preceding embodiments wherein the nucleic acid sequence that is a cleavage site for a nuclease is a cleavage site for a meganuclease and wherein the meganuclease recognition site is selected from a recognition site for I-Scel, I-Ppol, I-Dmol, I-Njal, I-Dirl, I-Bmol, BneMS4ORFIP, F- Cphl, F-ECOT3I, F-ECOT5I, F-EcoT5II, F-EcoT5IV, F-PhiU5I, F-Scel, F-Scell, F-TevI, F- TevII, F-TevIII, F-TevIV, H-Drel, H-Drel, I-AabMI, I-AchMI, I-Anil, I-ApeKI, I-BanI, I- BasI, I-Bth0305I, I-Bthll, I-BthORFAP, I-Ceul, I-Chul, I-Cmoel, I-Cpal, I-CpaII, I- CpaMI, I-Crel, I-Crell, I-CsmI, I-Cvul, I-Ddil, I-GpeMI, I-Gpil, I-Gzel, I-GzeII, I- HjeMI, I-Hmul, I-HmuII, I-Llal, I-Ltrl, I-LtrWI, I-MpeMI, I-Msol, I-NanI, I-Nfil, I-Nitl, I-OmiII, I-Onul, I-PakI, I-PanMI, I-PfoP3I, I-PnoMI, I-PogTE7I, I-Porl, I-Scal, I-SceII, I-SceIII, I-SceIV, I-SceV, I-SceVI, I-SceVII, I-SecIII, I-SmaMI, I-SpomI, I-SscMI, I- Ssp6803I, I-TevI, I-TevII. I-TevIII. I-TslI. I-TsIWI, I-TspO61I, I-Twol, I-Vdil41I, - Aval, PI-BciPI, PI-HvoWI, PI-MgaI„ PI-MleSI, PI-MtuI, Pl-PabI, Pl-Pabll, Pl-Pful, PI- PfuII, Pl-Pkol, Pl-PkoII, PI-PspI, PI-PspI, Pl-Scal, Pl-Scel, PI-Tful, PI-TfuII, PI-Thyl, PI-Tlil, PI-Tlill, PI-Tmal, PI-TmaKI, Pl-Zba, or has a sequence of any of SEQ ID NO: 1-142; or a sequence that is: at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NOs. 1-142.

[0191] 28. The insertion cassette of any of the preceding embodiments wherein the cassette comprises at least 2, 3, 4, 5 or more sequences that are cleavage sites for nucleases, optionally at least 2, 3, 4, 5 or more sequences that are cleavage sites for one or more meganucleases.

[0192] 29. The insertion cassette of embodiment 28 wherein the at least at least 2, 3, 4, 5 or more sequences that are cleavage sites for nucleases all have the same sequence and / or cleavage sites for the same nuclease. 30. The insertion cassette of embodiment 28 wherein at least at least 2 of the at least 2, 3, 4, 5 or more sequences that are cleavage sites for nucleases are different sequences and / or are cleavage sites for at least two different nucleases.

[0193] 31. The insertion cassette of any of the preceding embodiments wherein the insertion cassette is operably linked to an insertion cassette promoter that is able to drive transcription across the insertion cassette to produce an insertion cassette RNA transcript.

[0194] 32. The insertion cassette of embodiment 31 wherein the insertion cassette promoter is recognised by a DNA dependent RNA polymerase.

[0195] 33. The insertion cassette of embodiment 32 wherein the insertion cassette promoter is: a T7 promoter; a eukaryotic promoter; and / o a cell-type specific promoter, optionally a eukaryotic cell-type specific promoter.

[0196] 34. The insertion cassette of any of the preceding embodiments wherein the selectable marker promoter is selected from the group comprising or consisting of: a) a strong promoter; b) a constitutive promoter; c) a strong constitutive promoter; d) a CMV promoter or derivative thereof.

[0197] 35. The insertion cassette of any of the preceding embodiments wherein the insertion cassette: does not encode a protein comprising reverse transcriptase activity; and / or does not have the ability to self-excise.

[0198] 36. The insertion cassette of any of the preceding embodiments wherein the insertion cassette comprises no sequences that confer the ability to self-excise on the Group II intron.

[0199] 37. The insertion cassette of any of the preceding embodiments wherein the insertion cassette further comprises a portion comprising a second promoter operably linked to a nucleic acid sequence encoding an intron-encoded protein (IEP), optionally i) wherein the second promoter is a eukaryotic promoter, optionally a mammalian promoter, optionally a human promoter; and / or ii) the IEP is LtrA, optionally wherein LtrA comprises a nucleotide sequence substantially as set out in SEQ ID NO: 156; or a sequence that is: a)at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NO: 156; b)about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% identical to SEQ ID NO: 156; and / or c)75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 156.

[0200] 38. The insertion cassette of embodiment 37 wherein expression of the IEP causes: i) an RNA transcript from the insertion cassette promoter to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand; or ii) where the insertion cassette is an RNA molecule, the insertion cassette to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand.

[0201] 39. The insertion cassette of embodiment 37 or 38 wherein the IEP does not confer the ability to self-splice on the Group II intron.

[0202] 39a. The insertion cassette of any of the preceding embodiments, wherein the insertion cassette is encoded by a nucleic acid that further comprises a nucleotide sequence encoding an accessory protein, wherein the accessory protein is not comprised by the Group II intron.

[0203] 39b. The insertion cassette of embodiment 39a, wherein the nucleic acid encoding the accessory protein is operably linked to a second promoter.

[0204] 39c. The insertion cassette of embodiment 39b, wherein the second promoter is a different promoter to the first promoter.

[0205] 39d. The insertion cassette of embodiment 39b or 39c, wherein the second promoter is a eukaryotic promoter, optionally a mammalian promoter, optionally a human promoter. 39e. The insertion cassette of any one of embodiments 39b-39d, wherein the second promoter is selected from the group comprising or consisting of: CMV promoter, an EFla promoter, an SV40 promoter, a UBC promoter, a PGK promoter, a CAG promoter, an SFFV promoter, an MSCV promoter.

[0206] 39f. The insertion cassette of any one of embodiments 39a-39e, wherein the accessory protein is LtrA, optionally wherein LtrA comprises a nucleotide sequence substantially as set out in SEQ ID NO: 156; or a sequence that is: a)at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NO: 156; b)about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% identical to SEQ ID NO: 156; and / or c)75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 156;

[0207] 39g) The insertion cassette of any one of embodiments 39a-39f, wherein expression of the accessory protein causes: i) an RNA transcript from the insertion cassette promoter to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand; or ii) where the insertion cassette is an RNA molecule, the insertion cassette to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand.

[0208] 39h) The insertion cassette of any one of the preceding embodiments, wherein the insertion cassette is encoded by a nucleic acid that further comprises an origin of replication that is not capable of driving replication of the nucleic acid in the target cell.

[0209] 39i) The insertion cassette of embodiment 39h, wherein the origin of replication is a bacterial origin of replication; optionally selected from the group comprising or consisting of: a ColEl origin; a BBR1 origin.

[0210] 39j) The insertion cassette of any one of the preceding embodiments, wherein the insertion cassette is encoded by a nucleic acid that further comprises a selectable marker that is not suitable for selecting the target cell. 39k) The insertion cassette of embodiment 39j), wherein the selectable marker is a bacterial selectable marker; optionally selected from the group comprising or consisting of: an ampicillin resistance cassette, a chloramphenicol resistance cassette, a kanamycin resistance cassette, a streptomycin resistance cassette, a tetracycline resistance cassette, a spectinomycin resistance cassette, an erythromycin resistance cassette, a gentamycin resistance cassette, a rifampicin resistance cassette, a trimethoprim resistance cassette, a carbenicillin resistance cassette, a neomycin resistance cassette, a G418 / geneticin resistance cassette, a hygromycin resistance cassette, a paromomycin resistance cassette, a tobramycin resistance cassette, a puromycin resistance cassette, a blasticidin resistance cassette, a zeocin resistance cassette, a penicillin resistance cassette, and a lincomycin resistance cassette.

[0211] 391) The insertion cassette of embodiment 39j), wherein the selectable marker is a bacterial selectable marker; optionally selected from the group comprising or consisting of: hisA, hisB, hisD (histidine biosynthesis), leuB (leucine biosynthesis), ilvA (isoleucine / valine biosynthesis), metE (methionine biosynthesis), argA, argE (arginine biosynthesis), thrC (threonine biosynthesis), pyrF (UMP biosynthesis (uracil pathway)), purA (adenine biosynthesis), thyA (thymidine biosynthesis), galK (galactose utilization), and araA (arabinose utilization).

[0212] 40. The insertion cassette of any of the preceding embodiments wherein the insertion cassette is an RIMA molecule complexed with an IEP.

[0213] 41. The insertion cassette of any of the preceding embodiments wherein the target nucleic acid is : a genetic locus of a chromosome of a target cell; or mitochondrial DNA of a target cell; or an episomally maintained element such as a plasmid, cosmid or minichromosome optionally where the target cell is a mammalian cell.

[0214] 42. The insertion cassette of any of the preceding embodiments wherein the target nucleic acid is a target nucleic acid present in a target cell, optionally in a target mammalian cell. 43. The insertion cassette of any of the preceding embodiments wherein the insertion cassette is suitable for use with a mammalian cell, optionally a human cell.

[0215] 44. The insertion cassette of any of the preceding embodiments wherein the insertion cassette is not suitable for use with a prokaryotic cell.

[0216] 45. The insertion cassette according to any one of the preceding embodiments, wherein the group II intron comprises LI.LtrB from Lactococcus lactis, Ecl5 from Escherichia coli, Rmlntl from Sinorhizobium meliloti, or TeI3c from Thermosynechococcus elongatus.

[0217] 46. The insertion cassette of any one of the preceding embodiments, wherein the group II intron, selectable marker, sequence that is a cleavage site for a nuclease and / or IEP are codon optimised for a eukaryotic cell, optionally a mammalian cell, optionally a human cell.

[0218] 47. The insertion cassette of any of the preceding embodiments wherein: a)the selectable marker is located upstream (5') of the sequence that is a cleavage site for a nuclease; or b)the selectable marker is located downstream (3') of the sequence that is a cleavage site for a nuclease.

[0219] 47a. A nucleic acid comprising the insertion cassette of any one of embodiments 1-

[0220] 47.

[0221] 47b. The nucleic acid of embodiment 47a, further comprising a nucleotide sequence encoding an accessory protein according to any one of embodiments 39a-39g, optionally where the accessory protein is operably linked to a promoter that drives expression of only the accessory protein.

[0222] 47c. The nucleic acid of embodiment 47a or 47b, further comprising an origin of replication according to any one of embodiments 39h-39i.

[0223] 47d. The nucleic acid of any one of embodiments 47a-47c, further comprising a selectable marker according to any one of embodiments 39j-39l.

[0224] 48. An insertion complex comprising an RIMA transcript from the insertion cassette of any of the preceding embodiments and one or more proteins that allow the insertion of the transcript into the target nucleic acid, optionally an IEP as described in any of the preceding embodiments.

[0225] 49. An editing cassette comprising at least a first region that has homology to a first portion of a target nucleic acid, optionally also comprises at least a second region that has homology to a second portion of said target nucleic acid; wherein said first and second regions have sufficient homology to the first and second portions of the target nucleic acid so as to allow repair of a double stranded break using homology directed repair (HDR).

[0226] 50. The editing cassette nucleic acid of embodiment 49 wherein the cassette comprises a first region that has homology to a first portion of a target nucleic acid and a second region that has homology to a second portion of the target nucleic acid.

[0227] 51. The editing cassette nucleic acid of embodiment 49 or 50 wherein the cassette further comprises an intervening sequence located between the first region that has homology to a first portion of a target nucleic acid and the second region that has homology to a second portion of the target nucleic acid, optionally wherein the intervening sequence comprises an alternative sequence to that present in the target nucleic acid between the first region and the second region, optionally wherein the alternative sequence comprises an insertion, deletion, single nucleotide polymorphism, stop codon, or nonsense mutation with respect to the sequence of the target nucleic acid between the first and second portion.

[0228] 52. The editing cassette of any of embodiments 49-51 wherein the editing cassette comprising a further portion that is a nucleic acid that is an aptamer with binding specificity for a nuclease, optionally wherein the editing cassette is a single stranded DNA molecule.

[0229] 52a. The editing cassette of any one of embodiments 49-52, wherein the editing cassette is encoded by a nucleic acid that further comprises an origin of replication that is not capable of driving replication of the nucleic acid in the target cell.

[0230] 52b. The editing cassette of embodiment 52a, wherein the origin of replication is a bacterial origin of replication; optionally selected from the group comprising or consisting of: a ColEl origin; a BBR1 origin; a pl5A origin, a pSClOl origin, an R6K origin, an RK2 / RP4 origin, and an RSF1010 origin. 52c. The editing cassette of any one of embodiments 49-52b, wherein the insertion cassette is encoded by a nucleic acid that further comprises a selectable marker that is not suitable for selecting the target cell.

[0231] 52d. The editing cassette of embodiment 52c, wherein the selectable marker is a bacterial selectable marker; optionally selected from the group comprising or consisting of: an ampicillin resistance cassette, a chloramphenicol resistance cassette, a kanamycin resistance cassette, a streptomycin resistance cassette, a tetracycline resistance cassette, a spectinomycin resistance cassette, an erythromycin resistance cassette, a gentamycin resistance cassette, a rifampicin resistance cassette, a trimethoprim resistance cassette, a carbenicillin resistance cassette, a neomycin resistance cassette, a G418 / geneticin resistance cassette, a hygromycin resistance cassette, a paromomycin resistance cassette, a tobramycin resistance cassette, a puromycin resistance cassette, a blasticidin resistance cassette, a zeocin resistance cassette, a penicillin resistance cassette, and a lincomycin resistance cassette.

[0232] 52e. The editing cassette of embodiment 52c, wherein the selectable marker is a bacterial selectable marker; optionally a bacterial auxotrophic marker; optionally selected from the group comprising or consisting of: hisA, hisB, hisD (histidine biosynthesis), leuB (leucine biosynthesis), ilvA (isoleucine / valine biosynthesis), metE (methionine biosynthesis), argA, argE (arginine biosynthesis), thrC (threonine biosynthesis), pyrF (UMP biosynthesis (uracil pathway)), purA (adenine biosynthesis), thyA (thymidine biosynthesis), galK (galactose utilization), and araA (arabinose utilization).

[0233] 53. A nuclease expressing cassette comprising a nucleic acid sequence that encodes a nuclease operably linked to a promoter.

[0234] 54. The nuclease expressing cassette of embodiment 53 wherein the cassette comprises more than one nucleic acid sequence encoding more than one different nuclease.

[0235] 55. The nuclease expressing cassette of embodiment 53 or 54 wherein the nuclease or more than one different nuclease is a meganuclease, a ZFN or a TALEN, optionally wherein the meganuclease is selected from the group comprising or consisting of: I-Scel, I-Ppol, I-Dmol, I-Njal, I-Dirl, I-Bmol, BneMS4ORFIP, F-CphI, F-EcoT3I, F- ECOT5I, F-ECOT5II, F-EcoT5IV, F-PhiU5I, F-Scel, F-Scell, F-TevI, F-TevII, F-TevIII, F- TevIV, H-Drel, H-Drel, I-AabMI, I-AchMI, I-Anil, I-ApeKI, I-BanI, I-BasI, I-Bth0305I, I-BthII, I-BthORFAP, I-Ceul, I-Chul, I-Cmoel, I-Cpal, I-CpaII, I-CpaMI, I-Crel, I- Crell, I-CsmI, I-Cvul, I-Ddil, I-GpeMI, I-Gpil, I-Gzel, I-GzeII, I-HjeMI, I-Hmul, I- HmuII, I-Llal, I-Ltrl, I-LtrWI, I-MpeMI, I-Msol, I-NanI, I-Nfil, I-Nitl, I-Omill, I-Onul, I-PakI, I-PanMI, I-PfoP3I, I-PnoMI, I-PogTE7I, I-Porl, I-Scal, I-SceII, I-SceIII, I- ScelV, I-SceV, I-SceVI, I-SceVII, I-SecIII, I-SmaMI, I-SpomI, I-SscMI, I-Ssp6803I, I- TevI, I-TevII. I-TevIII. I-TslI. I-TsIWI, I-TspO61I, I-Twol, I-Vdil41I, -Aval, PI-BciPI, PI-HvoWI, PI-MgaI„ PI-MleSI, PI-MtuI, Pl-PabI, Pl-Pabll, Pl-Pful, Pl-PfuII, Pl-Pkol, PI- PkoII, PI-PspI, PI-PspI, Pl-Scal, Pl-Scel, PI-Tful, PI-TfuII, PI-Thyl, PI-Tlil, PI-Tlill, PI- Tmal, PI-TmaKI, Pl-Zba, or has a sequence of any of SEQ ID NO: 1-142; or a sequence that is: at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NOs. 1-142.

[0236] 56. The nuclease expressing cassette of any of embodiments 53-55 wherein the promoter is: a) an inducible promoter; b) a constitutive promoter; c) a eukaryotic promoter; d) a mammalian promoter; e) a human promoter; and / or f) a promoter that is not operable in a prokaryote.

[0237] 57. The nuclease expressing cassette of any of embodiments 53-56 wherein the nucleic acid sequence that encodes a nuclease comprises a portion that encodes a nuclear localisation signal (NLS), such that upon expression the nuclease protein is fused to a NLS.

[0238] 57a. The nuclease expressing cassette of embodiment 57, wherein the nuclear localisation signal (NLS) is selected from the group comprising or consisting of: SV40 NLS, c-Myc NLS, p53 NLS, nucleoplasmin NLS, yeast GAL4 NLS, Xenopus nucleoplasmin NLS, adenovirus E1A NLS, polyoma virus large T antigen NLS.

[0239] 57b. The nuclease expressing cassette of embodiment 57 or 57a, wherein the nuclear localisation signal (NLS) is a mammalian nuclear localisation signal (NLS); optionally wherein the mammalian nuclear localisation signal (NLS) is selected from the group comprising or consisting of: SV40 NLS, c-Myc NLS, p53 NLS, nucleoplasmin NLS; optionally has an amino acid sequence selected from the group comprising or consisting of: SEQ ID NO: 160 (SV40 Large T antigen NLS (PKKKRKV)), SEQ ID NO: 161 (c-Myc NLS (PAAKRVKLD)), SEQ ID NO: 162 (p53 NLS (PQPKKKPL)), SEQ ID NO: 163 (nucleoplasmin NLS (KRPAATKKAGQAKKKK)).

[0240] 57c. The nuclease expressing cassette of embodiment 57 or 57a, wherein the nuclear localisation signal (NLS) is from a non-mammalian cell but is operable in a mammalian cell; optionally wherein the nuclear localisation signal (NLS) from a non-mammalian cell that is operable in a mammalian cell is selected from the group comprising or consisting of: yeast GAL4 NLS, Xenopus nucleoplasmin NLS, adenovirus E1A NLS, polyoma virus large T antigen NLS; optionally has an amino acid sequence selected from the group comprising or consisting of: SEQ ID NO: 164 (yeast GAL4 NLS (MKLKRAARRRRRKR)), SEQ ID NO: 165 (adenovirus E1A NLS (KRPRP)), SEQ ID NO: 166 (polyoma virus large T antigen NLS (VSRKRPRP)).

[0241] 57d. The nuclease expressing cassette of any one of embodiments 53-57c, wherein the nuclease expressing cassette is encoded by a nucleic acid that further comprises an origin of replication that is not capable of driving replication of the nucleic acid in the target cell.

[0242] 57e. The nuclease expression cassette of embodiment 52d, wherein the origin of replication is a bacterial origin of replication; optionally selected from the group comprising or consisting of: a ColEl origin; a BBR1 origin.

[0243] 57f. The nuclease expression cassette of any one of embodiments 53-52e, wherein the nuclease expression cassette is encoded by a nucleic acid that further comprises a selectable marker that is not suitable for selecting the target cell.

[0244] 57g. The nuclease expression cassette of embodiment 52f, wherein the selectable marker is a bacterial selectable marker; optionally: a) wherein the bacterial selectable marker is selected from the group comprising or consisting of: an ampicillin resistance cassette, a chloramphenicol resistance cassette, a kanamycin resistance cassette, a streptomycin resistance cassette, a tetracycline resistance cassette, a spectinomycin resistance cassette an erythromycin resistance cassette, a gentamycin resistance cassette, a rifampicin resistance cassette, a trimethoprim resistance cassette, a carbenicillin resistance cassette, a neomycin resistance cassette, a G418 / geneticin resistance cassette, a hygromycin resistance cassette, a paromomycin resistance cassette, a tobramycin resistance cassette, a puromycin resistance cassette, a blasticidin resistance cassette, a zeocin resistance cassette, a penicillin resistance cassette, and a lincomycin resistance cassette; or b) wherein the bacterial selectable marker is a bacterial auxotrophic marker; optionally selected from the group comprising or consisting of: hisA, hisB, hisD (histidine biosynthesis), leuB (leucine biosynthesis), ilvA (isoleucine / valine biosynthesis), metE (methionine biosynthesis), argA, argE (arginine biosynthesis), thrC (threonine biosynthesis), pyrF (UMP biosynthesis (uracil pathway)), purA (adenine biosynthesis), thyA (thymidine biosynthesis), galK (galactose utilization), and araA (arabinose utilization).

[0245] 57h. A nucleic acid comprising the editing cassette of any one of embodiments 49- 52d.

[0246] 57i. The nucleic acid of embodiment 57h, further comprising the nuclease expression cassette of any one of embodiments 53-57g.

[0247] 57j. A nucleic acid comprising the nuclease expression cassette of any one of embodiments 53-57g.

[0248] 57k. The nucleic acid of embodiment 57j, further comprising the editing cassette of any one of embodiments 49-52d.

[0249] 571. The nucleic acid of any one of embodiments 57h-57k, further comprising a nucleotide sequence encoding an accessory protein according to any one of embodiments 39a-39g.

[0250] 57m. The nucleic acid of any one of embodiments 57h-57l, further comprising an origin of replication according to any one of embodiments 39h-39i.

[0251] 57n. The nucleic acid of any one of embodiments 57h-57m, further comprising a selectable marker according to any one of embodiments 39j-39k.

[0252] 58. An editing complex comprising the editing cassette of any of embodiments 49- 52 complexed with a nuclease protein, optionally wherein the editing cassette is the editing cassette of any one of embodiments 52-52d. 59. The editing complex of embodiment 58 wherein the nuclease protein is a meganuclease, a ZFN or a TALEN, optionally wherein the meganuclease is selected from the group comprising or consisting of:

[0253] I-Scel, I-Ppol, I-Dmol, I-Njal, I-Dirl, I-Bmol, BneMS4ORFIP, F-CphI, F-EcoT3I, F- ECOT5I, F-ECOT5II, F-EcoT5IV, F-PhiU5I, F-Scel, F-Scell, F-TevI, F-TevII, F-TevIII, F- TevIV, H-Drel, H-Drel, I-AabMI, I-AchMI, I-Anil, I-ApeKI, I-BanI, I-BasI, I-Bth0305I, I-BthII, I-BthORFAP, I-Ceul, I-Chul, I-Cmoel, I-Cpal, I-CpaII, I-CpaMI, I-Crel, I- Crell, I-CsmI, I-Cvul, I-Ddil, I-GpeMI, I-Gpil, I-Gzel, I-GzeII, I-HjeMI, I-Hmul, I- HmuII, I-Llal, I-Ltrl, I-LtrWI, I-MpeMI, I-Msol, I-NanI, I-Nfil, I-Nitl, I-Omill, I-Onul, I-PakI, I-PanMI, I-PfoP3I, I-PnoMI, I-PogTE7I, I-Porl, I-Scal, I-SceII, I-SceIII, I- ScelV, I-SceV, I-SceVI, I-SceVII, I-SecIII, I-SmaMI, I-SpomI, I-SscMI, I-Ssp6803I, I- TevI, I-TevII. I-TevIII. I-TslI. I-TsIWI, I-TspO61I, I-Twol, I-Vdil41I, -Aval, PI-BciPI, PI-HvoWI, PI-MgaI„ PI-MleSI, PI-MtuI, Pl-PabI, Pl-Pabll, Pl-Pful, Pl-PfuII, Pl-Pkol, PI- PkoII, PI-PspI, PI-PspI, Pl-Scal, Pl-Scel, PI-Tful, PI-TfuII, PI-Thyl, PI-Tlil, PI-Tlill, PI- Tmal, PI-TmaKI, Pl-Zba, or has a sequence of any of SEQ ID NO: 1-142; or a sequence that is: at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NOs. 1-142.

[0254] 60. The editing complex of embodiment 58 or 59 wherein the editing cassette comprises a portion that comprises one or more aptamers which have binding specificity for the nuclease, optionally wherein the editing cassette is a single-stranded DNA editing cassette that comprises a portion that comprises one or more aptamers which have binding specificity for the nuclease.

[0255] 61. The editing complex of any one of embodiments 58-60 wherein the nuclease is fused to an NLS, optionally wherein the nuclease is a meganuclease fused to an NLS.

[0256] 62. An editing and nuclease expressing cassette comprising the editing cassette of any one of embodiments 49-52d and the nuclease expressing cassette of any one of embodiments 53-57d on a single nucleic acid molecule, optionally on the same doublestranded DNA molecule, optionally on the same plasmid optionally wherein the editing and nuclease expressing cassette comprises more than one different nuclease expressing cassette.

[0257] 63. The insertion cassette of any of embodiments l-47d, the editing cassette of any of embodiments 49-52d, the nuclease expressing cassette of any of embodiments 53- 57n, the editing and nuclease expressing cassette of embodiment 62 wherein the cassette is: i) on a plasmid ii) on a vector iii) on a phagemid iv) on a cosmid v) part of a viral genome vi) on an artificial chromosome vii) on a non-replicative nucleic acid optionally on a non-replicative plasmid viii) is not maintained within a target cell or population of cells.

[0258] 64. The insertion cassette of any of embodiments l-47d and 63, the editing cassette of any of embodiments 49-52d and 63, the nuclease expressing cassette of any of embodiments 53-57n and 63, the editing and nuclease expressing cassette of embodiment 62 and 63 wherein the cassette is on a circular nucleic acid.

[0259] 65. The insertion cassette of any of embodiments l-47d and 63, the editing cassette of any of embodiments 49-52d and 63, the nuclease expressing cassette of any of embodiments 53-57n and 63, the editing and nuclease expressing cassette of embodiment 62 and 63 wherein the cassette is on a linear nucleic acid.

[0260] 66. The insertion cassette of any of embodiments l-47d and 63-65, the editing cassette of any of embodiments 49-52d and 63-65, the nuclease expressing cassette of any of embodiments 53-57n and 63-65, the editing and nuclease expressing cassette of embodiment 62 and 63-65 wherein the cassette is DNA.

[0261] 67. The insertion cassette of any of embodiments l-47d and 63-65, the editing cassette of any of embodiments 49-52d and 63-65, the nuclease expressing cassette of any of embodiments 53-57n and 63-65, the editing and nuclease expressing cassette of embodiment 62 and 63-65 wherein the cassette is RIMA.

[0262] 68. A mammalian nucleic acid editing system comprising any two or more of:

[0263] The insertion cassette according to any of embodiments l-47d;

[0264] The insertion complex of embodiment 48;

[0265] The editing cassette of any of embodiments 49-52d;

[0266] The nuclease expressing cassette of any of embodiments 53-57n;

[0267] The editing complex of any embodiments 58-61; and / or

[0268] The editing and nuclease expressing cassette of embodiment 62. 69. The mammalian nucleic acid editing system of embodiment 68 comprising:

[0269] The insertion cassette according to any of embodiments l-47d and: i) the editing cassette of any of embodiments 49-52d and the nuclease expressing cassette of any of embodiments 53-57n; ii) the editing and nuclease expressing cassette of embodiment 62; or iii) the editing complex of any of embodiments 58-61.

[0270] 69a. The mammalian nucleic acid editing system of embodiment 68 or 69, comprising:

[0271] A nucleic acid comprising an insertion cassette of any one of embodiments 1- 47 and a nucleotide sequence encoding an accessory protein of any one of embodiments 39a-39f, wherein the nucleic acid encoding the accessory protein is not comprised by the insertion cassette; and i) the editing cassette according to the any one of embodiments 49-52d invention and the nuclease expressing cassette according to any one of 53-57n; ii) the editing and nuclease expressing cassette according to embodiment 62; or iii) the editing complex of any one of embodiments 58-61.

[0272] 70. The mammalian nucleic acid editing system of embodiment 68 comprising:

[0273] The insertion complex of embodiment 48 and: i) the editing cassette of any of embodiments 49-52d and the nuclease expressing cassette of any of embodiments 53-57n; ii) the editing and nuclease expressing cassette of embodiment 62; or iii) the editing complex of any of embodiments 58-61.

[0274] 71. A kit comprising: a) any two or more of:

[0275] The insertion cassette according to any of embodiments l-47d;

[0276] The insertion complex of embodiment 48;

[0277] The editing cassette of any of embodiments 49-52d;

[0278] The nuclease expressing cassette of any of embodiments 53-57n;

[0279] The editing complex of any embodiments 58-61; and / or The editing and nuclease expressing cassette of embodiment 62; b) The insertion cassette according to any of embodiments l-47d and: i) the editing cassette of any of embodiments 49-52d and the nuclease expressing cassette of any of embodiments 53-57n; ii) the editing and nuclease expressing cassette of embodiment 62; or iii) the editing complex of any of embodiments 58-61; and / or c) The insertion complex of embodiment 48 and: i) the editing cassette of any of embodiments 49-52d and the nuclease expressing cassette of any of embodiments 53-57n; ii) the editing and nuclease expressing cassette of embodiment 62; or iii) the editing complex of any of embodiments 58-61.

[0280] 71a. A kit comprising:

[0281] A nucleic acid comprising an insertion cassette of any one of embodiments 1- 47d the invention and a nucleotide sequence encoding an accessory protein of any one of embodiments 39a-39f, wherein the nucleic acid encoding the accessory protein is not comprised by the insertion cassette; and i) the editing cassette according to the any one of embodiments 49-52d invention and the nuclease expressing cassette according to any one of 53-57n; ii) the editing and nuclease expressing cassette according to embodiment 62; or iii) the editing complex of any one of embodiments 58-61.

[0282] 72. The mammalian nucleic acid editing system of embodiment 68-70 or the kit of embodiment 71 or 71a wherein the editing system or the kit is suitable for use in a method of editing a target nucleotide sequence in a target cell.

[0283] 73. A cell comprising any one or more of: The insertion cassette according to any of embodiments l-47d;

[0284] The insertion complex of embodiment 48;

[0285] The editing cassette of any of embodiments 49-52d;

[0286] The nuclease expressing cassette of any of embodiments 53-57n;

[0287] The editing complex of any embodiments 58-61; and / or

[0288] The editing and nuclease expressing cassette of embodiment 62. 73a. A cell comprising:

[0289] A nucleic acid comprising an insertion cassette of any one of embodiments 1- 47d the invention and a nucleotide sequence encoding an accessory protein of any one of embodiments 39a-39f, wherein the nucleic acid encoding the accessory protein is not comprised by the insertion cassette; and i) the editing cassette according to the any one of embodiments 49-52d invention and the nuclease expressing cassette according to any one of 53-57n; ii) the editing and nuclease expressing cassette according to embodiment 62; or iii) the editing complex of any one of embodiments 58-61.

[0290] 74. The cell of embodiment 73 or 73a, wherein the cell is a eukaryotic cell; optionally wherein the cell is a mammalian cell, optionally a human cell.

[0291] 75. A method of editing a target nucleic acid present in a target cell, the method comprising:

[0292] (a) providing a target cell or a plurality of target cells;

[0293] (b)introducing into the target cell or plurality of target cells one or more insertion cassettes according to any of embodiments l-47d; and

[0294] (c)incubating the target cell or plurality of target cells under conditions that permit integration of the insertion cassette into the target nucleic acid.

[0295] 76. The method of embodiment 75 further comprising step (d) detecting the presence of the selectable maker, wherein detection of the selectable marker indicates that the insertion nucleic acid is present in the target cell, and selecting those cells in which the presence of the selectable marker is detected for further processing.

[0296] 77. The method of embodiment 76 wherein the selectable marker present in the insertion cassette according to any of embodiments l-47d is a fluorescent protein, optionally selected from the group comprising or consisting of: blue fluorescent protein, green fluorescent protein, yellow fluorescent protein, and red fluorescent protein and the step (d) comprises selecting cells that emit light, optionally fluoresce.

[0297] 78. The method of embodiment 77 wherein the step of selecting cells that express the fluorescent protein is performed using a fluorescence activated cell sorter (FACS). 79. The method of any of embodiments 70-73 wherein the method further comprises step (e) wherein step (e) comprises introducing into the cells selected for further processing in step (d): i) the nuclease expressing cassette of any of embodiments 53-57n and the editing cassette of any of embodiments 49-52d; ii) the editing and nuclease expressing cassette of embodiment 62; or iii) the editing complex of any of embodiments 58-61.

[0298] 80. The method of embodiment 79 wherein the method further comprises incubating the target cell or plurality of target cells under conditions that permit expression of the nuclease in the target cell or plurality of target cells.

[0299] 81. The method of any of embodiments 79 or 80 wherein the cell(s) are incubated under conditions that permit HDR of the double stranded break produced by the nuclease, using the editing cassette as the template for repair.

[0300] 82. The method of any of embodiments 75-81 further comprising step (f), selecting for cells which comprise the edited nucleic acid.

[0301] 83. The method of embodiment 82 wherein step (f) comprises selecting for cells that are viable and which do not express the selectable marker.

[0302] 84. The method of embodiment 83 wherein where the selectable marker present in the insertion cassette according to any of embodiments l-47d is a fluorescent protein, optionally selected from the group comprising or consisting of: blue fluorescent protein, green fluorescent protein, yellow fluorescent protein, and red fluorescent protein and the step (d) comprises selecting cells that fluoresce, the method step (f) comprises identifying and selecting cells that do not fluoresce.

[0303] 85. The method of any one of embodiments 75-84, wherein the target cell or each cell of the plurality of target cells is a eukaryotic cell; a mammalian cell, optionally a human cell.

[0304] 86. The method of any one of embodiments 75-85, wherein integration of the editing cassette into the target nucleic acid changes the sequence of the target sequence; optionally wherein the change comprises a insertion, deletion, single nucleotide polymorphism, stop codon, or nonsense mutation. 87. The method of any one of embodiments 75-86 wherein the method is an ex vivo method or an in vitro method.

[0305] 88. An ex vivo or in vitro method of editing one or a plurality of cells wherein said method comprises the method of any one or more of embodiments 75-87.

[0306] 89. An ex vivo or in vitro method of editing one or more mammalian cells, optionally one or more human cells, said method comprising:

[0307] (a)providing a mammalian cell or a plurality of mammalian cells that have been obtained from a mammal;

[0308] (b)introducing into the mammalian cell or plurality of mammalian cells one or more insertion cassettes according to any of embodiments l-47d; and

[0309] (c)incubating the mammalian cell in conditions that permit integration of the insertion nucleic acid into the target nucleic acid wherein the selectable marker is a fluorescent protein and the nucleic acid sequence that is a cleavage site for a nuclease is a cleavage site for a meganuclease.

[0310] 90. The method of embodiment 89 wherein the method further comprises step (d) detecting mammalian cells which express the fluorescent protein for further processing, optionally wherein expression of the fluorescent protein is performed using a FACS.

[0311] 91. The method of embodiment 90 further comprising: step (e) introducing into the cells selected for further processing in step (d): i) the nuclease expressing cassette of any of embodiments 53-57n and the editing cassette of any of embodiments 49-52d; ii) the editing and nuclease expressing cassette of embodiment 62; or iii) the editing complex of any of embodiments 58-61 wherein the nuclease is a meganuclease capable of cleaving the sequence that is a cleavage site for a nuclease, optionally wherein the meganuclease is fused to a nuclear localisation signal and step (f) incubating the mammalian cell or mammalian plurality of cells under conditions that a) permit expression of the meganuclease in the target cell or plurality of target cells; and b) permit homologous recombination between a region of the editing cassette that has homology to the target nucleic acid.

[0312] 92. The method of embodiment 91 comprising the further step (g) of selecting cells following incubation that are viable and that do not fluoresce, optionally wherein said selecting is performed using FACS.

[0313] 93. An edited cell that has been edited according to any of the methods of any one or embodiments 75-92.

[0314] 94. A population of edited cells according to embodiment 93 wherein the population of edited cells comprises: a) no cells with an off-target gene edit event; b) cells in which at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99% or 100% 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cells comprise a desired editing event; c) less than 40%, 35%, 30%, 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% cells that comprise an indel at the cleavage site; and / or d) fewer indels at the cleavage site as compared to CRISPR based editing methods.

[0315] 95. A pharmaceutical composition comprising an edited cell according to embodiment 93 or a population of edited cells according to embodiment 94, optionally wherein said composition comprises one or more pharmaceutically acceptable excipients or adjuvants.

[0316] 96. A therapeutic method of treatment comprising administering an edited cell according to embodiment 93 or a population of edited cells according to embodiment 94 or a pharmaceutical composition of embodiment 95 to a subject in need thereof.

[0317] 97. An edited cell according to embodiment 93 or a population of edited cells according to embodiment 94 or a pharmaceutical composition of embodiment 95 for use in a method of treatment wherein the method comprising administering the edited cell or population of edited cells or a pharmaceutical composition to a subject in need thereof.

[0318] 98. Use of an edited cell according to embodiment 93 or a population of edited cells according to embodiment 94 or a pharmaceutical composition of embodiment 95 for use in a method of manufacture of a medicament.

[0319] Figures

[0320] Figure 1 - Exemplary schematic describing an embodiment of the invention. As set out herein, variations of this embodiment are included within the invention, with the figure setting out the core editing strategy. The strategy of Figure 1 employs an insertion complex as described herein and an editing complex as described herein.

[0321] Figure 2 - Exemplary list of meganucleases and corresponding recognition site that are suitable for use with the invention.

[0322] Figure 3 - Exemplary schematic diagram of a two-step process for generating targeted chromosomal edits in line with the invention.

[0323] Figure 4 - Exemplary representation of an intron containing a vector used to mobilise the intron containing a meganuclease site into a target locus in HEK 293 cells.

[0324] Figure 5 - PCR screen from Example 1, demonstrating intron insertions in HEK 293 cells into a plasmid-harboured target site. PCR of extracted plasmids using an oligonucleotide which binds outside of the target site and one which binds within the inserted intron for each integration junction. Insertion of the intron generates a product size of 1.56 kb for the 5' junction and 1.36 kb for the 3' junction, with the target plasmid only (designated TP) and not producing an amplicon due to the absence of an inserted intron as no intron donor plasmid was co -transfected into these HEK 293 cells. A Thermo Scientific GeneRuler 1 kb Plus DNA Ladder was used for this agarose gel.

[0325] Figure 6 - PCR screen from Example 2, demonstrating I-Scel-induced homology directed repair (HDR) in HEK 293 cells of a plasmid-harboured intron, containing a meganuclease site (as generated in Example 1). PCR of extracted plasmids using an oligonucleotide which binds outside of the homology cassette locus and one which binds within the edit-inserted I-Crel site for each junction. Insertion of the intron generates a product size of 1.02 kb for the 5' junction and 1.11 kb for the 3' junction, with the intron- inserted plasmid only (designated IP) and editing plasmid only (designated EP) controls not producing an amplicon due to the absence HDR in HEK 293 cells. A Thermo Scientific GeneRuler 1 kb Plus DNA Ladder was used for this agarose gel.

[0326] Figure 7 - PCR screen from Example 3, demonstrating targeted intron insertions into the chromosomal LINE1 locus in HEK 293 cells, using introns that do or not harbour a GFP marker. PCR of extracted total DNA using pairs of oligonucleotides, one which binds outside of the target site and one which binds within the inserted intron for each integration junction. Two sets of oligonucleotides primers were used to confirm integration at each junction. Insertion of an intron into the genome generates a product size of 948 bp and 536 bp for the nested 5' junction screens for both intron variants, whereas the 3' junction screen generates a 1163 bp and 784 bp for the GFP-marked intron (designated M) and 734 bp and 355 bp products for the unmarked intron (designated NM). Wild type HEK 293 DNA (designated WT) does not produce a correctly sized amplicon at each junction due to the absence of an inserted intron. An Thermo Scientific GeneRuler 1 kb Plus DNA Ladder was used for this agarose gel.

[0327] Figure 8 - Exemplary schematic of plasmids used in the Examples, a) Exemplary plasmid (hFEl) comprising an insertion cassette provided herein, b) Exemplary targeting plasmid (pTR). c) Exemplary expression editing plasmid (hFE2).

[0328] Examples

[0329] Example 1: Targeting a plasmid in HEK 293 cells

[0330] In this example, a eukaryotic single plasmid system for intron insertion was used. This plasmid consists of LI.LtrB transcribed from a human Hl Pol III promoter. The intron LtrA ORF has been deleted, with I-Scel meganuclease recognition site inserted in place of the deleted ORF, whilst maintaining the DIVb stem structure. To enable expression from the Hl promoter, as transcripts are terminated at thymine repeats, the intron sequence was altered to avoid strings of four repeat thymine nucleotides and to maintain the RNA structure. Following a T6 PolIII termination sequence, the human codon optimised LtrA is expressed from the PolII cytomegalovirus immediate early (CMV) promoter, fused to C-terminal 2 x SV40 NLS and followed by a polyadenylation signal. In this example the group II intron targeting region was that of the native LI.LtrB and the target site was contained on a targeting plasmid. The target plasmid, pTR, contains a wild type LI.LtrB target site inserted between the catP resistance marker and pBBRl bacterial replicon. Targeting experiments were performed in HEK 293 cells. These cells were maintained in DMEM, with high glucose and pyruvate, supplemented with 10% fetal bovine serum and 10% penicillin / streptomycin, at 37°C with 5% CO2 following standard cell culture practices. For intron integration the HEK 293 cells were seeded on a 12-well plate to around 80% confluency and cells were co-transfected using Lipofectamine 2000 (Invitrogen). Equal quantities of the intron-containing plasmid and the targeting plasmid, totaling 1.5 pg of purified plasmid DNA was used. Cells were incubated for 16-20 hours before cells were fed with fresh media supplemented with 40-80mM MgClz for 24-48 hours before screening or reseeding in non-supplemented media.

[0331] For screening of insertion events, the targeted plasmids were extracted from the cells using Monarch® plasmid miniprep kit (NEB). The extracted plasmids were used as templates in standard PCR reactions, performed using Dreamtaq PCR mastermix under standard conditions. PCR amplification to confirm integration of the intron into the targeted locus used primers which bind to the 5'- and 3'-integration junctions. Bands of the correct size indicated successful insertions at the target site (Figure 5).

[0332] Example 2: Editing a Group II intron harboured I-Scel recognition site in HEK 293 cells

[0333] In this example a target plasmid containing an integrated intron, generated in example 1, was edited which contained a single I-Scel recognition sequence. To facilitate the edit, an I-Scel expression plasmid consisting of the constitutive PolII cytomegalovirus immediate early (CMV) promoter expressing an universal code equivalent I-Scel (from PSCM525, B.Dujon), with a C-terminal fusion of 2 x SV40 NLS and followed by a polyadenylation signal. This expression plasmid also contained an editing cassette, with 950 bp areas of homology designed that are either side of the intron insertion site separated by a I-Crel meganuclease recognition sequence. To confirm the generation of edited plasmids were absolutely dependent upon expression of I-Scel, a frameshifted I-Scel expression plasmid was also generated. Which was identical to the above expression plasmid, except a version with a premature stop codon in the I-Scel CDS (N30*) was used.

[0334] Editing experiments were performed in HEK 293 cells. These cells were maintained in DMEM, with high glucose and pyruvate, supplemented with 10% fetal bovine serum and 10% penicillin / streptomycin, at 37°C with 5% CO2 following standard cell culture practices. For intron integration the HEK 293 cells were seeded on a 12-well plate to around 80% confluency and cells were co-transfected using Lipofectamine 2000. Equal quantities of the editing expression plasmid and intron-containing target plasmid, totaling 1.5 pg of purified plasmid DNA was used. Cells were incubated for 16-20 hours before cells were fed with fresh media for 24-48 hours before screening.

[0335] For screening of edits, the target plasmids were harvested from cells using a Monarch® plasmid miniprep kit (NEB). The extracted plasmids were used as templates in standard PCR reactions, performed using Dreamtaq PCR mastermix under standard conditions. PCR amplification to confirm HDR of the I-Scel-induced DSBs to generate the edits used primers which bind to the 5'- and 3'-integration junctions. Bands of the correct size indicated successful insertions at the target site (Figure 6).

[0336] Example 3: Targeting a LINE1 genomic region in HEK 293 cells

[0337] In this example, group II intron targeting region was specific to the LINE1 gene of the human genome (GenBank Accession number: L19092.1). The Perutka algorithm (Perutka et al., 2004) was used to select the target site within LINE1, which was positioned between base pairs 2420|2421a, the 'a' denominating that the intron will insert in the anti-sense orientation relative to the gene target. The Perutka algorithm assigned this target insertion site a score of 9.43.

[0338] A eukaryotic single plasmid system for intron insertion was used to generate an intron insertion. This plasmid consists of LI.LtrB transcribed from a human Hl Pol III promoter. The intron LtrA ORF has been deleted, with an I-Scel meganuclease recognition site inserted in place of the deleted ORF, whilst maintaining the DIVb stem structure. Similarly an intron variant encoding a fluorescent marker was additionally created, in which alongside the meganuclease recognition site a miniGFP marker driven by a truncated CMV promoter and succeeded by an NRP-polyA signal was inserted in place of the deleted ORF. To enable expression from the Hl promoter, as transcripts are terminated at thymine repeats, the intron and marker sequences were altered to avoid strings of four repeat thymine nucleotides and to maintain the RNA structure. Following a T6 PolIII termination sequence, the human codon optimised LtrA is expressed from the PolII cytomegalovirus immediate early (CMV) promoter, fused to C-terminal 2 x SV40 NLS and followed by a polyadenylation signal.

[0339] Targeting experiments were performed in HEK 293 cells. These cells were maintained in DMEM, with high glucose and pyruvate, supplemented with 10% fetal bovine serum and 10% penicillin / streptomycin, at 37°C with 5% CO2 following standard cell culture practices. For intron integration the HEK 293 cells were seeded on a 12-well plate to around 80% confluency and cells were transfected with 1.5 pg of purified plasmid DNA using Lipofectamine 2000 (Invitrogen). Cells were incubated for 16-20 hours before cells were fed with fresh media supplemented with 40-80mM MgClz for 24 hours before screening or reseeding in non-supplemented media. For screening of insertion events, genomic DNA was extracted from harvested cells using a DNeasy blood and tissue kit (Qiagen). Standard PCR reactions were performed using the extracted genomic DNA using Dreamtaq PCR mastermix under standard conditions. PCR amplification to confirm integration of the intron into the targeted genomic locus used primers which bind to the 5'- and 3'-integration junctions. Bands of the correct size indicated successful insertions of both GFP-marked and unmarked intron at the target site (Figure 7).

[0340] Example 4: Strains and plasmids used in Examples 1-3 Table 1 - List of strains used herein

[0341] Strain Relevant characteristics Reference source

[0342] DH10B derivative - A(ara- eu) 7697 araD139 fhuA AiacX74 ga!K16 galELS

[0343] E. coli 10-beta el4~ ij>80diacZAM15 recAl NEB relAl endAl nupG rpsL (Str ) rph spoTl A(mrr-hsdRMS-mcrBC)

[0344] Human embryonic - .. ._. . i Culture Collection Strain ECACC kidney (HEK293)

[0345] Table 2 - List of plasmids used herein

[0346] Reference

[0347] Plasmid Relevant characteristics source

[0348] Forge editing vector with bla resistance marker on the backbone, targeting LI.LtrB WT This study, hFEl-WT-TR exons from Lactococcus lactis, containing an example 1

[0349] I-Scel recognition site.

[0350] Forge editing vector step 2 vector with bla resistance marker on the backbone, expressing I-Scel and an homology cassette This study, hFE2-I-SceI-pTR designed to edit the targeting plasmid, example 2 removing the intron and inserting a I-Crel recognition site. hFEl- Forge editing vector with bla resistance This study,

[0351] LINEl: :2420a marker on the backbone, targeting LINE1 example s

[0352] 2420a insertion site in human genome, containing an I-Scel recognition site.

[0353] Forge editing vector with bla resistance marker on the backbone, targeting LINE1 hFEl-LINEl- This study, 2420a insertion site in human genome,

[0354] GFP: :2420a example 3 containing a miniGFP marker and an I-Scel recognition site.

[0355] Sequences of the disclosure

[0356] SEQ ID NOs 1-142 are set out in Figure 2.

[0357] Table 3 - SEQ ID NOs: 143-154 & 158

[0358] SEQ ID

[0359] Oligo Function Nucleotide sequence 5'= >3'

[0360] NO

[0361] Screening primers:

[0362] Enables screening of the

[0363] Ll_scrn_Fl LINE1 genomic GCAAGTTGGATAAAGAGTCAAGA 143 locus, forward primer 1 Enables screening of the

[0364] Ll_scrn_F2 LINE1 genomic GCAATCCTAGTCTCTGATAAAACAG 144 locus, forward primer 2 Enables screening of the

[0365] Ll_scrn_Rl LINE1 genomic GCTTTGAGTGAGATTCTTAATCCTG 145 locus, reverse primer 1 Enables

[0366] Ll_scrn_R2 screening of the GCTTTACTTCCAACTATGTGGTCAA 146 LINE1 genomic locus, reverse primer 2

[0367] Enables screening of the

[0368] P-Fl targeting plasmid CACCTCGCTAACGGATTCACCG 147 locus, forward primer 1

[0369] Enables screening of the

[0370] P-Rl targeting plasmid CGCACATTTCCCCGAAAAGTGC 148 locus, reverse primer 1

[0371] Enables screening of the intron insertion

[0372] Intron-Fl GTACAATCTGTAGGAGAACCTATG 149 into the targeting locus, forward primer 1 Enables screening of the intron insertion

[0373] Intron-F2 GGGATACCCTAAACAAGAATGC 150 into the targeting locus, forward primer 2

[0374] Intron-F3 Enables AGGGATAACAGGGTAATGACG 158 screening of the intron insertion into the targeting locus, forward primer 3

[0375] Enables screening of the intron insertion

[0376] Intron-Rl CATAGGTTCTCCTACAGATTGTAC 151 into the targeting locus, reverse primer 1 Enables screening of the intron insertion

[0377] Intron-R2 ATTACCCTGTTATCCCTAACGC 152 into the targeting locus, reverse primer 2

[0378] Enables screening of the

[0379] Crel-Fl TTCAAAACGTCGTGAGACAGTTTGG 153 edit, forward primer 1

[0380] Enables screening of the Crel-Rl GTCTCACGACG 1 1 1 1 GAACCCAG 154 edit, reverse primer 1

[0381] >SEQ_ID_NO:_155_Eukaryotic_single_plasmid_system_LI.LtrB_with_LINEl_targetin g_sequence

[0382] ATAATTATCCTTAGTGTTCAAGTCTGTGCGCCCAGATAGGGTGTTAAGTCAAGTAGTTTAAGG

[0383] TACTACTCTGTAAGATAACACAGAAAACAGCCAACCTAACCGAAAAGCGAAAGCTGATACGG

[0384] GAACAGAGCACGGTTGGAAAGCGATGAGTTACCTAAAGACAATCGGGTACGACTGAGTCGC AATGTTAATCAGATATAAGGTATAAGTTGTGTTTACTGAACGCAAGTTTCTAATTTCGATTAAC

[0385] ACTCGATAGAGGAAAGTGTCTGAAACCTCTAGTACAAAGAAAGGTAAGTTAGGAGACTTGAC

[0386] TTATCTGTTATCACCACATTTGTACAATCTGTAGGAGAACCTATGGGAACGAAACGAAAGCGA

[0387] TGCCGAGAATCTGAATTTACCAAGACTTAACACTAACTGGGGATACCCTAAACAAGAATGCCT AATAGAAAGGAGGAAAAAGGCTATAGCACTAGAGCTTGAAAATCTTGCAAGGGTACGGAGT

[0388] ACTCGTAGTAGTCTGAGAAGGGTAACGCCCTTTACATGGCAAAGGGGTACAGTTATTGTGTA CTATAATTAAAAATTGATTAGGGAGGAAAACCTCAAAATGAAACCAACAATGGCAATTATAGA

[0389] AAGAATCAGTAATAATTCACAACGCG7TAGGGATAACAGGGTAATCGCG7AGTGAATTATT

[0390] ACGAACGAACAATAACAGAGCCGTATACTCCGAGAGGGGTACGTACGGTTCCCGAAGAGGG TGGTGCAAACCAGTCACAGTAATGTGAACAAGGCGGTACCTCCCTACTTCAC

[0391] Underlined - targeting region variable sequences (IBS2, IBS1, EBS2, EBS1).

[0392] Bold and underlined- I-Scel recognition site.

[0393] Italiscs - Spacer sequence around I-Scel site.

[0394] Bold- nucleotides altered to avoid strings of thymines and preserve RIMA structure.

[0395] >SEQ_ID_NO:_156_human_optimised_LtrA_sequence ATGAAGCCCACCATGGCCATCCTGGAGCGCATCAGCAAGAACTCACAGGAGAACATCGACGAGGTGTT

[0396] CACCCGCCTGTACCGCTACCTGCTGCGCCCCGACATCTACTACGTGGCCTACCAGAACCTGTACAGCAA

[0397] CAAGGGCGCCAGCACCAAGGGCATCCTGGACGACACCGCCGACGGCTTCAGCGAGGAGAAGATCAAG

[0398] AAGATCATCCAGAGCCTGAAGGACGGCACCTACTACCCCCAGCCCGTGCGCCGCATGTACATCGCCAA

[0399] GAAGAACAGCAAGAAGATGCGCCCCCTGGGCATCCCCACCTTCACCGACAAGCTGATCCAGGAGGCCG

[0400] TGAGGATCATCCTGGAGTCCATCTACGAGCCCGTGTTCGAGGACGTGAGCCACGGCTTCCGCCCCCAG

[0401] CGCAGCTGCCACACCGCCCTGAAGACCATCAAGCGCGAGTTCGGCGGCGCCAGGTGGTTCGTGGAGG

[0402] GCGACATCAAGGGCTGCTTCGACAACATCGACCACGTGACCCTGATCGGCCTGATCAACCTGAAGATCA

[0403] AGGACATGAAGATGAGCCAGCTGATCTACAAGTTCCTGAAGGCCGGCTACCTGGAGAACTGGCAGTAC

[0404] CACAAGACCTACAGCGGCACCCCCCAGGGCGGCATCCTGAGCCCCCTGCTGGCCAACATCTACCTGCA

[0405] CGAGTTGGACAAGTTCGTGCTGCAGCTGAAGATGAAGTTCGACCGCGAGAGCCCCGAGCGCATCACCC

[0406] CCGAGTACCGCGAGCTGCACAACGAGATCAAGCGCATCAGCCACCGCCTGAAGAAGCTGGAGGGCGA

[0407] GGAGAAGGCCAAGGTCCTGCTGGAGTACCAGGAGAAGCGCAAGAGGCTGCCCACCCTCCCCTGCACC

[0408] AGCCAGACCAACAAGGTGCTGAAGTACGTGCGCTACGCCGACGACTTCATCATCAGCGTGAAGGGCAG

[0409] CAAGGAGGACTGCCAGTGGATCAAGGAGCAGCTGAAGCTGTTCATCCACAACAAGCTGAAGATGGAGC

[0410] TGAGCGAGGAGAAGACCCTGATCACCCACAGCAGCCAGCCCGCCCGCTTCCTGGGCTACGACATCCGC

[0411] GTGCGCCGCAGCGGCACCATCAAGCGCAGCGGCAAGGTGAAGAAGCGCACCCTGAACGGCAGCGTGG

[0412] AGCTGCTGATCCCCCTGCAGGACAAGATCCGCCAGTTCATCTTCGACAAGAAGATCGCCATCCAGAAGA

[0413] AGGACAGCAGCTGGTTCCCCGTGCACCGCAAGTACCTGATCCGCAGCACCGACCTGGAGATCATCACC

[0414] ATCTACAACAGCGAGCTGAGGGGCATCTGCAACTACTACGGCCTGGCCAGCAACTTCAACCAGCTGAA

[0415] CTACTTCGCCTACCTGATGGAGTACAGCTGCCTGAAGACCATCGCCTCCAAGCACAAGGGCACCCTGAG

[0416] CAAGACCATCTCCATGTTCAAGGACGGCAGCGGCAGCTGGGGCATCCCCTACGAGATCAAGCAGGGCA

[0417] AGCAGCGCCGCTACTTCGCCAACTTCAGCGAGTGCAAGTCCCCCTACCAGTTCACCGACGAGATCAGC

[0418] CAGGCCCCCGTGCTGTACGGCTACGCCCGCAACACCCTGGAGAACCGCCTGAAGGCCAAGTGCTGCGA

[0419] GCTGTGCGGCACCAGCGACGAGAACACCAGCTACGAGATCCACCACGTGAACAAGGTGAAGAACCTGA

[0420] AGGGCAAGGAGAAGTGGGAGATGGCCATGATCGCCAAGCAGCGCAAGACCCTGGTGGTGTGCTTCCA

[0421] CTGCCACCGCCACGTGATCCACAAGCACAAG

[0422] >SEQ_ID_NO:_157_I-SceI_sequence

[0423] ATGAAAAACATCAAAAAAAACCAGGTAATGAACCTGGGTCCGAACTCTAAACTGCTGAAAGAATACAAA

[0424] TCCCAGCTGATCGAACTGAACATCGAACAGTTCGAAGCAGGTATCGGTCTGATCCTGGGTGATGCTTAC

[0425] ATCCGTTCTCGTGATGAAGGTAAAACCTACTGTATGCAGTTCGAGTGGAAAAACAAAGCATACATGGAC

[0426] CACGTATGTCTGCTGTACGATCAGTGGGTACTGTCCCCGCCGCACAAAAAAGAACGTGTTAACCACCTG

[0427] GGTAACCTGGTAATCACCTGGGGCGCCCAGACTTTCAAACACCAAGCTTTCAACAAACTGGCTAACCTG

[0428] TTCATCGTTAACAACAAAAAAACCATCCCGAACAACCTGGTTGAAAACTACCTGACCCCGATGTCTCTGG

[0429] CATACTGGTTCATGGATGATGGTGGTAAATGGGATTACAACAAAAACTCTACCAACAAATCGATCGTACT

[0430] GAACACCCAGTCTTTCACTTTCGAAGAAGTAGAATACCTGGTTAAGGGTCTGCGTAACAAATTCCAACTG

[0431] AACTGTTACGTAAAAATCAACAAAAACAAACCGATCATCTACATCGATTCTATGTCTTACCTGATCTTCTA

[0432] CAACCTGATCAAACCGTACCTGATCCCGCAGATGATGTACAAACTGCCGAACACTATCTCCTCCGAAACT

[0433] TTCCTGAAA

[0434] >SEQ_ID_NO:_159_miniGFP_sequence ATGGAGAAGTCCTTTGTGATTACAGACCCATGGCTGCCCGACTACCCTATTATTTCCGCTAGT

[0435] GACGGATTCCTGGAGCTGACCGAGTATTCTCGGGAGGAAATCATGGGAAGGAACGCACGCT

[0436] TCCTGCAGGGACCCGAGACTGACCAGGCCACCGTGCAGAAGATCCGGGACGCTATTCGAGA

[0437] TCGGAGACCTACCACAGTCCAGCTGATCAACTACACAAAGAGCGGCAAGAAGTTCTGGAATC TGTTGCACCTGCAGCCTGTGTTTGACGGAAAAGGCGGGCTGCAATATTTCATTGGAGTGCAG

[0438] CTGGTCGGCTCCGATCATGTCTAG

Claims

CLAIMS1. An ex vivo or in vitro method of editing one or more mammalian cells, optionally one or more human cells, said method comprising:(a) providing a mammalian cell or a plurality of mammalian cells that have been obtained from a mammal;(b) introducing into the mammalian cell or plurality of mammalian cells one or more insertion cassettes; and(c) incubating the mammalian cell in conditions that permit integration of the insertion nucleic acid into the target nucleic acid wherein the insertion cassette is a nucleic acid comprising at least: i) a sequence that encodes a group II intron; ii) a sequence that encodes a selectable marker operably linked to a selectable marker promoter; and iii) a nucleic acid sequence that is a cleavage site for a nuclease.

2. The method of claim 1, wherein the selectable marker is a fluorescent protein.

3. The method of claim 1 or 2, wherein the nucleic acid sequence that is a cleavage site for a nuclease is a cleavage site for a meganuclease.

4. The method of any one of claims 1-3 wherein the method further comprises step (d) detecting mammalian cells which express the fluorescent protein and selecting those cells for further processing, optionally wherein expression of the fluorescent protein is performed using a FACS.

5. The method of claim 4 further comprising step (e) introducing into the cells selected for further processing in step (d): i) a nuclease expressing cassette; ii) an editing and nuclease expressing cassette; or iii) an editing complex wherein: the nuclease expressing cassette comprises a nucleic acid sequence that encodes a meganuclease operably linked to a promoter, wherein the meganucleasecorresponds to the nucleic acid sequence that is a cleavage site for a nuclease present in the insertion cassette; the editing cassette comprises a nucleic acid with at least a first region that has homology to a first portion of the target nucleic acid, optionally also comprises at least a second region that has homology to a second portion of said target nucleic acid; wherein said first and second regions have sufficient homology to the first and second portions of the target nucleic acid so as to allow repair of a double stranded break using homology directed repair (HDR); the editing and nuclease expressing cassette comprises a nucleic acid sequence that encodes a nuclease operably linked to a promoter, wherein the meganuclease corresponds to the nucleic acid sequence that is a cleavage site for a nuclease present in the insertion cassette, and also has at least a first region that has homology to a first portion of the target nucleic acid, optionally also comprises at least a second region that has homology to a second portion of said target nucleic acid; wherein said first and second regions have sufficient homology to the first and second portions of the target nucleic acid so as to allow repair of a double stranded break using homology directed repair (HDR); the editing complex comprises a nucleic acid with at least a first region that has homology to a first portion of the target nucleic acid, optionally also comprises at least a second region that has homology to a second portion of said target nucleic acid; wherein said first and second regions have sufficient homology to the first and second portions of the target nucleic acid so as to allow repair of a double stranded break using homology directed repair (HDR), wherein the nucleic acid is complexed with a meganuclease that corresponds to the nucleic acid sequence that is a cleavage site for a nuclease present in the insertion cassette and step (f) incubating the mammalian cell or mammalian plurality of cells under conditions that a) permit expression of the meganuclease in the target cell or plurality of target cells; and b) permit homologous recombination between a region of the editing cassette that has homology to the target nucleic acid.

6. The method of claim 5 comprising the further step (g) of selecting cells following incubation that are viable and that do not fluoresce, optionally wherein said selecting is performed using FACS.

7. The method of any of the preceding claims wherein the sequence that encodes a selectable marker, selectable marker promoter and the nucleic acid sequence that is a cleavage site for a nuclease are arranged within the insertion cassette such that upon insertion of the Group II intron a target nucleic acid, optionally into a target cell genome, all of the sequence that encodes a selectable marker, selectable marker promoter and the nucleic acid sequence that is a cleavage site for a nuclease are inserted into the target cell genome / nucleic acid.

8. The method of any of claims 1-7 wherein the fluorescent protein is selected from the group comprising or consisting of: blue fluorescent protein, green fluorescent protein, yellow fluorescent protein, and red fluorescent protein, optionally wherein: a)the blue fluorescent protein is selected from the group comprising or consisting of: TagBFP, mTagBFP2, EBFP, EBFP2, Azutire, mKamalal, cyan fluorescent protein, ECFP, Cerulian, CyPet, mTurquoise, mTurquoise2, and AmCyanl; b)the green fluorescent protein is selected from the group comprising or consisting of: miniGFPl, GFP, EGFP, mNeonGreen, and mGFP; c)the yellow fluorescent protein is selected from the group comprising or consisting of: YFP, Citrine, Venue, and YPet; and / or d)the red fluorescent protein is selected from the group comprising or consisting of: mCherry, mCardinal, mScarlet, DsRed, mOrange, mRaspberry, mFruit, mK, TagRFP, mKate, mRuby, FusionRed, and DsRed-Express.

9. The method of any of the preceding claims wherein the sequence that encodes a selectable maker is only expressed once the transposition of the Group II intron and the selectable maker into a target nucleic acid has occurred.

10. The method of any of the preceding claims wherein the sequence that encodes a selectable marker comprises a removable intervening sequence arranged such that when the removable sequence is present expression from the selectable marker promoter does not produce a functional protein, and wherein removal of the removable intervening sequence allows expression of a functional protein or RNA, optionally wherein the selectable marker is arranged as a RAM, optionally wherein the removable intervening sequence is a group I intron, optionally wherein the group I intron is flanked by exons that enable self-splicing of the group I intron.

11. The method of any of the preceding claims wherein the selectable marker and / or the selectable marker promoter are present in the cassette in the antisense orientation, such that:a) direct transcription from the selectable marker promoter from the insertion cassette is not possible; and / or b) translation of a transcript produced from the selectable marker promoter from the insertion cassette does not produce a functional selectable marker protein or RNA.

12. The method of any of the preceding claims wherein meganuclease is selected from the group comprising or consisting of:I-Scel, I-Ppol, I-Dmol, I-Njal, I-Dirl, I-Bmol, BneMS4ORFIP, F-CphI, F-EcoT3I, F- ECOT5I, F-ECOT5II, F-EcoT5IV, F-PhiU5I, F-Scel, F-Scell, F-TevI, F-TevII, F-TevIII, F- TevIV, H-Drel, H-Drel, I-AabMI, I-AchMI, I-Anil, I-ApeKI, I-BanI, I-BasI, I-Bth0305I, I-BthII, I-BthORFAP, I-Ceul, I-Chul, I-Cmoel, I-Cpal, I-CpaII, I-CpaMI, I-Crel, I- Crell, I-CsmI, I-Cvul, I-Ddil, I-GpeMI, I-Gpil, I-Gzel, I-GzeII, I-HjeMI, I-Hmul, I- HmuII, I-Llal, I-Ltrl, I-LtrWI, I-MpeMI, I-Msol, I-NanI, I-Nfil, I-Nitl, I-Omill, I-Onul, I-PakI, I-PanMI, I-PfoP3I, I-PnoMI, I-PogTE7I, I-Porl, I-Scal, I-SceII, I-SceIII, I- ScelV, I-SceV, I-SceVI, I-SceVII, I-SecIII, I-SmaMI, I-SpomI, I-SscMI, I-Ssp6803I, I- TevI, I-TevII. I-TevIII. I-TslI. I-TsIWI, I-TspO61I, I-Twol, I-Vdil41I, -Aval, PI-BciPI, PI-HvoWI, PI-MgaI„ PI-MleSI, PI-MtuI, Pl-PabI, Pl-Pabll, Pl-Pful, Pl-PfuII, Pl-Pkol, PI- PkoII, PI-PspI, PI-PspI, Pl-Scal, Pl-Scel, PI-Tful, PI-TfuII, PI-Thyl, PI-Tlil, PI-Tlill, PI- Tmal, PI-TmaKI, Pl-Zba, or has a sequence of any of SEQ ID NO: 1-142; or a sequence that is: at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NOs. 1-142.

13. The method of any of the preceding claims where the insertion cassette comprises at least 2, 3, 4, 5 or more sequences that are cleavage sites for a meganuclease, optionally a) wherein the at least at least 2, 3, 4, 5 or more sequences that are cleavage sites for a meganuclease all have the same sequence and / or are cleavage sites for the same meganuclease; or b) wherein the at least 2 of the at least 2, 3, 4, 5 or more sequences that are cleavage sites for a meganuclease are different sequences and / or are cleavage sites for at least two different meganucleases.

14. The method of any of the preceding claims where the insertion cassette: does not encode a protein comprising reverse transcriptase activity; and / or does not have the ability to self-excise.

15. The method of any of the preceding claims where the insertion cassette further comprises a portion comprising a second promoter operably linked to a nucleic acid sequence encoding an intron-encoded protein (IEP), optionally i) wherein the second promoter is a eukaryotic promoter, optionally a mammalian promoter, optionally a human promoter; and / or ii) the IEP is LtrA, optionally wherein LtrA comprises a nucleotide sequence substantially as set out in SEQ ID NO: 156; or a sequence that is: a)at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NO: 156; b)about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% identical to SEQ ID NO: 156; and / or c)75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 156, wherein expression of the IEP causes: i) an RNA transcript from the insertion cassette promoter to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand; or ii) where the insertion cassette is an RNA molecule, the insertion cassette to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand.

16. The method of claim 13 wherein the IEP does not confer the ability to self-splice on the Group II intron.

17. The method of any of claims 1-14, wherein the insertion cassette is encoded by a nucleic acid that further comprises a nucleotide sequence encoding an accessory protein, wherein the accessory protein is not comprised by the Group II intron.

18. The method of claim 17, wherein the nucleic acid encoding the accessory protein is operably linked to a second promoter.

19. The method of claim 18, wherein the second promoter is a different promoter to the first promoter.

20. The method of any one of claims 17-19, wherein the accessory protein is LtrA, optionally wherein LtrA comprises a nucleotide sequence substantially as set out in SEQ ID NO: 156; or a sequence that is: a)at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or more identical to SEQ ID NO: 156; b)about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% identical to SEQ ID NO: 156; and / or c)75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to SEQ ID NO: 156; wherein expression of the accessory protein causes: i) an RNA transcript from the insertion cassette promoter to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand; or ii) where the insertion cassette is an RNA molecule, the insertion cassette to become inserted into a target locus in a target nucleic acid and optionally reverse transcribed to produce the complementary DNA strand.

21. The method of any of the preceding claims wherein the group II intron comprises LI.LtrB from Lactococcus lactis, Ecl5 from Escherichia coli, Rmlntl from Sinorhizobium meliloti, or TeI3c from Thermosynechococcus elongatus.

22. An insertion cassette that is a nucleic acid comprising at least: i) a sequence that encodes a group II intron; ii) a sequence that encodes a selectable marker operably linked to a selectable marker promoter; and iii) a nucleic acid sequence that is a cleavage site for a nuclease.

23. An insertion complex comprising an RNA transcript from the insertion cassette of claim 22 and one or more proteins that allow the insertion of the transcript into the target nucleic acid, optionally an IEP as described in any of the preceding claims.

24. An editing cassette comprising at least a first region that has homology to a first portion of a target nucleic acid, optionally also comprises at least a second region that has homology to a second portion of said target nucleic acid; wherein said first and second regions have sufficient homology to the first and second portions of thetarget nucleic acid so as to allow repair of a double stranded break using homology directed repair (HDR).

25. The editing cassette of claim 24 wherein the editing cassette comprising a further portion that is a nucleic acid that is an aptamer with binding specificity for a nuclease, optionally wherein the editing cassette is a single stranded DNA molecule.

26. An editing complex comprising the editing cassette of claim 24 or 25 complexed with a nuclease protein, optionally wherein the editing cassette is the editing cassette of claim 25.

27. A mammalian nucleic acid editing system comprising any two or more of:The insertion cassette according to claim 22;The insertion complex of claim 23;The editing cassette of any of claims 24 or 25;A nuclease expressing cassette comprising a nucleic acid sequence that encodes a meganuclease operably linked to a promoter, wherein the meganuclease corresponds to the nucleic acid sequence that is a cleavage site for a nuclease present in the insertion cassette; and / orThe editing complex of claim 26.

28. The mammalian nucleic acid editing system of claim 27 comprising:A) The insertion cassette according to any of claims 22 and: i) the editing cassette of any of claims 24 or 25 and the nuclease expressing cassette; or ii) the editing complex of claim 26; orB) The insertion complex of claim 23 and: i) the editing cassette of any of claims 24 or 25 and the nuclease expressing cassette; or ii) the editing complex of claim 26.

29. A nucleic acid comprising the insertion cassette of claim 22.

30. The nucleic acid of claim 29, further comprising a nucleotide sequence encoding an accessory protein, wherein the accessory protein is not comprised by the Group II intron.

31. The nucleic acid of claim 30, wherein the nucleic acid is operably linked to a second promoter.

32. The nucleic acid of any one of claims 30-31, wherein the accessory protein is an accessory protein of claim 20.

33. A population of edited cells that have been edited according to the method of any one of claims 1-21 wherein: a) the population of edited cells comprise no cells with an off-target gene edit event and / or b) the population of cells has 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cells comprise a desired editing event and no off-target editing events.

34. A pharmaceutical composition comprising a population of edited cells according to claim 33 optionally wherein said composition comprises one or more pharmaceutically acceptable excipients or adjuvants.

35. A population of edited cells according to claim 33 or a pharmaceutical composition of claim 24 for use in a method of treatment wherein the method comprising administering the population of edited cells or a pharmaceutical composition to a subject in need thereof.

Citation Information

Patent Citations

  • DNA molecules and methods

    US10544422B2

  • Recruitment of donor DNA from in VIVO assembled plasmids for saturation genome editing

    WO2024044767A2