Novel gene editing systems
The gene editing system uses a polypeptide and RNA molecule to achieve precise and efficient insertion of nucleic acid cargo at target sites in prokaryotic cells, addressing specificity and efficiency issues in existing technologies.
Patent Information
- Application Number
- PCT/AU2025/050149
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2025-02-21
- Publication Date
- 2025-08-28
AI Technical Summary
Existing gene editing systems lack specificity and efficiency in inserting nucleic acid cargo at designated target sites in genomes, often leading to undesired mutagenic consequences and low insertion efficiency, particularly for large nucleic acid cargos.
A gene editing system comprising a polypeptide and an RNA molecule that specifically bind to a target site, enabling the insertion of a nucleic acid cargo with a left and right end sequence to form a circular nucleic acid molecule, allowing scarless or near scarless insertion into prokaryotic cells.
Enables precise and efficient insertion of nucleic acid cargo at designated target sites without undesired mutagenic effects, supporting large nucleic acid cargo insertion with high specificity and efficiency.
Smart Images

Figure IMGF000083_0001 
Figure IMGF000041_0001 
Figure IMGF000042_0001
Abstract
Description
Novel gene editing systems
[0001] This application claims priority from Australian provisional patent applications 2024900444 and 2024901154, the entire contents of both of which are incorporated herein by reference.Field of the disclosure
[0002] The present disclosure relates to systems, vectors, and compositions for gene editing via the site-specific insertion of gene sequences, and methods of use thereof. In particular, the present disclosure relates to gene editing systems derived from insertion sequences and associated transposases.Background of the disclosure
[0003] Mobile genetic elements (MGE) found in bacteria encode a broad range of DNA-processing enzymes, some of which have been harnessed for gene editing applications. However, a majority of MGEs that rely upon a transposition mechanism, i.e. insertion sequences (IS) and transposons, exhibit little or no selectivity regarding the target site to be edited in a genome of interest, and little or no specificity for the orientation in which genetic material is inserted into a target site.
[0004] In contrast, gene editing systems that are not based upon mobile gene elements (eg CRISPR / Cas, TALENs) may exhibit target site selectivity, but often require the cleavage of DNA prior to insertion of a gene sequence at the target site. DNA cleavage triggers DNA repair processes (such as non-homologous end joining and homologous recombination), which can result in the unwanted insertion and / or deletion of nucleotides (eg high rates of indels; insertional mutagenesis) during insertion of the desired gene sequence into the target site. Gene insertion approaches relying upon double-stranded DNA cleavage are also hindered by low insertion efficiency and limitations regarding the size of the nucleic acid cargo that can be inserted.
[0005] Therefore, there is a need to develop gene editing systems to precisely insert nucleic acid cargo at a designated target site without having undesired mutagenic or deleterious consequences; particularly for insertion of large nucleic acid cargos and for where the target site location can be specified.
[0006] Reference to any prior art in the specification is not an acknowledgment or suggestion that this prior art forms part of the common general knowledge in any jurisdiction or that this prior art could reasonably be expected to be understood, regarded as relevant, and / or combined with other pieces of prior art by a skilled person in the art.Summary of the disclosure
[0007] The disclosure provides a gene editing system for inserting a nucleic acid cargo comprising a nucleotide sequence of interest into a genome of a prokaryotic cell, the system comprising:(i) a polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40, or a polynucleotide encoding the same; and ii) an RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell, and / or a polynucleotide sequence encoding the same; wherein the nucleic acid cargo comprises a left end (LE) nucleotide sequence and a right end (RE) nucleotide sequence to enable formation of a circular nucleic acid molecule; wherein the nucleic acid cargo comprises in the 5’ to 3’ direction: a left flank (LF) sequence corresponding to the target site in the genome of the cell, the LE, the nucleotide sequence of interest, the RE, and a right flank (RF) sequence corresponding to the target site in the genome of the cell; wherein the polypeptide, fragment or functional equivalent thereof of (i) and the RNA molecule of (ii) are capable of forming a gene editing complex for enabling the insertion of the nucleic acid cargo into the target site within the genome.
[0008] The disclosure further provides a gene editing system for inserting a nucleic acid cargo comprising a nucleotide sequence of interest into a genome of a prokaryotic cell, the system comprising:(i) a polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40, or a polynucleotide encoding the same; and ii) an RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell, and / or a polynucleotide sequence encoding the same; wherein the nucleic acid cargo comprises a left end (LE) nucleotide sequence and a right end (RE) nucleotide sequence that are concatenated to form of a circular nucleic acid molecule (eg artificial minicircle vector); wherein prior to concatenation, the nucleic acid cargo comprises in the 5' to 3’ direction: a left flank (LF) sequence corresponding to the target site in the genome of the cell, the LE, the nucleotide sequence of interest, the RE, and a right flank (RF) sequence corresponding to the target site in the genome of the cell; wherein the polypeptide, fragment or functional equivalent thereof of (i) and the RNA molecule of (ii) are capable of forming a gene editing complex for enabling the insertion of the nucleic acid cargo into the target site within the genome.
[0009] In any embodiment of any aspect of the disclosure herein, the gene editing system enables scarless, or near scarless, insertion of the nucleic acid cargo into the target site within the genome.
[0010] In any embodiment of any aspect of the disclosure herein, the polypeptide, fragment or functional equivalent thereof has an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the amino acid sequence set forth in any one of SEQ ID NOs: 33-40.
[0011] In preferred embodiments, the polypeptide, fragment or functional equivalent thereof is encoded by a polynucleotide sequence comprising or consisting of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 41 -48.
[0012] In any embodiment of any aspect of the disclosure herein, the polypeptide, fragment or functional equivalent thereof is encoded by a polynucleotide sequence comprising or consisting a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence as set forth in any one of SEQ ID NOs: 1 -16.
[0013] In any embodiment of any aspect of the disclosure herein, the polypeptide, fragment or functional equivalent thereof is a transposase (Tnp), preferably derived from an IS 1111 or IS 110 insertion sequence. In some embodiments, the transposase is selected from: TnpEc11 , TnpKpn4, TnpPst6, TnpPal 1 , TnpEc21 , TnpXne4, TnpPa25, or TnpMch6. In some embodiments, the transposase is TnpEc11 , TnpKpn4, TnpPst6, TnpPal 1 , or TnpEc21 . In some embodiments, the transposase is TnpEd 1 , TnpKpn4, or TnpEc21 . Preferably, the transposase is TnpKpn4 or TnpEc21 .
[0014] In any embodiment of any aspect of the disclosure herein, the polypeptide, fragment or functional equivalent thereof of (i), is an isolated, synthetic or recombinant protein, preferably a purified protein.
[0015] In any embodiment of any aspect of the disclosure herein, the polypeptide, fragment or functional equivalent thereof is or has been codon optimised for expression in a prokaryotic cell.
[0016] In any embodiment of any aspect of the disclosure herein, the RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell, and / or a polynucleotide sequence encoding the same of (ii), is derived from a native or a modified IS 1111 or IS 110 insertion sequence that encodes an RNA molecule. In some embodiments, the IS 1111 or IS 110 insertion sequence is selected from: ISEd 1 , ISKpn4, ISPst6, ISPal 1 , ISEc21 , ISXne4, ISPa25, or ISMch6. In some embodiments, the insertion sequence is ISEd 1 , ISKpn4, ISPst6, ISPal 1 , or ISEc21 . In some embodiments, the insertion sequence is ISEd 1 , ISKpn4, or ISEc21 . Preferably, the insertion sequence is ISKpn4 or ISEc21. Exemplary insertion sequences are shown in Tables 1 -2.
[0017] In any embodiment of any aspect of the disclosure herein, the RNA molecule that specifically binds or is capable of specifically binding to a target site within thegenome of the cell, and / or a polynucleotide sequence encoding the same of (ii), is an isolated, synthetic or recombinant RNA molecule or polynucleotide.
[0018] In any embodiment of any aspect of the disclosure herein, the polynucleotide encoding an RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to an insertion sequence set forth in any one of SEQ ID NOs: 1 -16; optionally any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, or 16.
[0019] In any embodiment of any aspect of the disclosure herein, a polypeptide, fragment or functional equivalent thereof and / or the RNA molecule are encoded by an insertion sequence selected from: ISEc11 , ISKpn4, ISPst6, ISPal 1 , ISEc21 , ISXne4, ISMch6, ISPa25. In some embodiments, the insertion sequence is ISEc11 , ISKpn4, ISPst6, ISPal 1 , or ISEc21 . Preferably, the insertion sequence is ISEc11 , ISKpn4, or ISEc21 ; more preferably ISEc11 or ISKpn4.
[0020] In some embodiments, the polynucleotide encoding an RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 17-24.
[0021] In preferred embodiments, the polynucleotide encoding an RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 25-32.
[0022] In any embodiment of any aspect of the disclosure herein, the RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 50-57.
[0023] In any embodiment of any aspect, the system does not require a “core” sequence of 1 -5 nucleotides, such as the nucleotides “CT”, to be present at theboundaries of any of the LE, RE, LF or RF elements, for example between the LF and LE elements, or the RE and RF elements.
[0024] In any embodiment of any aspect of the disclosure described herein, the target site comprises one or more sequences comprising or consisting of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the target site of an insertion sequence set forth in any one of SEQ ID NOs: 1 -16; optionally any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, or 16. In some embodiments, the target site comprises one or more sequences comprising or consisting of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 58-131 ; optionally any one of SEQ ID NOs: 66, 67, 76, 82, 112, 115, 127 and 129.
[0025] A target site may comprise a native target site for a polypeptide / RNA combination as herein described, or may include any non-native target site. For example, it will be understood by the skilled person that it is possible to re-direct (ie “reprogramme”) any genome editing system as described herein, to enable genome editing at a non-native target site (eg a different target site to one specified in any of SEQ ID NOs: 58-131 ; optionally any one of SEQ ID NOs: 66, 67, 76, 82, 112, 115, 127 and 129). Typically, all that is required is for the skilled person to identify a desired site in the genome into which the nucleic acid cargo is to be inserted, and having identified the desired site, i) modifying the target-site binding regions of the RNA molecule so that they are complementary to the nucleic acid sequence of the desired target site; and ii) adding the target site sequence to the 5’ and 3’ ends of the nucleic acid cargo.
[0026] In preferred embodiments, the non-native target site may be any nucleic acid sequence in the target genome, typically comprising between at least about 5 - 15 nucleotides and having at least 5 nucleotides. Native target sites of IS 1111 and IS 110 family members could be used to design non-native target sites. The non-native target site may comprise or consist of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQID NOs: 58-131 ; optionally any one of SEQ ID NOs: 66, 67, 76, 82, 112, 115, 127 and 129.
[0027] In some embodiments, the non-native target site may be modified such that transposition efficiency is improved. The skilled person will appreciate appropriate methods to assess transposition efficiency, including the assays described herein for example in Figure 6D; Examples 7-9.
[0028] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:50 or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 17), wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 58-66, optionally SEQ ID NO:66.
[0029] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:51 , or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 18); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 67-75, optionally SEQ ID NO:67.
[0030] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:52, or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 19); wherein the target site comprises or consists of the polynucleotide set forth in any one of SEQ ID NOs: 76-81 , optionally SEQ ID NO:76.
[0031] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:53, or a polynucleotide encoding the same (such as SEQ ID NO: 20);wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 82-111 , optionally SEQ ID NO:82.
[0032] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:54, or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 21 ); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 112-114, optionally SEQ ID NO:1 12.
[0033] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:55, or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 22); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 115-126, optionally SEQ ID NO:1 15.
[0034] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:56, or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 23); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 127-128, optionally SEQ ID NO:127.
[0035] In some embodiments, the gene editing system may comprise:(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:57, or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 24); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 129-131 , optionally SEQ ID NO:129.
[0036] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 33 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:50, or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 17); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 58-66, optionally SEQ ID NO:66.
[0037] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 34 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:51 , or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 18); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 67-75, optionally SEQ ID NO:67.
[0038] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 35 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:52 or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 19); wherein the target site comprises or consists of the polynucleotide set forth in any one of SEQ ID NOs: 76-81 , optionally SEQ ID NO:76.
[0039] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 36 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:53 or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 20); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 82-111 , optionally SEQ ID NO:82.
[0040] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 37 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:54 or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 21 ); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 112-114, optionally SEQ ID NO:1 12.
[0041] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 38 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:55, or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 22); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 115-126, optionally SEQ ID NO:1 15.
[0042] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 39 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:56 or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 23); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 127-128, optionally SEQ ID NO:127.
[0043] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO: 40 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising or consisting of the sequence set forth in SEQ ID NO:57 or functional equivalent thereof or a polynucleotide encoding the same (such as SEQ ID NO: 24); wherein the target site comprises or consists of the polynucleotide sequence set forth in any one of SEQ ID NOs: 129-131 , optionally SEQ ID NO:129.
[0044] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO:33 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NQ:50 or encoded by a polynucleotide comprising or consisting of the sequence set forth in SEQ ID NO:17, or a polynucleotide encoding the same; wherein the nucleotides at positions 31 to 40 and 57 to 65 (or nucleotides at a position equivalent thereto) of SEQ ID NO:25 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0045] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO:34 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NO:51 or encoded by a polynucleotide comprising or consisting of the sequence set forth in SEQ ID NO:18, or a polynucleotide encoding the same; wherein the nucleotides at positions 32 to 38 and 57 to 63 (or nucleotides at a position equivalent thereto) of SEQ ID NO:26 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0046] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO:35 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NO:52 or encoded by a polynucleotide comprising or consisting of the sequence set forth in SEQ ID NO:19, or a polynucleotide encoding the same; wherein the nucleotides at positions 30 to 35 and 54 to 62 (or nucleotides at a position equivalent thereto) of SEQ ID NO:27 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0047] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO:36 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NO:53 or encoded by a polynucleotide comprising orconsisting of the sequence set forth in SEQ ID NO:20, or a polynucleotide encoding the same; wherein the nucleotides at positions 32 to 38 and 70 to 78 (or nucleotides at a position equivalent thereto) of SEQ ID NO:28 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0048] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO:37 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NO:54 or encoded by a polynucleotide comprising or consisting of the sequence set forth in SEQ ID NO:21 , or a polynucleotide encoding the same; wherein the nucleotides at positions 31 to 40 and 57 to 63 (or nucleotides at a position equivalent thereto) of SEQ ID NO:30 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0049] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO:38 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NO:57 or encoded by a polynucleotide comprising or consisting of the sequence set forth in SEQ ID NO:24, or a polynucleotide encoding the same; wherein the nucleotides at positions 32 to 37 and 56 to 64 (or nucleotides at a position equivalent thereto) of SEQ ID NO:32 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0050] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NO:39 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NO:55 or encoded by a polynucleotide comprising or consisting of the sequence set forth in SEQ ID NO:22, or a polynucleotide encoding the same; wherein the nucleotides at positions 13 to 21 and 36 to 43 (or nucleotides at a position equivalent thereto) of SEQ ID NO:29 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0051] In some embodiments, the gene editing system may comprise:(i) a polypeptide, fragment, or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in SEQ ID NQ:40 or a polynucleotide encoding the same; and(ii) an RNA molecule comprising a sequence at least about 70% identical to the sequence set forth in SEQ ID NO:56 or encoded by a polynucleotide comprising or consisting of the sequence set forth in SEQ ID NO:23, or a polynucleotide encoding the same; wherein the nucleotides at positions 13 to 22 and 39 to 43 (or nucleotides at a position equivalent thereto) of SEQ ID NO:31 are modified to be complementary to a non-native target site so as to enable binding of the RNA molecule to the non-native target site.
[0052] In any embodiment of any aspect of the disclosure described herein, the system may comprise one or more expression vectors encoding one or more of the system components. The skilled person will appreciate that one or more components of the gene editing system may be provided by a vector expressing said one or more components. In some embodiments, the gene editing system will comprise a combination of vectors encoding the components of the gene editing system. For example, the gene editing system may comprise a first expression vector encoding the polypeptide, fragment or functional equivalent thereof of (i), and a second expression vector encoding the RNA molecule of (ii). In other words, the nucleic acid cargo is notrequired to encode the polypeptide, fragment or functional equivalent thereof of (i), or the RNA molecule of (ii), because either or both of these components can be provided separately.
[0053] In another embodiment, the gene editing system is encoded by a single expression vector.
[0054] In any embodiment, the system may comprise one or more expression vectors encoding one or more of the system components, wherein the one or more expression vectors comprise one or more regulatory elements. For example, the expression vector may comprise one or more of the following: promoter (e.g. T7, T5, Sp6, araBAD, trp, Ptac, PI, T3, lac, endogenous promoter from an IS, endogenous promoter from a cargo gene), enhancer, ribosome binding site, transcription termination sequence (eg T7 terminator, polyadenylation signals), a ribozyme sequence (such as a hepatitis delta virus(HDV) or HDV-like ribozyme sequence, hammerhead ribozyme).
[0055] Preferably, the expression vector encoding an RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell comprises a 3’ ribozyme sequence downstream of the RNA-encoding sequence, most preferably a HDV ribozyme sequence. Preferably, expression of a polynucleotide encoding an RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell is driven by a promoter (eg T7, T5, endogenous promoter) that promotes transcription of multiple copies of the molecule.
[0056] In preferred embodiments, the gene editing system does not require an additional enzyme, such as an integrase, to ensure stable integration of the nucleic acid cargo into the target site.
[0057] The nucleic acid cargo to be inserted comprises towards its 5’ end a nucleotide sequence derived from the left end (LE) of an insertion sequence, and towards its 3’ end a nucleotide sequence derived from the right end (RE) of an insertion sequence. The LE is not required to be adjacent to or part of a sequence encoding the polypeptide, fragment or functional equivalent thereof of (i), or the RNA molecule of (ii). The RE is not required to be adjacent to or part of a sequence encoding the polypeptide, fragment or functional equivalent thereof of (i), or the RNA molecule of (ii).
[0058] In any embodiment of any aspect of the disclosure, the left end (LE) and right end (RE) sequences may comprise or consist of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the LE and RE sequences present in an insertion sequence defined by any one of SEQ ID NOs: 1 -16; optionally any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, or 16. Exemplary LE and RE sequences are shown in Tables and 10.
[0059] In any embodiment of any aspect of the disclosure, the left end (LE) and right end (RE) sequences may comprise or consist of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the sequence defined by any one of SEQ ID NOs: 132-147.
[0060] In some embodiments, the nucleic acid cargo may comprise an inverted repeat (IR) sequence of about 3-10 nucleotides in length within or near to the LE and / or RE sequences, including within or near to both the LE and RE sequences, such as within 4, 5, 6, 7, 8, 9 or 10 nucleotides of the LE or within 2, 3, 4, 5, 6 7 or 8 nucleotides from the RE. In some embodiments, the nucleic acid cargo may comprise IR sequences within the LE and RE sequences.
[0061] IR sequences are present in IS 1111 and IS 110 insertion sequences; with longer sub-terminal inverted repeats being a distinctive feature of the IS 1111 family compared to the IS 110 family (see Figures 1A, 2A). IR sequences may be derived from a native IS 1111 or IS 110 insertion sequence, such as those present in the insertion sequences defined by any one of SEQ ID NOs:1 -16; optionally any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, or 16. The skilled person will appreciate how to identify an inverted repeat sequence and introduce an inverted repeat into a nucleic acid cargo as desired using cloning techniques.
[0062] In any embodiment of any aspect of the disclosure, the LE and / or RE sequence may correspond to a native LE sequence and a native RE sequence, respectively, derived from an insertion sequence eg ISEc21 . Representative native LE and RE sequences are underlined in Table 2 and shown in Table 10.
[0063] In other embodiments of any aspect of the disclosure, the length of the LE sequence and / or the RE sequence may be modified to modify transposition efficiency. The skilled person will appreciate suitable methods to (a) shorten or lengthen the LE and / or RE sequence within a polynucleotide; and (b) determine that transposition efficiency (such as using an assay described herein; see eg Examples 2, 7-9).
[0064] In any embodiment of any aspect of the disclosure described herein, the LE sequence is at least 3 nucleotides, at least 4 nucleotides, least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 33 nucleotides, at least 34 nucleotides, or at least 35 nucleotides in length. In some embodiments, the LE sequence is between at least 7 nucleotides and at least 16 nucleotides in length.
[0065] In any embodiment of any aspect of the disclosure described herein, the RE sequence is at least 3 nucleotides, at least 4 nucleotides, least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, at least 30 nucleotides, at least 33 nucleotides, at least 34 nucleotides, or at least 35 nucleotides in length. In some embodiments, the RE sequence is between at least 3 nucleotides and at least 16 nucleotides in length.
[0066] In any aspect of the disclosure, the nucleic acid cargo comprises at its 5’ end, 5’ of the LE sequence, and its 3’ end, 3’ of the RE sequence, a nucleotide sequence corresponding to the target site in the genome of the cell. The 5’ end sequence can bereferred to as the left flank (LF) or left target (LT) sequence. The 3’ end sequence can be referred to as the right flank (RF) or the right target (RT) sequence.
[0067] In any embodiment of any aspect of the disclosure described herein the left flank and right flank sequences of the nucleic acid cargo may comprise or consist of a nucleotide sequence that is at least 50%, at least 55%, at least 60%, at least 65%, least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the a target site nucleotide sequence set forth in any one of SEQ ID NOs: 58-131 ; optionally selected from any of SEQ ID NOs: 58, 66, 67, 76, 82, 112, 1 15, 127 and 129.
[0068] In preferred embodiments of any aspect of the disclosure, the nucleic acid cargo comprises, from 5’ to 3’, LF-LE-nucleotide sequence of interest-RE-RF, wherein the LE and RE enable formation of the minicircle but the LE and RE are not artificially joined or concatenated in either order (RE-LE or LE-RE). The nucleic acid cargo may also be represented as 5’-LF-LE-X-RE-RF-3’, where ‘X’ represents a nucleotide sequence or sequences of interest for insertion. Advantageously, this configuration, enabling formation of the minicircle, permits in cis or in trans insertion scenarios. Preferably, the formation of the minicircle ensures that the nucleic acid cargo is ultimately inserted in its true 5’ to 3’ orientation (ie LE-X-RE in the 5’ to 3’ direction), without inversion of the nucleic acid cargo. Preferably, when the minicircle is formed RE-LE are not directly adjacent. The polypeptide and RNA molecule may be expressed separately.
[0069] In other embodiments using an artificial minicircle, the nucleic acid cargo comprises, from 5’ to 3’, LE-nucleic acid cargo - RE, wherein the LE and RE are artificially joined / concatenated to form the minicircle. Joined / concatenated may include directly linked, or indirectly linked by a sequence of less than about 20 basepairs, less than about 15 basepairs, less than about 10 basepairs, less than about 5 basepairs, or between 1 -10 basepairs, such as 1 -5 basepairs. The artificially joined / concatenated LE and RE sequences may act as a promoter for expression of the polypeptide and RNA molecule, provided sequences encoding the polypeptide and RNA molecule are positioned downstream of the concatenated LE and RE sequences. In preferred embodiments, the LE and RE sequences are concatenated LE-RE. In other embodiments, they may be joined RE-LE.
[0070] In any embodiment of any aspect of the disclosure described herein, the nucleic acid cargo comprises at 5’ of the LE sequence a nucleotide sequence corresponding to at least 4, least 5, at least 6, at least 7, at least 8, or at least 9 contiguous nucleotides of the 5’ end of the target site sequence.
[0071] In any embodiment of any aspect of the disclosure described herein, the nucleic acid cargo comprises 3’ of the RE sequence, a nucleotide sequence corresponding to at least 4, at least 5, at least 6, at least 7, at least 8, or at least 9 contiguous nucleotides of the 3’ end of the target site sequence.
[0072] The sequences of the nucleic acid cargo that are 5’ to the LE sequence and 3’ of the RE sequence correspond to the target site. In some embodiments the target site and the polypeptide, fragment or functional equivalent thereof of the gene editing system are derived from the same insertion sequence. For example, if the polypeptide, fragment or equivalent thereof is derived from TnpEc11 , then the target site, and the sequences of the nucleic acid cargo that are 5’ to the LE sequence and 3’ of the RE sequence, are also derived from ISEc11 . In other words, the gene editing system comprises polypeptide, fragment or equivalent thereof is derived from an insertion sequence, wherein the target site is derived from the native target site of the same insertion sequence.
[0073] In alternative embodiments, the target site and the polypeptide, fragment or functional equivalent thereof of the gene editing system are each derived from a different insertion sequence. In other words, the target site is not derived from the native target sequence of the insertion sequence from which the polypeptide, fragment or functional equivalent thereof is derived. In these embodiments, the RNA molecule, or polynucleotide encoding the same, may be reprogrammed such that it specifically binds or is capable of binding a target site that is not its native target site. For example, an RNA molecule of the gene editing system derived from the insertion sequence ISEc11 may be modified in the region which binds to the target site, so as to specifically bind to or be capable of specifically binding to a non-native target site of an insertion sequence that is not ISEc11 . The skilled person will be able to determine which region of the RNA molecule binds to the target site (and therefore which region requires modification to enable binding to a different eg, non-native, target site). For example, Table 8 defines the regions of the RNA molecule which are complementary to the target site. If the RNAmolecule of the gene editing system is reprogrammed to specifically bind to or be capable of specifically binding to a non-native target site, the nucleic acid cargo is reprogrammed accordingly to comprise sequences corresponding to the non-native target site bound or capable of being bound by the RNA molecule of the gene editing system.
[0074] In any embodiment of any aspect described herein, the nucleic acid cargo may comprise one or more nucleotide sequences of interest for insertion into the target site, which are positioned between the LE and RE sequences. The sequence or sequences for integration / insertion into the target site may be endogenous or exogenous to the cell. Examples of a sequence to be inserted include DNA encoding a protein, peptide, or a non-coding RNA (e.g., a microRNA, long non-coding RNA, shRNA). Thus, the nucleotide sequence of interest may be operably linked to an appropriate control / regulatory sequence or sequences. Alternatively, the sequence or sequences to be integrated may provide a regulatory function (eg insertion of a promoter upstream of a gene sequence). The nucleic acid cargo may be provided within an expression vector or as an isolated, purified, synthetic or recombinant polynucleotide.
[0075] In any embodiment of any aspect described herein, the nucleic acid cargo comprises a single cargo gene of interest. In other embodiments, the nucleic acid cargo comprises a gene cassette containing multiple cargo genes (including operons and / or regulons).
[0076] In any embodiment of any aspect described herein, the nucleotide of interest may comprise a coding sequence encoding a protein, peptide, or RNA molecule. The nucleotide of interest may comprise a gene sequence to be expressed by the cell. The nucleotide of interest may comprise one or more regulatory elements for modulating gene expression. For example, the nucleotide of interest may comprise one or more of the following: promoter (e.g. T7, T5, Sp6, araBAD, trp, Ptac, PI, T3, lac, endogenous promoter from an IS, endogenous promoter from a cargo gene), enhancer, ribosome binding site, transcription termination sequence (eg T7 terminator, polyadenylation signals), a ribozyme sequence (such as a hepatitis delta virus(HDV) or HDV-like ribozyme sequence, hammerhead ribozyme).
[0077] In any embodiment of any aspect described herein, the nucleotide of interest may comprise a mutated, truncated, or dominant negative sequence for insertion intothe target site, for the purpose of modifying or disrupting expression of the native sequence within the cell. A mutated sequence may comprise nucleotide substitutions, deletions and / or insertions.
[0078] In any embodiment of any aspect described herein, the system comprises one or more expression vectors wherein expression of (a) a polynucleotide encoding the RNA molecule of the gene editing system, and / or expression of (b) a polynucleotide encoding the polypeptide, fragment or functional equivalent thereof of the gene editing system, and / or expression of (c) the nucleic acid cargo are operably linked.
[0079] In some embodiments, a single promoter may drive expression of the nucleotide of interest and / or expression of (a) a polynucleotide encoding the RNA molecule and / or expression of (b) a polynucleotide encoding the polypeptide, fragment or functional equivalent thereof of the gene editing system; embedded within one or more intron sequences (e.g., each in a different intron, two or more in at least one intron, or all in a single intron; each in a different operon, two or more in at least one operon, or all in a single operon). In other embodiments, the nucleotide of interest, and / or (a) a polynucleotide encoding the RNA molecule of the gene editing system and / or (b) a polynucleotide encoding the polypeptide, fragment or functional equivalent thereof of the gene editing system; may be operably linked and expressed from the same promoter.
[0080] In any embodiment of any aspect described herein, the nucleic acid cargo may comprise a tnp gene encoding the associated Tnp polypeptide or fragment thereof. The nucleic acid cargo may comprise a polynucleotide sequence encoding the RNA molecule. Inclusion of a tnp encoding sequence and a corresponding seekRNA encoding sequence in the nucleic acid cargo may enable the nucleic acid cargo to further transpose, and optionally self-replicate, once inserted into the genome of the cell. Alternatively, the nucleic acid cargo is inserted without including a tnp encoding sequence and / or a seekRNA encoding sequence, such that the insertion of the cargo into the genome is immobile (ie no longer able to transpose).
[0081] In any embodiment of any aspect described herein, the nucleic acid cargo may comprise one or more suitable markers. Examples of suitable markers include restriction sites, one more nucleic acid sequences encoding a detectable label and / or selectable marker. For example, the nucleic acid cargo may comprise a reportersequence to indicate successful insertion of the nucleic acid cargo into the genome. The selectable marker may be an antibiotic resistance gene, such as for chloramphenicol resistance (cat), kanamycin resistance, ampicillin resistance. The detectable label may be a fluorescent protein, such as GFP, eGFP, RFP, BFP, dsRed, mCherry. Such a marker may make it easy to screen for targeted integrations.
[0082] The nucleic acid cargo may be at least about 5 nucleotides in size. Advantageously, the nucleic acid cargo may range in size from small (eg basepairs, bp; less than 1 kilobase) to large nucleic acid cargo (eg multi-kilobase, kb). In some embodiments, the nucleic acid cargo is between 5 nucleotides (ie 5 bases) and about 3 kilobases in size. In some embodiments, nucleic acid cargo is at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, or at least 950 bases in size.
[0083] Preferably, the nucleic acid cargo is a large nucleic acid cargo (eg multikilobase). In some embodiments, the nucleic acid cargo is at least one kilobase (kb), at least 2kb, at least 2.5kb, at least 3kb, at least 3.5kb, at least 4kb, at least 4.5kb, at least 5kb, at least 5.5kb, at least 6kb, at least 6.5 kb, at least 7kb, at least 7.5kb, at least 8kb, at least 8.5kb, at least 9kb, at least 9.5kb, at least 10kb, at least 11 kb, at least 12kb, at least 13kb, at least 14kb, at least 15kb, at least 16kb, at least 17kb, at least 18kb, at least 19kb or at least 20kb in size. Most preferably, the nucleic acid cargo is at least 3kb in size.
[0084] In some embodiments, the system further comprises: iii) an exogenous nucleic acid cargo comprising one or more nucleotide sequences of interest to be inserted into the target site within the genome; wherein the exogenous nucleic acid cargo comprises at its 5’ and 3’ ends a nucleotide sequence corresponding to the target site in the genome of the cell; wherein the polypeptide, fragment or functional equivalent thereof of (i) and the RNA molecule of (ii) are capable of forming a gene editing complex for enabling the insertion of the exogenous nucleic acid cargo into the target site within the genome.
[0085] In some embodiments, the system comprises one or more expression vectors encoding the nucleic acid cargo.
[0086] In some embodiments, the system comprises one or more expression vectors encoding the polypeptide, fragment or functional equivalent thereof of (i); the RNA molecule of (ii); and the nucleic acid cargo of (iii); or any combination thereof. In some embodiments, the gene editing system may comprise a vector encoding the polypeptide, fragment or functional equivalent thereof of (i); an expression vector encoding the RNA molecule of (ii); and an expression vector encoding the nucleic acid cargo. In other embodiments, the polypeptide, fragment or functional equivalent thereof of (i); the RNA molecule of (ii); and the nucleic acid cargo of (iii) may be encoded by a single expression vector.
[0087] In some embodiments, the system does not comprise an exogenous nucleic acid cargo. For example, the cell for gene editing endogenously expresses a nucleic acid cargo for insertion, whereby the nucleic acid cargo is capable of being inserted into the genomic target site once the system is introduced into the cell.
[0088] In another aspect, the disclosure provides an isolated, synthetic and / or recombinant polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40.
[0089] In any embodiment of any aspect of the disclosure described herein, the polypeptide, fragment or functional equivalent thereof may be modified to increase stability and / or activity compared to the unmodified polypeptide, fragment or functional equivalent thereof. Examples of modifications include but are not limited to methylation, phosphorylation, amino acid substitution (eg selectively removing or reducing identified protease cleavage sites; random mutagenesis to identify more stable and / or highly expressed variants), multimerisation (including dimerization eg homodimerization). In some embodiments, the RNA molecule of the gene editing system or polynucleotide encoding the same, is provided in excess compared to the polypeptide, fragment or functional equivalent thereof of the gene editing system, to improve the stability of the polypeptide, fragment or functional equivalent thereof (via the negatively charged RNA molecule forming a complex with the relatively positively-charged polypeptide, fragment or functional equivalent thereof).
[0090] In any embodiment the polypeptide, fragment, or functional equivalent thereof comprises a fusion polypeptide wherein the fusion protein comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 33-40.
[0091] In any embodiment of any aspect of the disclosure described herein, the polypeptide, fragment, or functional equivalent thereof may further comprise a marker such as a fluorescent detectable label, for example, for detection of delivery into the cell of interest. The skilled person will appreciate suitable markers as used by those in the art.
[0092] In any embodiment of any aspect of the disclosure described herein, the polypeptide, fragment, or functional equivalent thereof may comprise one or more tags to improve expression, stability, solubility and / or purification of the polypeptide, fragment, or functional equivalent thereof. A tag may be added to the N-terminus and / or C-terminus of the polypeptide, fragment, or functional equivalent thereof. Examples of protein tags include maltose binding protein (MBP), STREPtag II, Myc, His-tag, S-tag, SUMO, FLAG-tag, GST, HA, HBH, TAP, V5, TRX. Optionally, the tag is a purification tag. The skilled person will appreciate suitable tags as used by those in the art.
[0093] In any embodiment, the polypeptide, fragment or functional equivalent thereof may comprise a fragment or peptide that has been identified to have transposase activity by any of the methods described herein.
[0094] In another aspect, the disclosure provides an isolated, synthetic or recombinant polynucleotide encoding an RNA that specifically binds or is capable of specifically binding to a target site within the genome of the cell, comprising or consisting of the nucleotide sequence set forth in any one of SEQ ID NOs: 17-24. The isolated, synthetic or recombinant polynucleotide encoding the RNA may comprise or consist of the nucleotide sequence set forth in any one of SEQ ID NOs: 1 -16; optionally any one of SEQ ID NOs: 1 , 3, 5, 7, 9, 1 1 , 13 or 15.
[0095] In another aspect, the disclosure provides an isolated, synthetic or recombinant RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell, encoded by a sequence comprising or consisting of the nucleotide sequence set forth in any one of SEQ ID NOs: 25-32. In some embodiments, the synthetic or recombinant RNA molecule is encoded by a sequence comprising orconsisting of the nucleotide sequence set forth in any one of SEQ ID NOs: 17-24. In some embodiments, the synthetic or recombinant RNA molecule comprises or consists of the nucleotide sequence set forth in any one of SEQ ID NOs: 50-57.
[0096] In any embodiment of any aspect of the disclosure described herein, the RNA molecule of the gene editing system may be between at least 50 bases in length and about 200 nucleotides in length, such as about 50 nucleotides, about 55 nucleotides, about 60 nucleotides, about 65 nucleotides, about 70 nucleotides, about 75 nucleotides, about 80 nucleotides, about 85 nucleotides, about 90 nucleotides, about 95 nucleotides, about 100 nucleotides, about 105 nucleotides, about 110 nucleotides, about 1 15 nucleotides, about 120 nucleotides, about 125 nucleotides, about 130 nucleotides, about 135 nucleotides, about 140 nucleotides, about 145 nucleotides, about 150 nucleotides, about 155 nucleotides, about 160 nucleotides, about 165 nucleotides , about 170 nucleotides, about 175 nucleotides, about 180 nucleotides, about 185 nucleotides, about 190 nucleotides, about 195 nucleotides, or about 200 nucleotides in length. In some embodiments, the RNA molecule of the gene editing system in between about 58 nucleotides and about 180 nucleotides in length. In preferred embodiments, the RNA molecule is less than about 200 nucleotides, less than about 190 nucleotides, less than about 180 nucleotides, less than about 170 nucleotides, less than about 160 nucleotides, less than about 150 nucleotides, or less than about 140 nucleotides in length. More preferably, the RNA is between about 58 and about 96 nucleotides in length, such as between 74 to 96 nucleotides in length.
[0097] In any embodiment of any aspect of the disclosure described herein, the RNA molecule of the gene editing system may be truncated or lengthened such that transposition efficiency is improved. Typically, all that is required is for the skilled person to i) identify a seek RNA molecule complementary to the nucleic acid sequence of the desired target site; and ii) add or remove nucleotides from the region of the RNA molecule that is complementary to the target site (the bottom and / or top strand) to influence unpairing / pairing in the RNA sequence secondary structure; and iii) determine if the modified seekRNA has increased or decreased insertion efficiency. Exemplary seekRNA sequences and their secondary structures are provided herein. Transposition efficiency may be improved by at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, or at least about 40%, or higher. The skilled person will appreciate appropriate methods to assesstransposition efficiency, including the assays described herein for example in Figures 6C-D, Examples 7-9. The RNA molecule may be the full-length native seekRNA, or a native shorter seekRNA (for example, the RNA molecules outlined in Table 8). The RNA molecule may be a shorter seekRNA corresponding to the most enriched peak sequence, as identified by RNA sequencing. In some embodiments, the RNA molecule that has been modified from a native full-length seekRNA or a native short / peak seekRNA, such that the RNA molecule has increased pairing in the region complementary to the target site. In some embodiments, the RNA may be shortened and / or comprise an unpaired overhang in the region complementary to the target site.
[0098] In any embodiment of any aspect of the disclosure described herein, the RNA molecule of the gene editing system may be chemically modified to increase stability and / or activity compared to the unmodified RNA molecule. Examples of RNA chemical modifications include, without limitation, incorporation of 2'-O-methyl (M), 2'-O-methyl 3'phosphorothioate (MS), 2'-O-methyl 3'thioPACE (MSP), inclusion of one or more pseudouridines or a tRNA scaffold sequence at one or more terminal nucleotides, or any other stabilising RNA moiety known to those in the art.
[0099] In any embodiment of any aspect of the disclosure described herein, the RNA molecule of the gene editing system may further comprise a marker such as a fluorescent detectable label (eg fluorescent light-up aptamer (FLAP), for example, for detection of delivery into the cell of interest.
[0100] In any aspect of the disclosure, the RNA of the gene editing system forms a double-stranded secondary structure comprising: a terminal hairpin (including a first double-stranded duplex structure comprising about 4-10 paired nucleotides that is the stem of the hairpin, and a loop comprising about 3-7 nucleotides); and at least one internal mismatched nucleotide (ie single unpaired basepair) and / or bulge (ie a series of unpaired basepairs) immediately downstream of the first double-stranded structure; and at least one additional double-stranded duplex structure downstream of the internal mismatched nucleotide and / or bulge. Preferably, at least some of the unpairednucleotides of the internal mismatches and / or bulges are complementary to one or more sequences of the target sequence.
[0101] The at least one additional double-stranded duplex structure may connect two mismatches / bulges.
[0102] The RNA molecule may comprise a 5’ and / or 3’ overhang. The RNA molecule may comprise a 5’ overhang, optionally of 9-18 nucleotides in length, such as 9-12 nucleotides in length. The RNA molecule may comprise a 3’ overhang, optionally of 1 -7 nucleotides in length, such as 2-4 nucleotides in length.
[0103] In certain embodiments, the RNA molecule may not comprise an overhang.
[0104] Examples of the secondary structures of the RNA molecules for use according to the disclosure are provided in Figures 3e, 3g, 3i, 4, 6, and 8-12 herein; and Table 8. From these secondary structures, and the information provided in these figures and table, the skilled person will be able to readily determine how to suitably reprogramme the RNA molecules of the gene editing system for binding to alternative target sites. The skilled person is provided suitable methods herein to determine a successfully reprogrammed RNA molecule. The skilled person will further appreciate, from the information provided herein, that when reprogramming the RNA molecules, the sequence complementary to the target site should not form a tightly-bound doublestranded duplex (ie should be at least partially unpaired from its corresponding bottom strand), such that the region can interact with the target site.
[0105] In a further aspect, the disclosure provides an expression vector comprising polynucleotide sequence or sequences encoding one or more components of the systems as described herein.
[0106] In any embodiment of any aspect of the disclosure, the expression vector may be a viral or non-viral vector. Non-viral expression vectors include but are not limited to mRNA vectors, linear DNA (eg doggybone DNA) plasmids (such as pSV, pUC1 , pBluescript, pCMV, pCDF Duet-1 plasmids), phagemids, phage vectors such as M13 phage vector, bacmids, cosmids, and bacteriophage (eg bacteriophage lambda). In preferred embodiments, the expression vector is a plasmid. In certain embodiments, the plasmid may be conjugative or non-conjugative (eg for use in difficult to transform bacteria such as Streptomyces).
[0107] In a further aspect, the disclosure provides a non-naturally occurring or engineered composition comprising a polypeptide or functional fragment thereof as described herein; an RNA molecule of the gene editing system as described herein; a polynucleotide as described herein; and / or a vector as described herein. In some embodiments, the composition may further comprise a delivery system such as non- viral vectors (eg nanoparticles, lipid nanoparticles, exosomes, extracellular vesicles, microvesicles, liposomes, polyplexes), to assist delivery of the gene editing system or one or more system components to a genome of a prokaryotic cell.
[0108] In a further aspect, the disclosure provides a method for genome editing, comprising providing to a prokaryotic cell a gene editing system or one or more system components as described herein.
[0109] In a further aspect, the disclosure provides a method for inserting a nucleic acid cargo into a genome of a prokaryotic cell, comprising providing to the cell a gene editing system or one or more system components as described herein.
[0110] In a further aspect, the disclosure provides a method for inserting a nucleic acid cargo into a genome of a prokaryotic cell, comprising introducing into the cell an expression vector encoding one or more components of the systems as described herein.
[0111] In any aspect, the method may further comprise determining whether the nucleic acid cargo, and / or nucleotide of interest, has been inserted into the target site. Integration detection may comprise PCR amplification of the nucleic acid cargo and / or nucleotide of interest, restriction digest, DNA sequencing, and / or detection of a suitable marker wherein the nucleic acid cargo comprises a marker (eg detectable label, antibiotic resistance gene).
[0112] The methods of the disclosure may be in vitro or ex vivo. In some embodiments, the method comprises sampling a cell or population of prokaryotic cells from an environment (eg mammalian tissue; faecal sample; environmental sample), and modifying the cell or cells to have a nucleic acid cargo insertion. In some embodiments, the methods of the disclosure may be used in a method of generating an in vitro or ex vivo model of a disease or disorder caused by or suspected to be caused by a prokaryotic cell, such as an infectious disease. In some embodiments, the methods ofthe disclosure may be used in a method of identifying the function of a sequence such as a gene. For example, the methods of the disclosure may be used in a genetic complementation assay to assess a mutation or mutations that may contribute to a particular phenotype (eg antibiotic resistance, increased pathogenicity).
[0113] In some embodiments, the system may be provided to the cell in a single composition comprising the gene editing system and the nucleic acid cargo. For example, the system may be delivered in the form of an expression vector that comprises one or more polynucleotides encoding the polypeptide, fragment or functional equivalent thereof of (i); encoding the RNA of (ii); and encoding the nucleic acid cargo.
[0114] In alternative embodiments, the system may be delivered to the cell as one or more separate components. For example, the cell may be first provided with the polypeptide, fragment or functional equivalent thereof of (i), or polynucleotide encoding said polypeptide, fragment or functional equivalent thereof; and secondly provided the RNA molecule of (ii), or polynucleotide encoding said RNA molecule.
[0115] In some embodiments, the system or one or more components of the system, may be provided to the prokaryotic cell utilising a delivery system such as non-viral vectors (eg nanoparticles, lipid nanoparticles, exosomes, extracellular vesicles, microvesicles, liposomes, polyplexes), electroporation, or gene gun / biolistics. The skilled person will be familiar with appropriate delivery systems available in the art.
[0116] In any embodiment of any aspect of the disclosure, the prokaryotic cell may be a bacterial cell or archaea. The skilled person will be familiar with prokaryotic cell types that are suitable for gene editing applications. A bacterial cell may be a gram positive or gram negative bacterium. The bacterial cell may be from the genus: Enterococcus, Enterobacter, Legionella, Burkholderia, Yersinia, Pediococcus, Citrobacter, Rickettsia, Wolbachia, Nocardia, Mycoplasma, Bacillus, Lactobacillus, Bifidobacterium, Clostridioides, Clostridium, Corynebacterium, Streptococcus, Streptomyces, Staphylococcus, Camplyobacter, Eubacterium, Acinetobacter, Pseudomonas, Vibrio, Leptospira, Shigella, Clostridioides, Salmonella, Haemophilus, Helicobacter, Mycobacterium, Listeria.
[0117] Bacterial cells include but are not limited to: Escherichia coli, Acinetobacter baumanii, Klebsiella pneumonia, Pseudomonas aeruginosa, Staphylococcus aureus (including MRSA), Lactobacillus acidophilus, Vibrio cholerae, Salmonella, Neisseria gonorrhoeae, Mycobacterium tuberculosis, Bacillus cereus, Bacillus subtilis, Chlamydia trachomatis, Enterococcus faecalis, Campylobacter jejuni, Group A Streptococcus, Streptococcus pneumoniae, Streptococcus pyogenes, Streptococcus dysgalactiaem Corynebacterium diptheriae, Klebsiella spp, Citrobacter spp, Shigella flexneri, Shigella sonnei, Clostridium botulinum, Clostridium tetani, Clostridioides difficile, Clostridium perfringens, Clostridium sordellii, Neisseria meningitides, Haemophilus influenza, Group B Streptococcus including Streptococcus agalactiae, Helicobacter pylori, Listeria monocytogenes, Burkholderia cenocepacia,l Legionella pneumophilia, Staphylococcus sp,, Yersinia pestis, and Porphyromonas gingivalis.
[0118] In some embodiments, the prokaryotic cell may be a pathogen.
[0119] The nucleic acid cargo introduced to the cell by the present disclosure may be such that the cell and progeny of the cell are engineered for improved production of biologic products such as an antibody, antibiotic, starch, alcohol or other desired cellular output. The nucleic acid cargo inserted into the genome of the cell by the present disclosure may be such that the cell and progeny of the cell include an alteration that changes the biologic product produced.
[0120] In any embodiment of any aspect of the disclosure, the genome of the cell is or has been modified to include the target site sequence or additional copies of the target site sequence.
[0121] In any embodiment, the cell may comprise an identified genetic mutation or mutations that alters the cell’s function. For example, gene mutations can result in production of improper gene products in improper amounts resulting in dysfunction; alternatively gene mutations can result in a gain of function (eg antibiotic resistance; increased pathogenicity). In some embodiments, the nucleic acid cargo to be inserted provides a gene, genes or non-coding sequence / s that correct, rescue, compensate for, replace, or reduce the effects of a mutation in a cell. In some embodiments, the nucleic acid cargo to be inserted provides additional copies of a gene or genes, to compensate or rescue a gene mutation or mutations in the cell. In other embodiments, the nucleicacid cargo may provide a gene, genes or non-coding sequence / s which modulate a molecular signalling pathway associated with a gene mutation.
[0122] The skilled person will appreciate genetic mutations that would be appropriate for such gene editing applications, such as generating in vitro models for infectious disease, antibiotic resistance. The skilled person would be able to identify suitable mutated genes from publically available online databases (eg NCBI, ENSEMBL), and it is within their remit to design a suitable nucleic acid cargo using the gene editing system described herein, that would edit the genetic mutation upon insertion of the nucleic acid cargo into the genome of a cell.
[0123] In a further aspect of the disclosure, the disclosure provides an in vitro method of modelling an infectious disease, comprising administering to the cell a gene editing system described herein to correct, compensate for, replace, or reduce the effects of a gene mutation in the cell.
[0124] In some embodiments, the infectious disease is a bacterial infection selected from: Escherichia coli, Acinetobacter baumanii, Klebsiella pneumonia, Pseudomonas aeruginosa, Staphylococcus aureus (including MRSA), Vibrio cholerae, Salmonella, Neisseria gonorrhoeae, Mycobacterium tuberculosis, Bacillus cereus, Bacillus subtilis, Chlamydia trachomatis, Enterococcus faecalis, Campylobacter jejuni, Group A Streptococcus, Streptococcus pneumoniae, Streptococcus pyogenes, Streptococcus dysgalactiaem Corynebacterium diptheriae, Klebsiella spp, Citrobacter spp, Shigella flexneri, Shigella sonnei, Clostridium botulinum, Clostridium tetani, Clostridioides difficile, Clostridium perfringens, Clostridium sordellii, Neisseria meningitides, Haemophilus influenza, Group B Streptococcus including Streptococcus agalactiae, Helicobacter pylori, Listeria monocytogenes, Burkholderia cenocepacia, I Legionella pneumophilia, Staphylococcus sp,, Yersinia pestis, and Porphyromonas gingivalis.
[0125] In a further aspect of the disclosure, the disclosure provides a kit comprising the gene editing system described herein, or components thereof described herein. In some embodiments, the kit comprises:(i) a polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40, or a polynucleotide encoding said polypeptide, fragment or functional equivalent thereof; andii) a RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome, or a polynucleotide sequence encoding said seek RNA molecule.
[0126] The kit may further comprise a nucleic acid cargo for insertion into a target site in the genome of a cell. The kit may further comprise oligonucleotide primers for enabling confirmation or detection of a successful insertion of the cargo, eg oligonucleotide primers for amplifying and / or sequencing a LE and / or RE end sequence. The kit may further comprise oligonucleotide primers for amplifying a LE and / or RE end sequence, in order to generate a nucleic acid cargo suitable for use with a gene editing system described herein. The kit may further comprise oligonucleotide primers for amplifying a target sequence or sequences or part thereof, in order to generate a nucleic acid cargo suitable for use with a gene editing system described herein. In some embodiments, the kit comprises a vector or combination of vectors as described herein.
[0127] In a further aspect, the disclosure provides an engineered prokaryotic cell comprising a gene editing system or components thereof described herein; or obtained by the methods described herein. The engineered cell may be isolated.
[0128] In some embodiments, the engineered cell may be a host cell for recombinant production of a polypeptide or peptide (such as E.coli, B. subtilis).
[0129] In some embodiments, the engineered cell may be used as a probiotic (such as Lactobacillus and Bifidobacterium species
[0130] In some embodiments, the engineered cell may be used for fermentation processes, such as those used to produce: a recombinant polypeptide or peptide; fermented food or beverage product; a biofuel; hydrogen. Fermented food or beverage products include for example cultured milk, yoghurt, cheese, alcoholic drinks (eg wine, beer, cider), tempeh, miso, kimchi, sauerkraut, kombucha. Fermented biofuels include ethanol, butanol and propanol bioalcohols.
[0131] The disclosure further provides for a population of engineered cells comprising an engineered cell described herein, and for progeny of an engineered cell described herein.
[0132] As used herein, except where the context requires otherwise, the term "comprise" and variations of the term, such as "comprising", "comprises" and "comprised", are not intended to exclude further additives, components, integers or steps.
[0133] Further aspects of the present disclosure and further embodiments of the aspects described in the preceding paragraphs will become apparent from the following description, given by way of example and with reference to the accompanying drawings.
[0134] It will be understood that the disclosure disclosed and defined in this specification extends to all alternative combinations of two or more of the individual features mentioned or evident from the text or drawings. All of these different combinations constitute various alternative aspects of the disclosure.Brief description of the drawings
[0135] Figure 1 - IS1111 and IS110 are two distinct families of insertion sequences. A. Diagram of IS 1111 and IS 110 family members organization. IS 1111 contains a non-coding region (NCR) at the 3' end, while in IS 110 family members the NCR is at the 5' end. ISs are nestled between left and right end (LE, RE) sequences. Subterminal inverted repeat (sTIR) at the LE and RE are prevalent on IS 1111 ISs. B-C. Domain structural representations of ISEc11 and ISEc21 as examples of IS 1111 and IS 110 family members, respectively. AlphaFold 2 (AF2) model prediction of the transposase monomers consisting of a RuvC, a coiled-coil domain (CC), a variable region (V) and an un characterized C-terminal domain (CTD) (pfam: PF02371 ). D.Simplified phylogenetic tree analysis of transposase sequences from IS 1111 and IS 110 family members, rooted by midpoint. The scale bar represents a 20% divergence in protein sequence. Arrow points to ISs studied in this work.
[0136] Figure 2 - ISEc11 of IS1111 family, carries a ncRNA required for transposition. A. Diagram of pDonor plasmids encoding ISEc11 and ISEc11 ANCR, lacking the non-coding region (nucleotides 1 170-1374). ISEc11 forms an intermediary minicircle by connecting the left and right ends (LE and RE) and forming a promoter while ISEc11 ANCR is not able to form minicircle. Minicircle formation is confirmed by PCR employing outward facing primers R and F and Sanger sequencing of the PCR products. The IS ends is indicated by an arrow, whereas the sTIRs within the RE andLE are denoted by IR. B. Experimental workflow of transposition assay. Plasmids encoding ISEc11 (pDonor) was co-transformed with plasmids encoding the target sequence of ISEc11 (pTarget) into E. coli. Transposition products result in the formation of a plnsert plasmid with ISEd 1 inserted within its target site. C. In vitro transposition can be detected using primers sets (a and b) complementary to the target plasmid and the transposase gene on each end of the IS and target site insertion. Full length ISEd 1 , but not ISEd 1 ANCR, can insert into its target, ISEd 1 ANCR can be rescued upon the co-transformation of E. coli with a plasmid containing the NCR (pncRNA). Sequencing of the PCR products confirms the correct target site insertion point. D. Removal of either left flank (ALF) or right flank (ARF) of the split target site around the ISEd 1 in the pDonor abolishes transposition.
[0137] Figure 3 - Characterisation of seeker RNA bound to of IS1111 family members. A. SDS-PAGE of purified TnpEd 1 alone or in the presences of the ISEd 1 . TnpEd 1 yields were greatly enhanced by the presence of the ncRNA encoded within the IS sequence. B. SYBR GOLD-stained denaturing polyacrylamide gel electrophoresis of purified TnpEd 1 alone or in the presence of the corresponding ISEd 1 after digestion with DNase or RNase. C. Schematic of the RNA purification and sequencing. D. Small RNA-seq reads mapped to the ISEd 1 sequence and the seeker RNA is part of the ncRNA with a main peak containing a shorter seekRNA. E-F. Same as described in (d) applied to ISPal 1 , ISKpn4 respectively. G. Folded structure prediction of the seekRNA contains the target site of ISEd 1 with complementary sequence to the top (F) and bottom strand (R). H-l. Same as described in (g) applied to ISPal 1 , ISKpn4 respectively.
[0138] Figure 4 - ISEc11 can be engineered to carry cargo. A. An example of natural reprogramming of ISEd 1 with a homolog ISXne4 (72% sequence identity of transposase). Schematic of LE and RE including the sTIRs sequences aligned. Target sites of ISEd 1 and ISXne4, dashed line indicates the insertion point. F and R correspond to the complementary sequences matched to the seekRNA of ISEd 1 and ISXne4. Folded structure prediction of ISEd 1 and ISXne4 seekRNA show a conserved sequence between the variants along with the complementary sequences to their target site spaced by a conserved hairpin. B. Schematic of transposition of ISEd 1 alone or with a cargo gene (catA 1 - 750 bp) inserted upstream of the tnp after 50 bp from the start of the IS or downstream of the tnp 46 bp from the end of IS. Mini-circle formationand transposition into the pTarget was detected for all 3 pDonor plasmids. C. Cargo (catA1) movement with all elements of the ISEc1 1 separated into different plasmids, pTnp (containing the tnpEd 1 under the control of the T7 promoter), pSeekRNA containing the full seekRNA (under the control of a T7 promoter followed by a HDV ribozyme) and a pCargo containing the cat flanked by 84 bp of LE and 71 bp of RE of the ISEc11 along with 4 and 7 bp from the flanking target at each ends (LF and RF). Mini-circle formation can be observed once the three plasmids are co -transformed into E. coli, while transposition into a pTarget plasmid can be detected only upon cotransformation of all 4 plasmids into E. coli. PCR products were sequenced to confirm the insertion point and direction. D. Schematic of the mCherry reporter assay. The pDonor plasmid contains a T7 promoter preceding tnpEd 1 and the natural or re-programmed seekRNA followed by a HDV ribozyme with a strong T7 terminator at the 3’end. The mini IS contains mCherry without promoter surrounded by the LE and left flanking (LF) sequence and the RE and right flanking sequence (RF). pTarget includes a target (LF and RF abutted) preceded by a T7 promoter (bent arrow). In each assay, the LF and RF in the donor and target plasmids are identical and appropriate to the matches in the seekRNA. E. Transposition efficiency of the mCherry mini IS by TnpEd 1 with the natural (ISEd 1 ) or reprogrammed seekRNA (M1 and M2) or the peak ISEd 1 seekRNA. The target sequences used are shown with vertical dashed line marking the insertion point. Black dots indicate triplicate experimental determinations with standard deviation shown as bar. Transposition frequencies are on the right. F. Sequences of the ISEd 1 target and modified targets (M1 and M2) with the base-pairing sequences present in the seekRNA below the top strand sequence and above the bottom strand sequence. Arrows point to the 3’ of the DNA target site in both strands T (top) and B (bottom). Dashed line indicates the insertion point of the target site.
[0139] Figure 5 - ISEc21 of IS110 family contains a ncRNA required for transposition. A. Diagram of pDonor plasmids encoding ISEc21 and ISEc21 ANCR, lacking the non-coding region (bases 20-150). ISEc21 forms an intermediary mini-circle by connecting the left and right ends (LE and RE) and forming a promoter while ISEc21 ANCR is not able to form minicircle. Minicircle formation is confirmed by PCR employing outward facing primers R and F and Sanger sequencing of the PCR products. The IS ends is indicated by an arrow, whereas the RE and LE are denoted. B. Experimental workflow of in vitro transposition assay. Plasmids encoding ISEc21 (pDonor) were cotransformed with plasmids encoding the target sequence of ISEc21 (pTarget) into E.coli. Transposition products results in the formation of a plnsert plasmid with ISEc11 inserted within its target site. C. In vitro transposition can be detected using primer sets complementary to the target plasmid and the transposase gene on each end of the IS and target site insertion (a and b). Full length ISEc21 , but not ISEc21 ANCR, can insert into its target, ISEc21 ANCR can be rescued upon the co-transformation of E. coli with a plasmid containing the NCR (pncRNA). D. Sequencing of the PCR products confirm the correct target insertion point. E. Removal of left flank and right flank (ALFRF) of the split target site around the ISEc21 in the pDonor abolishes transposition, while the split target can be shortened to 5 bp at each side of the flanking sides and transposition is retained.
[0140] Figure 6 - Characterisation and reprogramming of the seekRNA fromISEc21. A. TnpEc21 was purified after being expressed alone or in the presence of the ISEc21 . Small RNA-seq results from the purified RNP complexes. Alignment of the small RNA-seq reads to the ISEc21 plasmid sequence identify the seeker RNA boundaries within the ISEc21 on the 5’ end of the transposase. Two distinct peaks were observed with the main peak (peak 2) seekRNA containing the complementary sequence to ISEc21 target site (R). Insertion point into the target sequence is indicated by the dashed line. Folded structure prediction of peak 2 (74 nt) seekRNA from ISEc21 with complementary sequence to its target site is indicated. B. Schematic of a mCherry reporter assay showing a pDonor plasmid containing the tnpEc21 and seekRNA under the control of a T7 promoter followed by a HDV ribozyme in the 3’end, a strong T7 terminator and the mCherry reporter gene without promoter is inserted between the left flanking sequence (LF) and LE upstream of the cargo gene and the RE and right flanking sequence (RF) downstream of the cargo. C. Upon co-transformation of the pDonor with pTarget into E. co / / transposition was observed with the full ncRNA, the full length seekRNA and the shortened seekRNA (58 nt) via PCR and Sanger sequencing. D. Reprogramming. New sequences are in the target-complementary regions of the shortened seekRNA, in the target and in the LF and RF of the mini IS. Transposition was detected by PCR and sequencing of the plnsert boundaries and detected by mCherry fluorescence.
[0141] Figure 7 - SDS-PAGE and nuclease digest gel analysis. Coomassie stained SDS-PAGE gel shows size and purity of fusion transposases containing an N-terminal 6X-His-MBP tagged, and C-terminal StrepTag(ll). A. TnpKpn4 (85.7 kDa); B. TnpPa11(83.9 kDa); C. TnpEc21 (86 kDa) purified alone or in the presence of their corresponding ISs (see methods and Figure 2). Post-purification, protein samples were digested with DNAse or RNAse and resolved in a 7 M Urea polyacrylamide denaturing gel and stained with SYBR GOLD.
[0142] Figure 8 - IS1111 small RNA-seq and seekRNA folded structure similarities and differences compared with their target site. A. Small RNA-seq reads of the TnpKpn4 expressed with the ISKpn4 complex were aligned to the full length ISKpn4 sequence. The full seekRNA is 151 nt in length and peak seekRNA is 82 nt. B. Small RNA-seq reads of the TnpPst6 expressed with the ISPst6 complex were aligned to the full length ISPst6 sequence. The full seekRNA is 148 nt in length and peak seekRNA is 86 nt. C-D. The transposase of ISKpn4 and ISPst6 are 86% identical. Transposase from ISPa25 and ISKpn4 are 46% identical. Alignment of the NCR of ISPa25 and ISKpn4 indicates a high level of identity (B; bold and black). The mapped seekRNA to the target site is shown (bold and red).
[0143] Figure 9 - ISEc11 and ISPa11 seekRNA boundaries, folded structure and target mapping and reprograming of IS1111 ISs. A. Small RNA-seq reads of the TnpEc11 expressed with ISEc11 were aligned to the full length ISEc11 sequence. The full seekRNA is 154 nt in length and spans from 74-227 bp after the stop codon in ISEc11 . Folded structure prediction shows two hairpin structures with internal loops. The peak seekRNA is 82 nt in length and spans from 82-164 bp after the stop codon. B. Small RNA-seq reads of the TnpPal 1 expressed with ISPal 1 were aligned to the full length ISPal 1 sequence. The full seekRNA is 172 nt in length and spans 39-210 bp after the stop codon in ISPal 1 . Folded structure prediction shows two hairpin structures with internal loops within each. The peak seekRNA is 96 nt in length and spans from 86- 161 bp after the transposase stop codon. The Weblogo shows the consensus target sequence which is complementary mapped onto the peak seekRNA. The cleavage and insertion point is marked with a dashed line on both strands. C-D. ISEc11 seekRNA modifications for reprogramming into new target sites. Folded structure model of the seekRNA and peak seekRNA of ISEc1 1 with the bases involved in base pairing with the target site sequence shown in green (top -T) strand (enclosed in a dashed line box) and blue (bottom -B) strand. The DNA sequence of the target site is shown for both DNA strands with the arrow pointing towards the 3’ of the corresponding strand and dashed lines indicates the insertion point of the IS. The sequence of the donor LF- LE and RE-RF of ISEc11 can be recognized by the seekRNA with base pairing nucleotides shown in green for top strand (enclosed in a dashed line box) and blue for the bottom strand. E-F. Folded structure model of the reprogrammed seekRNA of ISEc11 with the bases involved in base pairing with the modified target site M1 (E) and M2 (F) showing the top (T) strand (enclosed in a dashed line box) and the bottom (B) strand. The corresponding sequence of the new target site M1 and M22 are shown for both DNA strands with the arrow pointing towards the 3’ of the corresponding DNA strand and dashed lines indicating the insertion point of the IS. The sequence of the donor LF-LE and RE-RF of ISEc11 can be recognized by the seekRNA with base pairing nucleotides shown in green for top strand (enclosed in a dashed line box) and blue for the bottom strand.
[0144] Figure 10 - Sequence alignment the NCR region of ISEc11 and ISXne4 shows a natural reprogramming. Transposases from ISEd 1 and ISXne4 are 71 .6% sequence identical. A-C. Alignment of the NCR of ISEd 1 and ISXne4 starting from the stop codon of the tnp is shown. D. Experimentally determined seekRNA for ISEd 1 , and predicted seekRNA for ISXne4 are highlighted as bold and black from 82-164. The seekRNA fold is similar, however, each IS targets a different site. Identical bases are highlighted and enclosed in a dashed line box.
[0145] Figure 11 - ISEc21 seekRNA folded structure and target site mapping and a naturally occurring reprogramming ISMch6. A. Small RNA-seq reads of the TnpEc21 RNP complex aligned to the full-length ISEc21 sequence. The full seekRNA folded structure prediction is shown below with folds for the two shorter RNAs below. The peak 1 seekRNA spans from 8-88 from the start of the ISEc21 . The peak 2 seekRNA, is 74 nt in length and spans from 89-163 bp. Consensus target sequence mapping shows that the bottom (B, blue) and top (T, green; enclosed in a dashed line box) strands maps to the peak 2 seekRNA. B. Comparison of the seekRNAs for ISEc21 and ISMch6. Transposases from ISEc21 and ISMch6 are 84% identical. A predicted ISMch6 seekRNA is derived from the alignment of the ISEc21 and ISMch6 NCRs from the start of ISs to the start codon of the transposases. The region of ISEc21 seekRNA is bold and the target matching regions are shaded grey with differences in ISMch6 underlined or in black.
[0146] Figure 12 - ISEc21 seekRNA modifications for reprogramming into new target site. The peak 2 seekRNA 74 nt (89-163) of ISEc21 folded structure is shownwith the native target site complements shown in blue (B) or green (T; enclosed in a dashed line box) on the seekRNA. Reprograming of the seekRNA into M1 and M2 are shown on the bottom strand (B) on a shorter 52 nt seekRNA. Addition of 15 nt to the 5’end of the short seekRNA M1 , resulted in a total of 16 nt base pairing as indicated. M2 contains 4 extra nucleotides added to the 5’end of the seekRNA bringing the total base pairing with the target to 21 nt. The new targets are shown below.
[0147] Figure 13 - Phylogenetic tree of transposases sequences from members of IS110 and IS1111 ISs. Phylogenetic tree shown as a circle was generated on MEGA using a maximum likelihood neighbour-joining tree default settings and rooted using the midpoint. The MSA used to generate the tree is comprised of 196 sequences curated from the ISFinder to include a representative of the cluster with >70 sequence identity and include Piv sequence was aligned using Clustal Omega. First family members ( IS110 and IS 111 / ), Piv and the ISs used in this study are highlighted.
[0148] Figure 14 - Transposition efficiency of ISEc11 seekRNA and reprogramming using mCherry. A. Plots of cells containing a target plasmid (pTarget) and a pDonor and using ISEc11 seekRNA full and mCherry as a cargo (listed above the plots) were gated for all cells (FSC-A and SSC-A), single cells (FSC-A and FSC-H) and mCherry expression shown in each plot. The transposition efficiency of the ISEc11 as the percentage of the total single cells expressing mCherry. B. Plots of cells co-transformed with reprogramed pDonor and pTarget plasmids containing a the reprogramed seekRNA and the corresponding target and donor sequence for M1 or M2 and the peak seekRNA with native target of ISEc11 gated for all cells, single cells and mCherry expression shown in each plot. The transposition efficiency of the reprogramed ISEc11 and peak seekRNA.
[0149] Figure 15 - Transposition efficiency of ISEc21 following truncation of LE and RE. Plots of cells containing a target plasmid (pTarget) and a pDonor and using ISEc21 seekRNA full and mCherry as a cargo (listed above the plots) were gated for all cells (FSC-A and SSC-A), single cells (FSC-A and FSC-H) and mCherry expression shown in each plot. The transposition efficiency of the ISEc21 measured as the percentage of the total single cells expressing mCherry. Shortening the LE and RE to contain 16 bases (shown in B) increases the transposition efficiency from 1 1% (original full-length LE and RE, shown in A) to 30%. Shortening the LE and RE to 7 bases hasan efficiency of 17% (shown in C) while removal of the LE and RE sequences completely resulted in only 2% transposition efficiency (shown in D). This indicates that the optimal length of the LE and RE sequences for ISEc21 is between 7-16 bases.
[0150] Figure 16 - Transposition of ISEc11 following truncation of LE and RE.Truncation of the LE and RE sequences of ISEc11 was investigated using an mCherry transposition assay as described herein. A. shows the PCR products indicating minicircle formation B. shows the PCR products indicating transposition into target pSFA59; uninduced (no IPTG) or induced (IPTG added to induce transposase and seekRNA expression). pSFA107 contains the native full-length LE and RE sequences and serves as a positive control. Removal of the LE and RE sequences entirely (pSFA152) abolished minicircle formation and transposition. The combination of removing the inverted repeat (IR) sequences, truncating the LE to 7 bases, and truncating the RE to 3 bases (pSFA151 ) also abolished minicircle formation and transposition. This indicates that a minimum of the IR plus a 7-base LE and a 3-base RE are required for ISEc11 transposition.Sequence information
[0151] Table 1 : Insertion sequences
[0152] Table 2: DNA sequences of complete insertion sequences.
[0153] Table 3: Non-coding regions (NCRs) nucleotide sequences and the ncRNA- encoding regions of the NCR as determined by small RNA-Seq.
[0154] Table 4: Subset of total NCR sequence that encodes peak seek RNA sequences as determined by small RNA-Seq
[0155] Table 5: Exemplary IS 1111 transposase amino acid sequences
[0156] Table 6: Exemplary IS 110 transposase amino acid sequences
[0157] Table 7: Transposase encoding nucleotide sequences
[0158] Table 8: Exemplary seekRNA sequences
[0159] Table 9: Exemplary target sequences.
[0160] Table 10: Exemplary LE and RE sequences.Detailed description of the embodiments
[0161] The disclosure provides for gene editing systems derived from \S 1111 and IS 110 insertion sequences, specifically derived from the insertion sequences designated ISEc11 , ISKpn4, ISPst6, ISPal 1 and ISEc21 , and components for said gene editing systems. The disclosure further provides for methods of use thereof, compositions, cells, and kits comprising said gene editing systems.IS1 111 and IS110 families
[0162] Insertion Sequences (IS) in the IS 1111 and IS110 families have been identified to encode an unusual transposase (Tnp) type, referred to as DEDD transposases due to the presence of an N-terminal RuvC-like (also referred to as RuvC fold) catalytic domain. Each IS1111 and IS110 family member, or group of related family members, exhibit specificity for a different target site, where they are inserted in only one orientation. This characteristic indicates a flexible target site selection mechanism, which could be effectively modified for gene editing applications that require insertion of sequences into specific target sites.
[0163] However, to date, the IS 1111 and IS110 families have been insufficiently characterised for programmable genetic insertion. Without understanding the mechanisms behind IS 1111 and IS110 target site selection and sequence insertion, the ISs have not been able to be readily adapted into effective gene editing systems for specific insertion of genetic material into genomic target site.
[0164] A comprehensive and detailed analysis of over 50 IS identified as related to IS 1111 conducted 20 years ago confirmed a proposal (Lauf et al. (1999) Identification and characterisation of IS 1383, a new insertion sequence isolated from Pseudomonas putida strain H. FEMS Microbiol. Lett., 170: 407-412) that there were at least two families headed by IS 1111 (the founder of this family) and IS110 (Partridge and Hall (2003) The IS1111 family members IS4321 and IS5075 have sub-terminal inverted repeats and target the terminal inverted repeats of Tn21 family transposons. J.Bacteriol., 185(21 ):6371 -6384). A similar comprehensive analysis of the IS110 group is not available. Further information on these families can be found in the online resource TnPedia (https: / / tncentral.ncc.unesp.br / TnPedia / index.php) and in the ISFinder online database (https: / / isfinder.biotoul.fr / ; albeit where IS 1111 and IS110 family members are currently group together under “ IS 1 10 family”).
[0165] IS 7111 family member insertion sequences recovered from available sequences can be distinguished from IS110 family members by the presence of subterminal inverted repeats (here designated sTIR), usually 11 -13 bp in length with a perfect match. The IS 1111 family transposases align well with the transposases IS110 and its relatives in the N-terminal catalytic (DEDD, RuvC) domain, but less well at the C- terminus. The presence of a non-coding region (NCR) of significant length downstream of the transposase gene top was noted in all IS1111 cases (See Partridge and Hall (2003) The IS1111 family members IS4321 and IS5075 have sub-terminal inverted repeats and target the terminal inverted repeats of Tn21 family transposons. J. Bacteriol., 185(21 ): 6371 -6384).
[0166] The inventors show that an NCR is usually found downstream of the transposase gene in the IS1111 type ISs and upstream in the IS110 type ISs. To determine how these NCRs were involved in target selection, the inventors used four IS1111 family members and one IS110 family member, each targeting different sequences, to demonstrate that the NCR determines a short seeker RNA (seekRNA)that co-purifies with the transposase protein and is essential for transposition of the IS or a cargo flanked by IS ends from and to its preferred target site.
[0167] By identifying the precise role of seekRNAs for IS movement and correct target site selection, the inventors have developed gene editing systems, components, and methods of use thereof for the specific insertion of nucleic acid cargo into a desired target site.
[0168] The present disclosure provides a significant advance to gene editing. The present disclosure provides gene editing systems, components and methods of thereof that are particularly advantageous for gene editing approaches that require the specific and precisely-oriented insertion of a donor nucleic acid into a target site within the genome of a cell, wherein the donor nucleic acid cargo may comprise one or more polynucleotide sequences. The present disclosure is particularly advantageous for insertion of large nucleic acid cargo into a genome without introducing significant “scars” by insertional mutagenesis. Notably, a number of the system components described herein are newly discovered and characterised for the insertion of genetic material for the first time. Furthermore, the gene editing systems described herein do not rely upon additional enzymes such as integrases for genetic integration.Definitions
[0169] As used herein, the phrase “sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical” provides basis for a sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71 %, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identical to the specified sequence.
[0170] As used herein, a “sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%... identical” to a specified sequence includes a sequence that is at least 50%, at least 51%, at least 52%, at least 53%, at least 54%, at least 55%, atleast 56%, at least 57%, at least 58%, at least 59%, at least 60%, at least 61%, at least 62%, at least 63%, at least 64%, at least 65%, at least 66%, at least 67%, at least 68%, at least 69%, at least 70% identical to the specified sequence.
[0171] “ Insertion sequence” (IS) as defined herein refers to a DNA mobile genetic element that encode the gene or genes necessary for IS transposition (including transposase) between left and right ends (LE, RE). The IS may have left and right inverted repeat (IR) sequences within or near the LE and RE, respectively.
[0172] “Transposon” as defined herein refers to a composite DNA mobile genetic element that encodes a nucleic acid cargo (ie comprising a nucleic acid of interest) to be inserted into a genome, flanked by left and right inverted insertion sequences.
[0173] “Transposase” (Tnp; also known as a transposition enzyme) as defined herein refers to a polypeptide, fragment or functional equivalent thereof that catalyses the movement of transposable mobile genetic elements, eg insertion sequences or transposons.
[0174] “Target sequence” or “target site” as used interchangeably herein is defined as the site within the genome of interest into which a genetic insertion is or will be made.
[0175] “Gene editing system” or genome editing system” as used interchangeably herein refer to a system for modifying a genetic sequence or sequences within the genome of a cell.
[0176] “Genetic insertion”, “genome insertion”, genomic integration” or “genetic integration” as interchangeably used herein is defined as incorporation of genetic material into a prokaryotic genome. The genetic material may comprise a protein-coding sequence or part thereof. The genetic material may comprise one or more gene sequences or fragments thereof. The genetic material may encode an RNA molecule or part thereof. The genetic material may comprise regulatory elements such as promoter sequences. The genome is prokaryotic (including loci that may be located on plasmids in the cytoplasm).
[0177] “ Nucleic acid cargo”, “cargo”, or “donor nucleic acid” or “donor” as used interchangeably herein refer to a polynucleotide sequence or sequences that is / are or will be inserted into the genome of a prokaryotic cell. A nucleic acid cargo will beunderstood to comprise a “nucleic acid of interest” that is to be inserted into the genome of a prokaryotic cell.
[0178] As used herein, the terms “cargo gene” or “cargo genes” or “cargo nucleic acid” may be used interchangeably and refer to a gene sequence or sequences that is / are or will be inserted into the genome of a cell; that does not encode the polypeptide, fragment or functional equivalent thereof of the gene editing system, or encode the RNA molecule of the gene editing system. The nucleic acid cargo may comprise one or more cargo genes, and may further comprise a polynucleotide sequence or sequences encoding the polypeptide, fragment or functional equivalent thereof of the gene editing system, and / or a polynucleotide sequence or sequences encoding the RNA molecule of the gene editing system.
[0179] “Non-coding region” (NCR) as defined herein refers to a nucleic acid sequence located 3’ or 5’ to the open reading frame (ORF) of the transposase gene in the native IS 1111 or IS 110 insertion sequence, extending from the ORF to the end of the IS. The NCR does not transcribe a protein-coding sequence, but instead transcribes non-coding RNA.
[0180] “Non-coding RNA” (ncRNA) as defined herein refers to the RNA transcribed from an IS / / / / or IS 110 NCR.
[0181] “Seeker RNA” or “seekRNA” as used interchangeably herein refers to the ncRNA derived from an IS / / / / or IS / 10 NCR, which includes a sequence or sequences that is / are complementary to the target site sequence.
[0182] “Full seekRNA”, “full-length seekRNA”, or “long seekRNA” are used interchangeably herein to refer to the native full-length seekRNA sequence encoded by an NCR.
[0183] “short seekRNA” or “peak seekRNA” may be used to refer to a seekRNA corresponding to the most enriched RNA sequence from peaks identified by RNA sequencing. There may be numbered peaks, eg peakl or peak2.
[0184] “Vector” or “construct” as used interchangeably herein refer to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include, but are not limited to, nucleic acid molecules that are single-stranded,double-stranded, or partially double-stranded; nucleic acid molecules that comprise one or more free ends, no free ends (e.g., circular); nucleic acid molecules that comprise DNA, RNA, or both; and other varieties of polynucleotides known in the art. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication). Other vectors are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. It will be appreciated by those skilled in the art that the design of a vector for expression of one or more components will be influenced by factors such as the cell into which the vector is to be introduced, and the level of expression desired.
[0185] “Regulatory element” as used herein is intended to include promoters, enhancers, and other elements that control sequence expression. Regulatory elements may direct constitutive or inducible expression. Regulatory elements may also direct temporal expression.
[0186] “In trans” is used herein to describe an insertion experiment or scenario in which each of the gene editing system components (which may include the nucleic acid cargo) are separated in different locations; including different locations of the same plasmid or encoded by different plasmids or vectors. For example, the polypeptide, fragment or functional equivalent thereof can be provided as purified protein; along with the nucleic acid cargo encoded by a first expression vector (which may be referred to as pDonor), and the polynucleotide encoding the RNA molecule of the gene editing system carried by a second expression vector. In another example, the RNA molecule of the gene editing system, or the polynucleotide encoding the RNA molecule of the gene editing system, may be provided as a purified polynucleotide; along with the nucleic acid cargo encoded by a first expression vector; along with the polypeptide, fragment or functional equivalent thereof encoded by a second expression vector.
[0187] “In cis” is used herein to describe an insertion experiment or scenario in which the nucleic acid cargo comprises: the sequence or sequences to be inserted; polynucleotide sequence encoding the polypeptide, fragment or functional equivalent thereof of the gene editing system; and a polynucleotide encoding the RNA molecule of the gene editing system. In this way, “in cis" mimics a native IS arrangement, wherein the IS comprises: the Tnp gene, the NCR encoding the RNA molecule that specificallybinds of is capable of binding the target site, and a sequence or sequences for insertion into the target site; situated between the LE and RE sequences.
[0188] “Left end” (LE) as used herein describes a sequence as used herein describes a sequence 5' of the end of the nucleic acid cargo to be inserted into a target site, which enables formation of a minicircle during transposition.
[0189] “Right end” (RE) as used herein describes a sequence 3’ of the end of the nucleic acid cargo to be inserted into a target site, which enables formation of a minicircle during transposition.
[0190] “Left flank” (LF, also referred to as “left target” LT) as used herein describes a sequence comprising a 5’ sequence of the target site, wherein the 5’ sequence begins at the 5’ end of the target site sequence and ends at the point where the nucleic acid cargo is inserted into the target site (the insertion / excision point).
[0191] “Right flank” (RF, also referred to as “right target” RT) as used herein describes a sequence comprising a 3’ sequence of the target site, wherein the 3’ sequence begins at the point where the nucleic acid cargo is inserted into the target site (the insertion / excision point) and ends at the 3’ end of the target site sequence.
[0192] It will be appreciated that reference to the LF or RF in the target site in any genome, plasmid or linear DNA, refers to the left side or right side sequence, respectively of the target site relative to the insertion point.
[0193] “Transposition” as used herein defines the movement of a transposable mobile gene element (ie IS, transposon, or designed nucleic acid cargo), in which first the gene element is excised from its first location, before it is then inserted into a target site in a directional manner with specific orientation. Before insertion, formation or provision of a minicircle or circular intermediary is required. Preferably, the gene editing system provides a nucleic acid cargo that comprises a left end (LE) nucleotide sequence and a right end (RE) nucleotide sequence that enables formation of a minicircle. This allows a gene editing process more reminiscent of natural insertion of a native insertion sequence, instead of beginning the gene editing process by providing an artificial minicircle construct representing an already-formed minicircle intermediate. Enabling formation of a minicircle in trans or in cis results in scarless or near-scarless insertion, and also permits greater control regarding the orientation of insertion.
[0194] As used herein, the terms “base” and “nucleotide” may be used interchangeably. In other words, when defining the length of a nucleotide sequence, it will be appreciated that a length defined in bases (eg 5 bases), will correspond to that length when defined in nucleotides (eg 5 nucleotides).
[0195] “Minicircle” or “circular intermediate” or “circular intermediary” as used interchangeably herein, refers to the circular intermediate nucleic acid molecule formed by the junction of the 5’ LE and 3’ RE sequences. The minicircle comprises the nucleotide sequence or sequences situated between the LE and RE, optionally the nucleotide sequence or sequences situated between the LF-LE and the RE-RF. In some embodiments, the minicircle comprises from 5’ to 3’: LE - nucleic acid cargo - RE; wherein the LE and RE are joined. In preferred embodiments, the minicircle comprises from 5’ to 3’: LF - LE - nucleic acid cargo - RE - RF; wherein the 5’ and 3’ ends are joined to form the circular molecule.
[0196] “Scar DNA” as used herein defines DNA that contains secondary changes (“scars”) that are introduced into the DNA sequence during gene editing. Scar DNA includes undesired insertions due to insertional mutagenesis. Examples include insertion of overhangs or in-dels into the target site when cleaved DNA is ligated together, or insertion of plasmid backbone DNA into a target site during recombination. Scar DNA may cause frameshift mutations resulting in aberrant gene expression. Scar DNA may also be immunogenic.
[0197] “Scarless” or “scar-less” gene editing techniques, as referred to herein, refer to gene editing that permits modification of a genomic sequence without introducing scars and generating scar DNA.
[0198] “Near scarless” refers to gene editing techniques that result in very rare introduction of scars into DNA; such as less than about 50%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, less than about 5% likely to introduce scar DNA during a gene editing insertion; preferably less than about 25% likely. “Near scarless” DNA may comprise a DNA scar of less than about 200 basepairs, less than about 150 basepairs, less than about 100 basepairs, less than about 50 basepairs, less than about 40 basepairs, less than about 30 basepairs, less than about 20 basepairs, less than about 15 basepairs, less than about 10 basepairs, or fewer basepairs in length.Examples
[0199] Example 1- Materials and methodsPhylogenetic analysis and modelling of IS110 and IS1111 family members
[0200] The protein sequence of annotated members of the IS 110 and IS 1111 family members (349 sequences) were curated from ISfinder database (See Siguier et al. (2006) ISfinder: the reference centre for bacterial insertion sequences. Nucleic Acids Res 34, D32-36).
[0201] Full-length protein sequences that contained the DEDD motif and correct start codon were clustered at 70% identity and one member of each cluster was selected to generate a sequence alignment. Multiple sequence alignment of the selected members of the IS110 and IS1111 family members along with Piv comprising a total of 197 sequences was performed on MEGA (Tamura et al. (2021 ) MEGA11 : Molecular Evolutionary Genetics Analysis version 11 . Molecular Biology and Evolution 38:3022- 3027) using default settings of Clustal Omega.
[0202] Phylogenetic tree was generated on Mega using Neighbour-joining tree default settings and rooted using the midpoint. The multiple sequence alignment (MSA) was further analysed with WebLogo tool to identify conserved residues. A percent identity matrix was created using Clustal Omega. 3D structure prediction of proteins sequences was performed using AlphaFold2 and visualised on Pymol using ISEd 1 and ISEc21 as representative members of the IS110 and IS11 11 family members. Protein sequence conservation was analysed using Consurf tool at https: / / consurfdb.tau.ac.il / .Genome mining of consensus target site of IS110 and IS1111 family members
[0203] Sequences for 349 total IS110 and IS1 111 family members were extracted from ISFinder database. DNA sequences were searched against the preformatted non- redundant nucleotide BLAST database (version date 2022-11 -21 ) using BLAST+ version 2.13.0, restricting to bacterial taxonomic IDs. The BlastN output was filtered by E value of 0, 100% identity, and 95% identity; subject length >=100,000 and <=10,000,000; query coverage 100%. This resulted in the curation of GenBank coordinates where each IS was inserted. From those coordinates and extra 200 base pairs were selected from each side of the identified IS. A custom script was run to queryagainst the new + / -200 and a new database was created of IS+ / -200 bp on each side. The flanking sequences were concatenated as pre-insertion sites. All pre-insertion sites for each member of the IS 110 and IS 1111 family members were grouped and filtered after an MSA. Identical sequences were removed in order to generate a non-redundant MSA and unique insertion events. The resulting unique flanking sequences MSA were analysed by a custom script to generate a WebLogo to find the consensus target sequence for each IS.Molecular cloning and plasmids
[0204] E. coli strain DH5a was used in this study for molecular cloning and transposition assays. The sequence of selected members of the IS 110 (ISEc21 ) and IS 1111 (ISEc11 , ISKpn4, ISPal 1 , ISPst6) family members were retrieved from ISfinder, searched using BlastN to locate an insertion point of each member and 100 bps from the flanking sides were added to each IS sequence (IS+1 OObp). Each IS+100bp were synthesised by IDT as a gBIock and the sequences were cloned into pUC19 under the control of lac promoter in BamHI site. The target sequences were created as described above and the DNA sequences corresponding to the target sites for each IS where ordered as a gBIock (IDT) and cloned into pRSF Duet-1 . The chloramphenicol resistance gene (catA1 gene) containing its endogenous promoter, and the mCherry gene were ordered as gBIocks (IDT) and cloned into pCDF Duet-1 . All molecular cloning was performed by the Gibson cloning method unless specified otherwise.Purification of Tnp-RNP complex
[0205] The transposase gene (tnp) was cloned into pCDF-Duet-1 with an N-terminal 6xHis-Maltose-Binding-Protein (MBP), a TEV cleavage site and a C-terminal Strep- Tagll, to produce a tnp plasmid. For in vitro RNA transcription, the entire IS sequence were cloned into pUC19 under T7 promoter. BL21 (DE3) E. coli cells were transformed with either the tnp plasmid alone or co-transformed with the tnp plasmid and the corresponding IS plasmids by electroporation using 100 ng of each plasmid. Colonies were grown in 500 mL of LB media supplemented with Spectinomycin (100 ug / mL) and Ampicillin (100 ug / mL) until OD600nm of 0.6-0.8 when 0.5 mM IPTG (isopropyl p-D-1 - thiogalactopyranoside) was added to the cell culture to induce the expression of protein and RNA and grown for an additional 16 h at 18 °C. The cells were pelleted by centrifugation at 5000 g for 10 min, resuspended in Lysis buffer (100 mM Tris-HCI pH8.0, 500 mM NaCI, 5% Glycerol, 1 mM TCEP) and lysed by sonication. Cell debris was removed by centrifugation (30 min at 17,000 g), clarified supernatant was loaded onto 1 mL of StrepTactin-XT (high capacity) affinity chromatography column (IBA Lifesciences) and washed with 5 column volumes of Lysis buffer. The protein was eluted with addition of 50 mM Biotin to the Lysis buffer. The fractions containing the transposase were combined and concentrated using 30 kDa Amicon centrifugal concentrators and stored at -80 °C.
[0206] The transposase alone and transposes expressed in the presence of the corresponding IS samples were then used for nucleic acid digest gels and small RNA sequencing. Protein samples were prepared as standard analysis for SDS-PAGE by adding NuPAGE™ LDS sample loading buffer and boiled for 10 min to denature, then run on NuPAGE™ 4-12% Bis-Tris Protein Gels (Invitrogen) at 200 V for 20 mins in MES running buffer (Invitrogen). The gels were stained with Coomassie Brilliant Blue and destained in water.
[0207] To analyse nucleic acids bound to the transposase, 10 pL of protein alone or co-purified complex with nucleic acids were incubated with either 1 pL of DNase I (18 U / pL) or 1 pL of RNase (10 U / pL) and incubated at 37 °C for 30 min. After the digest, samples were mixed with 2X RNA loading dye (95% formamide, 10 mM EDTA, bromophenol blue, Xylene Cyanol), heated at 98 °C for 2 min to denature before running on TBE-Urea (7 M) 7% (19:1 ) denaturing polyacrylamide gel using 1X TBE running buffer and nucleic acids were stained with SYBR GOLD.Small RNA-Sequencing and analysis
[0208] 100 pL of Tnp protein alone and Tnp-nucleic acid RNP complex was incubated with 5 pL of DNase (20 U / pL) and incubated for 30 min at 37 °C, then 5 pL of Proteinase K was added and incubated for 30 min at 37 °C. Once reaction was complete, the nucleic acids was ethanol precipitated and the pellet washed with 70 % ethanol. The nucleic acid pellet was resuspended in 15 pL MOW. 2 pg of purified RNA was used as total input for preparing RNA libraries for sequencing. Collibri Stranded RNA Library Prep Kit for Illumina Systems (Thermo Fisher Scientific) was used and the libraries were prepared using manufacturer’s protocol and sequenced using an iSeq100 System (Illumina). These reads were then aligned to the sequence of theircorresponding IS using the Burrows-Wheeler Aligner alignment tool. The coverage plot was visualised using Integrative Genomics Viewer (IGV) or Tablet v1 .21 .02.08.Mini-circle formation
[0209] Mini-circle formation experiments were conducted using full-length IS sequences+100bp flanking sequences cloned into pUC19. E. coli strain DH5a was transformed with plasmid containing the full IS sequence and cultured in LB media for 16 h at 37 °C, then subsequent standard plasmid isolation was carried out (Bioline). DNA template was diluted to 10 ng / pL and 1 pL was used to perform inverse PCR using primers facing outwards from the IS. Exemplary primer sets are shown in Table 11 .PCR was performed in a final reaction volume of 20 pL using Phusion™ reaction buffer, 0.2 mM of dNTPs, 0.5 pM of each primer, and 1 U of Phusion™ polymerase.Thermocycler conditions consisted of denaturation cycle (98 °C, 2 min) followed by 35 cycles of denaturation (98 °C, 15 s), annealing (52 °C, 20 s), extension (72 °C, 20 s), and a final cycle of amplification (72 °C, 40 s). The PCR product was then run on a 1% agarose gel, then isolated using PCR gel extraction kit (Bioline) and sequenced to identify the mini-circle junction product.
[0210] Table 11 - Exemplary oligonucleotides used to detect mini-circle formationIn vitro transposition assay
[0211] In vitro transposition assay was carried out upon co-transforming E. co / / strain DH5a with the pDonor plasmid containing the entire IS+1 OObp flanking sequences in pUC19 (Amp resistance) and the pTarget plasmid containing the target sequence specific to each IS cloned into pRSF Duet-1 (Kan resistance). Co-transformations were performed by electroporation. The transformants were plated onto agar plates and grown overnight at 37 °C. Colonies were grown into 2 mL of LB media supplemented with the antibiotics for 16 h at 37 °C. Standard plasmid isolation was carried out (Bioline). The resulting plasmid DNA templates were diluted to 10 ng / pL and 10 ng was used as template for PCR to detect transposition. Primer pair sets used were complementary to the IS (R) in the tnp and the pTarget plasmid backbone (F) (Table 10). The alternate pair (F) on IS and (R) on pTarget plasmid backbone was also performed. The PCR reaction was performed as the standard condition described in the mini-circle formation section, with the following changes, extension (72 °C, 30 s), and final cycle of amplification (72 °C, 60 s). The PCR products were run on 1 % agarose gel and the corresponding band was gel extracted using PCR gel extraction kit (Bioline) and further sequenced by Sanger method using one of the PCR primers. The sequencing result was analysed by MAFFT alignments on the target sequence and target sequence+IS left and right ends to confirm that correct insertion had occurred.Rescued transposition
[0212] Rescued transposition assays were performed by co-transforming the pTarget plasmid containing the target site for the IS, pDonor IS ANCR containing the IS sequence + 100bp flanking sequence cloned into pUC19, pNCR containing the corresponding NCR sequence of the IS followed by the HDV sequence was cloned under the T7 promoter into pCDF Duet-1 vector by electroporation into E. coli BL21 (DE3) cells. The transformants were cultured in LB media until ODeoo of 0.4-0.6 was reached then induced using 0.5 mM IPTG and grown at 25 °C for 16 h before harvesting. Standard plasmid isolation was carried out (Bioline). The resulting plasmid DNA templates were diluted to 10 ng / pL and 10 ng was used as template for PCR to detect transposition. Primer pair sets used were complementary to the IS (R) in the tnp and the pTarget plasmid backbone (F). The alternate pair (F) on IS and (R) on pTarget plasmid backbone was also performed. The PCR reaction was performed as thestandard condition described in the mini-circle formation section, with the following changes, extension (72 °C, 30 s), and final cycle of amplification (72 °C, 60 s). The PCR products were run on 1% agarose gel and the corresponding band was gel extracted using PCR gel extraction kit (Bioline) and further sequenced by Sanger method using one of the PCR primers. The sequencing result was analysed by MAFFT alignments on the target sequence and target sequence+IS left and right ends to confirm that correct insertion had occurred.In vitro transposition assay in trans
[0213] Transposition experiments with each element of the IS cloned into a different plasmid were performed as follows: pTarget plasmid containing the target site for the IS, pDonor containing the catA 1 gene flanked by 50 bp of LE (left end) and RE (right end) of the IS, as well as 50 basepairs of the LF (left flank) and RF (right flank) of the insertion site cloned into pUC19. The transposase tnp gene and NCR sequence encoding seekRNA followed by HDV sequence were cloned into pCDF Duet-1 downstream of the T7 promoter. The plasmids were electroporated as described, and grown in 2 mL of LB culture at 37 °C until ODeoo 0.6-0.8 was reached, then protein and seekRNA expression was induced by adding 0.5 mM IPTG to the culture media and grown for 4 h at 25 °C. The culture was harvested and plasmid purification and transposition detection was performed as described in the mini-circle formation session using primers that are complementary to the cargo gene (in this case, catA 1) (F primer) and the pTarget plasmid backbone (R primer) (Table 12). mCherry reporter transposition assay in trans
[0214] To investigate the efficacy of various engineered seekRNA truncations and mutations, a mCherry reporter transposition assay was created using a novel pDonor’ and pTarget’ plasmids. The pDonor’ contains the tnp and the seekRNA-encoding sequence under the control of a T7 promoter, followed by a strong T7 terminator sequence to prevent downstream gene expression. This configuration ensures that the mCherry gene flanked by 50 bp of LE (left end) and RE (right end) of the IS as well as 100 bp of the LF (left flank) and RF (right flank) of the insertion site cloned after the T7 terminator is not expressed. pTarget plasmid contains the IS target sequence cloned in front of the T7 promoter. mCherry expression and fluorescence can be detected if transposition has occurred. Exemplary primer oligonucleotides are shown in Table 12.
[0215] 100 ng of each plasmid (pDonor and pTarget) were either individually transformed or co-transformed into E. coli BL21 (DE3) cells via electroporation. Cells were cultured on agar plates containing Spectinomycin and Kanamycin, incubated overnight at 37 °C in 2 mL LB media. Subsequently, a fresh 2 mL LB medium was inoculated with the initial culture and grown until reaching an ODeoo of 0.6-0.8, then induced with 0.5 mM IPTG and incubated for 4 hours at 25 °C.
[0216] Cell assays involved transferring 100 pL of each culture into a 96-well Corning clear bottom plate. Cell density was quantified via absorbance at 600 nm, while mCherry fluorescence was measured using a TECAN infinite M1 OOOPro plate reader in bottom reading mode, with excitation at 587 nm and emission at 610 nm, each with a bandwidth of 5 nm, and the optimal gain set to 100%. To adjust fluorescence measurements for cell density, fluorescence per cell (FOD) was calculated as the fluorescence intensity divided by the OD600. The percent increase in mCherry fluorescence attributable to transposition was determined by the formula:
[0217] This analysis was validated through at least three independent transformations and co-transformations as biological triplicates and Sanger sequencing was performed to validate the transposition of the mCherry cargo gene.Quantification of mCherry reporter transposition
[0218] Transposition frequency was measured by Flow Cytometry using BD LSRFortessa™ X-20 Flow Cytometer. 100 ng of each plasmid (pDonor and pTarget) was co-transformed into E. coli BL21 (DE3) cells via electroporation. Cells were plated on fresh agar plates containing Spectinomycin, Kanamycin and 0.1 mM IPTG to induce expression of Transposase and seekRNA, as well as expression of mCherry after transposition. The plates were incubated for 16 hours at 37 °C, followed by 4 hours at room temperature. Entire agar plates were scraped containing hundreds of colonies and resuspended and mixed evenly in 1 mL of LB media. Cells were diluted 1 in 2 in PBS (Phosphate Buffered Saline, pH 7.4) and run on the Flow Cytometer. Around 35000- 70000 cells were run for each sample until at least 25000 events were recorded, which were gated for single live cells. The cells were counted using forward and side scatter channels, and mCherry fluorescence intensity was detected per cell. The transpositionfrequency was measured according to the number of single cells with high levels of mCherry fluorescence intensity over total number of single cells. Each set of samples was created by three independent transformations as biological repeats, and the transposition frequency was plotted as bar graphs. Sanger sequencing was performed to validate the transposition of the mCherry cargo gene. Flow cytometry data was analysed using 880 FlowJo™ v10.10 Software (BD Life Sciences).
[0219] Table 12 - Exemplary oligonucleotides used to detect transposition
[0220] Example 2 - Characterisation of IS 1111 and IS 110 family features
[0221] The inventors observed that the NCR of IS110 and some of its relatives typically sits upstream of the transposase-encoding gene tnp instead of downstream as in the IS 1111 type (Figure 1 A). The upstream NCR appears to be characteristic of the IS 110 family, albeit with some exceptions.
[0222] Using AlphaFold2, the structures of several transposases derived from each family were modelled, and ISEc11 , as a representative for the IS 1111 family, and ISEc21 , as a representative of the IS 110 family, are shown in Figures 1 B and 1 C. The transposases all consist of an N-terminal DEDD (RuvC fold) catalytic domain and a C- terminal domain with no well characterised homologues (includes Pfam PF02371 ) separated by a coiled coil with a variable region between the a-helices. The variable region, which is composed of small a-helices and loops, is longer in the IS110 type transposases. AlphaFold2 also predicts the formation of a dimer or tetramer via interactions between the coiled coils placing the variable region of one monomer next to the RuvC domain of the second monomer.
[0223] A phylogeny of the transposases constructed from the set of IS listed in ISFinder, curated to include only a single representative for each group of transposases with >70% pairwise amino acid identity (Figure 13), revealed that the \S 1111 family members, classified based on the presence of sTIR (sub-terminal inverted repeats) of appropriate length, and IS / 10 members are clearly separate. A simplified phylogenetic tree (Figure 1 D) highlights some of the exceptions. In the tightly clustered group of IS 1111 relatives all but one, ISMtspI 7, had a downstream NCR in contrast to its closest relative ISPye21. The more diverged IS ISHvoW and ISAcp5 have long NCRs both up and downstream. For the IS in the IS110 family, the NCR was found mainly upstream of tnp but it is downstream in some e.g. ISYps2. Hence, the inventors concluded that the IS1111 family is indeed distinct from the IS 110 family and that the mechanisms used by these two families may differ significantly, particularly with regard to end recognition.
[0224] The most unusual feature of the IS 7111 and IS 110 families is that the target recognised by each IS or group of IS is not the same. A computational pipeline to identify a consensus target site was developed and applied to the IS sequences found in ISFinder under the IS1 10 family (Figure 12). In some cases, a consensus was notgenerated because there were insufficient different locations, and in others it appears that the ends listed in ISFinder may be incorrect and need adjustment.
[0225] Example 3 - The downstream NCR in IS 1111 is essential for transposition.
[0226] ISEc11 , an IS 1111 family member that recognises a short target (eg GTGAAAATACTG, SEQ ID NO: 66) that is not part of a potentially folded region, has been previously shown to form a circular intermediate in which an active promoter is generated at the junction and to transpose to its preferred target (See Prosseda et al. (2006) Plasticity of the P junc promoter of ISEc11 , a new insertion sequence of the IS 1111 family. J. Bacterio / ., 188: 4681 -4689). Here, ISEc11 (cloned in pUC) produced a circular intermediate that was readily detected using PCR (Figure 2A). In contrast, when the 3’-NCR was deleted (bp 1 170-1374 of 1443 bp removed) leaving the transposase gene and ends intact, the circular form was not produced. However, when the 270 bp fragment that includes the NCR (preceded by a T7 promoter and followed by an HDV ribozyme) was supplied in a separate plasmid, the circular form was again produced (data not shown). Likewise, transposition of ISEc11 to its preferred target in a separate plasmid required the presence of the NCR in cis or in trans (Figure 2B). In the product, the IS was cleanly inserted into the target and as expected no additional bases were generated (Figure 2C). Previously, a 4 bp target site duplication was claimed (See Prosseda et al. (2006) Plasticity of the P junc promoter of ISEc11 , a new insertion sequence of the IS 1111 family. J. Bacterio / ., 188: 4681 -4689) but here one copy of the 4 bp was assigned to within the IS. Replacement of target bases flanking the donor on either side also prevented both circle formation and transposition (Figure 2D), demonstrating that intact target-derived sequence on both sides of the donor is needed for excision to form the circular intermediate.
[0227] Example 4 - IS 1111 transposases co- purify with NCR-derived RNA.
[0228] The ISEc11 transposase, TnpEc11 , was expressed (using a HisMBP and Strep Tagil fusion), either with or without the full length ISEd 1 present in the same cells to supply the NCR, and then purified. The yield of TnpEc11 was poor in the absence of the NCR region but improved when the complete ISEd 1 (preceded by a T7 promoter) was present to supply the NCR (Figure 3A) The TnpEd 1 expressed in the presence of the NCR, purified with a nucleic acid and distinct bands were clearly present. The nucleic acid was digested by RNase but not by DNase (Figure 3B). RNA extracted from theaffinity purified RNP complex (Figure 3C) was sequenced, and the reads were aligned with the ISEc11 sequence where they mapped clearly to the 3’-NCR (Figure 3D). Hence, RNA transcribed from the essential NCR, was specifically associated with the transposase.
[0229] To confirm that this property is found more widely in IS 1111 family members, three further IS that target different sites were examined. ISKpn4, ISPst6 and ISPal 1 have previously been shown to form a circular intermediate and are found at a specific location in a potentially folded structure, an attC site of integron associated gene cassettes for ISKpn4 (See Post and Hall (2009) Insertion sequences in the IS1111 family that target the attC recombination sites of integron-associated gene cassettes. FEMS Microbiol. Lett., 290:182-187) and ISPst6 (See Tetu and Holmes (2008) A family of insertion sequences that impacts integrons by specific targeting of gene cassette recombination sites, the IS1111 -attC group. J. Bacteriol., 190: 4959-4970) and a Pseudomonas REP for ISPal 1 (See Partridge and Hall (2003) The IS1111 family members IS4321 and IS5075 have sub-terminal inverted repeats (IRs) and target the terminal inverted repeats of Tn21 family transposons. J. Bacteriol., 185(21 ): 6371 - 6384). As for ISEc11 , TnpKpn4 TnpPst6 and TnpPal 1 all purified with RNA (Figure 7) and this RNA mapped to the region downstream of tnp (Figures 3E-F; Figures 8-9).
[0230] Example 5 - seekRNA enables target selection for IS 1111 family members and can be programmed for selective gene editing
[0231] To locate the target-determining region or regions, the inventors searched in the folded structure of the predominant 82 bp band associated with TnpEc11 for matches to the forward and / or the reverse strand of the target sequence. Two short segments, one matching the forward strand and one matching the reverse, were found. These matches overlap in the target (Figure 3G). As these matches can explain the ability to select a specific target, the NCR-derived RNA was named a seeker RNA or seekRNA. The folded structure for the longer RNA is in Figure 8. The folded structures of the seekRNA corresponding to the predominant peak for ISPal 1 , ISKpn4 and ISPst6 were also examined (Figures 3H and 3I, Figures 8-9). Predicted folds for the corresponding long seekRNAs are in Figure 8-9. Again, two short segments matching the forward and the reverse strand of the target sequence were found in the sequence of the predominant 96 or 82 bp band. The ISPst6-derived seekRNA contained the same matches as theISKpn4 seekRNA (Figure 8) as expected, given that TnpPst6 is 86% identical (92% similar) to TnpKpn4 and they insert into the same position in the same target.
[0232] The inventors also examined the predicted NCR of ISPa25 that also targets this site but encodes a significantly diverged transposase (TnpKpn4 and TnpPa25 are 46% identical; 60% similar). The region in the DNA sequences where they most closely match includes the region corresponding to the short seekRNA of ISKpn4 (Figure 8). The predicted seekRNA for ISKpn4 (was folded and found to include the same stretches of sequence matching the target (Figure 8). Hence the seekRNA does not appear to be confined to interaction only with a specific group of very closely related transposases.
[0233] ISEc11 and ISXne4 represent an example of natural re-programming of the seekRNA. When the predicted seekRNA of ISXne4, a relatively close relative of ISEc11 (TnpEc11 and TnpXne4 are 72% identical; 88% similar), was compared to the seekRNA of ISEc11 (Figure 10), a high level of identity was observed but the target-matching bases in the seek RNA of ISEc11 were not found (Figure 4A). Examination of the location of the few known copies of ISXne4 revealed that it was surrounded by a different sequence and that this target matched the differing regions (Figure 4A). This confirms that the two target-matches have been correctly identified and demonstrates re-programming.
[0234] To re-programme the IS to move to a different target, it was necessary to change both the sequences corresponding to the target that flank the donor IS and the target matching sequences in the seekRNA. For ISEc11 , two new targets were tested using movement of a mini IS containing the mCherry gene without an upstream promoter bounded by the IS ends (LE 50 bp, RE 46 bp) and flanks containing the target (Fig. 4D). 244 The TnpEc11 and the appropriate long (154 nt) seekRNA were supplied in the donor plasmid. Movement to targets (Fig. 4E and Fig. 9) that were preceded by a T7 promoter enabled expression of the mCherry, measured by FACS sorting (Figure 14). Using the wildtype target and seekRNA, transposition occurred at a frequency of about 15% (Fig. 4E-F). When the portion of the target that flanks the IS on the right was altered and the corresponding changes were made in the seekRNA (see Fig. 9), transposition of the mCherry to the new M1 target occurred at about 23% frequency. In the second case, the target on both sides of the donor mCherry mini IS andcorresponding positions in the seekRNA were altered, and again the transposition frequency to the new M2 target was 15%. Hence, the long seekRNA could be programmed to move the IS to a different location. Using the same assay, the short 82 nt seekRNA (peak in Fig. 3d) was also tested and the mCherry mini IS moved even more efficiently (42% transposition; Fig. 4E, bottom line) indicating that this length is sufficient to support IS movement.
[0235] In order to demonstrate that a nucleic acid cargo can be moved by these IS, the catA 1 chloramphenicol resistance gene (745 bp) was introduced at the start (after bp 50) and at the end (after bp 1397) of ISEc11 (Figure 4B). In both cases, movement was detected. The centre of ISEd 1 was also replaced by the catA 1 gene with an upstream promoter leaving the ends intact and the transposase and NCR supplied in trans. When the transposase and NCR were supplied in trans, again movement was detected as seen for mCherry (Figure 4D-F), indicating that the system can be harnessed to mobilise different cargos (Figure 4C).
[0236] Example 6 - The upstream NCR in IS110 is essential for transposition.
[0237] ISEc21 is an IS110 family member with an upstream NCR (Figure 5A). The target listed in ISFinder was confirmed computationally and here determined to be a linear target after it was traced to a conserved region within certain IS3 family members, namely the codons for the second of the DDE in the catalytic domain of the transposase.
[0238] As for ISEd 1 , ISEc21 formed a circular intermediate in which the ends were abutted and a promoter generated (Figure 5A). When the upstream NCR was deleted (bp 20-150 removed), mini-circle formation no longer occurred (Figure 5A). The NCR was also needed in cis or in trans for transposition (Figures 5B and 5C) and movement cleanly into the target supplied without generating a duplication of bases at the target site was detected (Figure 5D). When the number of target bases matching the consensus was reduced to 5 / 6 on the left and 5 on the right of the donor IS, both minicircle formation and transposition still occurred (Figure 5E) but removal of further bases on each side abolished movement again indicating the importance of the presence of the target sequence surrounding the donor IS.
[0239] Example 7 - IS110 transposases co-purify with NCR-derived RNA.
[0240] The TnpEc21 transposase purified with an RNA (Figure 7) that was recovered and sequenced. The sequences mapped to the upstream NCR (Figure 6A, Figure 11 ). Again, a shorter seekRNA was more abundant than the long RNA (see Figure 11 for the fold of the long seekRNA). Bases matching one strand of the target were found in the short 74 nt seekRNA (Figure 6B). This contrasts with the situation in the short seekRNAs from all of the IS1111 family members tested here, where short matches to overlapping stretches in the target were found.
[0241] A case of natural re-programming was also found in an ISEc21 relative. The transposase of ISMch6 is 72% identical (85% similar) to TnpEc21 (Figure 6C) but is found in a modified target. Comparison of the DNAs of the NCR revealed three altered bases in the target-defining region of the folded short seekRNA for ISEc21 (Figure 1 1 ) and these differences were also found in the ISEc21 target.
[0242] Example 8 - seekRNA enables target selection for IS110 family members and can be programmed for selective gene editing
[0243] Using a donor plasmid with the components in the configuration shown in Figure 6B, the 74 nt ISEc21 seekRNA was found to be functional when shortened even further to 58 nt (Figure 6C) and the target-matched region in this shortened seekRNA form was reprogrammed to move mCherry, surrounded by the outer IS ends and flanked by the appropriate target site, to its normal site and to two different new sites. Formation of the circular intermediate and transposition only occurred when the target site surrounding the donor was the same as the target offered and seekRNA included appropriate nucleotides to detect that target (Figure 6D and Figure 12).
[0244] Example 9 - LE and RE truncation can modify transposition efficiency
[0245] A transposition assay, as described herein (see eg Figure 2), was performed to assess transposition when the LE and RE of ISEc21 were shortened. Plots of cells containing a target plasmid (pTarget) and a pDonor and using ISEc21 seekRNA full and mCherry as a cargo (listed above the plots) were gated for all cells (FSC-A and SSC-A), single cells (FSC-A and FSC-H) and mCherry expression shown in each plot. The transposition efficiency of the ISEc21 was detected as the percentage of the total single cells expressing mCherry.
[0246] Shortening the LE and RE sequences of ISEc21 to each be 16 bases in length (removing nucleotides from the end inwards) increased transposition efficiency compared to the native LE and RE sequences, increasing efficiency by 19% from 11 % to 30% as measured by %mCherry positive cells (Figure 14A, B). Further shortening the LE and RE sequences of ISEc21 to each be 7 nucleotides long also saw an improvement in transposition efficiency compared to the native ISEc21 LE and RE sequences, increasing efficiency by 8% from 11 % to 17% (Figure 15A, C). Complete removal of the LE and RE sequences resulted in only 2% transposition efficiency (Figure 15D) - indicating that whilst transposition of ISEc21 can still occur in the absence of both LE and RE sequences, efficiency is low. The results demonstrate that for LE and RE sequences derived from ISEc21 , the optimal length of these sequences lies between 7-16 nucleotides in length.
[0247] Shortening the LE and RE sequences of ISEc1 1 affected minicircle formation and transposition. Minicircle formation was assessed by PCR employing outward facing primers R and F and Sanger sequencing confirming the PCR products (Figure 16A). A transposition assay was performed to assess transposition, as described herein (see eg Figure 2). Plasmids encoding ISEc11 (pDonor) were co-transformed with plasmids encoding the target sequence of ISEc11 (pTarget) into E. coli. Addition of IPTG induced expression of transposase and seekRNA. Transposition products result in the formation of a plnsert plasmid with ISEc11 inserted within its target site. In vitro transposition was detected using primers sets complementary to the target plasmid and the cargo sequence (mCherry) and target site insertion. PCR products indicated transposition into target pSFA59 (Figure 16B).
[0248] Removal of the LE and RE ends of ISEc11 resulted in no minicircle formation or transposition of ISEd 1 (pSFA152; Figure 16). Removal of the IR for the LE and RE of ISEc11 , along with truncation of the LE to 7 bases in length and truncation of the RE to 3 bases in length, also resulted in no minicircle formation or transposition of ISEd 1 (pSFA151 ; Figure 16).
Claims
CLAIMS1 . A gene editing system for inserting a nucleic acid cargo comprising a nucleotide sequence of interest into a genome of a prokaryotic cell, the system comprising:(i) a polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40, or a polynucleotide encoding the same; and(ii) an RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell, and / or a polynucleotide sequence encoding the same; wherein the nucleic acid cargo comprises a left end (LE) nucleotide sequence and a right end (RE) nucleotide sequence to enable formation of a circular nucleic acid molecule; wherein the nucleic acid cargo comprises in the 5’ to 3’ direction: a left flank (LF) sequence corresponding to the target site in the genome of the cell, the LE, the nucleotide sequence of interest, the RE, and a right flank (RF) sequence corresponding to the target site in the genome of the cell; wherein the polypeptide, fragment or functional equivalent thereof of (i) and the RNA molecule of (ii) are capable of forming a gene editing complex for enabling the insertion of the nucleic acid cargo into the target site within the genome.
2. The gene editing system of claim 1 , the gene editing system enables scarless, or near scarless, insertion of the nucleic acid cargo into the target site within the genome.
3. The gene editing system of claim 2, wherein the nucleic acid cargo is in the form of a circular nucleic acid molecule.
4. The gene editing system of claim 1 , wherein the polypeptide, fragment or functional equivalent thereof has an amino acid sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%,at least 99% or 100% identical to the amino acid sequence set forth in any one of SEQ ID NOs: 33-40.
5. The gene editing system of claim 1 , wherein the polypeptide, fragment or functional equivalent thereof is encoded by a polynucleotide sequence comprising or consisting a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence as set forth in any one of SEQ ID NOs: 41 -48.
6. The gene editing system of claim 1 , wherein the polypeptide, fragment or functional equivalent thereof is encoded by a polynucleotide sequence comprising or consisting a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence as set forth in any one of SEQ ID NOs: 1 -16; optionally any one of SEQ ID NOs: 2, 4, 6, 8, 10, 1 , 14, or 16.
7. The gene editing system of claim 1 , wherein the polynucleotide encoding an RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 25-32.
8. The gene editing system of claim 1 , wherein the polynucleotide encoding an RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 17-24.
9. The gene editing system of claim 1 , wherein the polynucleotide encoding an RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 1 -16; optionally any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, or 16.
10. The gene editing system of claim 1 , wherein the RNA molecule comprises or consists of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, atleast 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 50-57.11 . The gene editing system of claim 1 , wherein the target site comprises one or more sequences comprising or consisting of a nucleotide sequence that is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99% or 100% identical to the nucleotide sequence set forth in any one of SEQ ID NOs: 58-131 ; optionally any one of SEQ ID NOs: 58, 66, 67, 76, 82, 1 12, 115, 127 and 129.
12. The gene editing system of claim 1 , wherein the system comprises one or more expression vectors encoding one or more of the system components.
13. The gene editing system of claim 1 , wherein the system further comprises:(iii) an exogenous nucleic acid cargo comprising one or more nucleotide sequences of interest to be inserted into the target site within the genome; wherein the exogenous nucleic acid cargo comprises a left end (LE) nucleotide sequence and a right end (RE) nucleotide sequence to enable formation of a circular nucleic acid molecule; wherein the exogenous nucleic acid cargo comprises 5in the 5’ to 3’ direction: a left flank (LF) sequence corresponding to the target site in the genome of the cell, the LE, the nucleotide sequence of interest, the RE, and a right flank (RF) sequence corresponding to the target site in the genome of the cell; wherein the polypeptide, fragment or functional equivalent thereof of (i) and the RNA molecule of (ii) are capable of forming a gene editing complex for enabling the insertion of the exogenous nucleic acid cargo into the target site within the genome.
14. The gene editing system of claim 13, wherein the system comprises one or more expression vectors encoding the polypeptide, fragment or functional equivalent thereof of (i); the RNA molecule of (ii); and / or the exogenous nucleic acid cargo of (iii); or any combination thereof.
15. An isolated, synthetic and / or recombinant polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40.
16. An isolated, synthetic or recombinant polynucleotide comprising or consisting of the nucleotide sequence set forth in any one of SEQ ID NOs: 41 -48.
17. An isolated, synthetic or recombinant polynucleotide comprising or consisting of the nucleotide sequence set forth in any one of SEQ ID NOs: 25-32.
18. An isolated, synthetic or recombinant polynucleotide comprising or consisting of the nucleotide sequence set forth in any one of SEQ ID NOs: 50-57.
19. An expression vector comprising a polynucleotide sequence or sequences encoding the gene editing system of any one of claims 1 -14.
20. An expression vector of claim 17, comprising the polynucleotide of any one of claim 16 or claim 17.21 . An expression vector comprising a polynucleotide sequence or sequences encoding:(i) a polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40; and / or(ii) an RNA molecule comprising or consisting of the nucleotide sequence set forth in any one of SEQ ID NOs: 50-57.
22. The expression vector of claim 21 , further comprising (iii) a polynucleotide sequence or sequences encoding a nucleic acid cargo for insertion into a target site within a genome of a prokaryotic cell, wherein the nucleic acid cargo comprises a left end (LE) nucleotide sequence and a right end (RE) nucleotide sequence to enable formation of a circular nucleic acid molecule, and wherein the nucleic acid cargo comprises in the 5’ to 3’ direction: a left flank (LF) sequence corresponding to the target site in the genome of the cell, the LE, the nucleotide sequence of interest, the RE, and a right flank (RF) sequence corresponding to the target site in the genome of the cell.
23. A non-naturally occurring or engineered composition comprising the gene editing system of any one of claims 1 -14; the polypeptide, fragment or functional equivalent thereof of claim 15; the polynucleotide of any one of claims 16-18; and / or the expression vector of any one of claims 19-22.
24. A method for genome editing, comprising providing to a prokaryotic cell the gene editing system of any one of claims 1 -14; the polypeptide, fragment or functional equivalent thereof of claim 15; the polynucleotide of any one of claims 16-18; and / or the expression vector of any one of claims 19-22.
25. A method for inserting a nucleic acid cargo comprising a nucleotide sequence of interest into a genome of a prokaryotic cell comprising providing to a prokaryotic cell the gene editing system of any one of claims 1 -14; the polypeptide, fragment or functional equivalent thereof of claim 15; the polynucleotide of any one of claims 16-18; and / or the expression vector of any one of claims 19-22.
26. The method of claim 25, wherein the insertion of the nucleic acid cargo into the genome is scarless.
27. A method of generating an in vitro or ex vivo model of a disease or disorder caused by or suspected to be caused by a prokaryotic cell, comprising providing to a prokaryotic cell the gene editing system of any one of claims 1 -14; the polypeptide, fragment or functional equivalent thereof of claim 15; the polynucleotide of any one of claims 16-18; and / or the expression vector of any one of claims 19-22.
28. An engineered prokaryotic cell comprising the gene editing system of any one of claims 1 -14; the polypeptide, fragment or functional equivalent thereof of claim 15; the polynucleotide of any one of claims 16-18; and / or the expression vector of any one of claims 19-22.
29. An engineered prokaryotic cell obtained by the method of any one of claims 24 to 27.
30. A kit comprising the gene editing system of any one of claims 1 -14; the polypeptide, fragment or functional equivalent thereof of claim 15; the polynucleotide of any one of claims 16-18; and / or the expression vector of any one of claims 19-22.31 . A gene editing system for inserting a nucleic acid cargo comprising a nucleotide sequence of interest into a genome of a prokaryotic cell, the system comprising:(i) a polypeptide, fragment or functional equivalent thereof comprising or consisting of the amino acid sequence set forth in any one of SEQ ID NOs: 33-40, or a polynucleotide encoding the same; and ii) an RNA molecule that specifically binds or is capable of specifically binding to a target site within the genome of the cell, and / or a polynucleotide sequence encoding the same; wherein the nucleic acid cargo comprises a left end (LE) nucleotide sequence and a right end (RE) nucleotide sequence that are concatenated to form of a circular nucleic acid molecule (eg artificial minicircle vector); wherein prior to concatenation, the nucleic acid cargo comprises in the 5’ to 3’ direction: a left flank (LF) sequence corresponding to the target site in the genome of the cell, the LE, the nucleotide sequence of interest, the RE, and a right flank (RF) sequence corresponding to the target site in the genome of the cell; wherein the polypeptide, fragment or functional equivalent thereof of (i) and the RNA molecule of (ii) are capable of forming a gene editing complex for enabling the insertion of the nucleic acid cargo into the target site within the genome.
Citation Information
Patent Citations
Programmable DNA transposases for nucleic acid manipulation
WO2024119154A1