Compositions and methods for genome editing
Patent Information
- Application Number
- JP2023578194
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-25
- Filing Date
- 2022-06-24
- Publication Date
- 2025-06-26
AI Technical Summary
Existing genome editing technologies, such as CRISPR/Cas9 and transposon-encoded CRISPR-Cas systems, are limited in their ability to recombinantly modify entire exons and are restricted by the size of the DNA strands, lacking control and precision in sequence substitution.
A molecular complex comprising two single-stranded nucleic acid molecules with specific binding sites for a transposase, allowing site-specific recombination and sequence substitution, independent of target sequence length, using transposase-mediated tagmentation.
Enables precise and controlled recombinant editing of genetic sequences, allowing for the substitution of targeted regions with user-defined sequences, overcoming size limitations and sequence-specific constraints of previous methods.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to compositions and methods for genome editing.
[0002] The desire to understand the effects of modifying the genetic information of living cells dates back to the earliest days of genetics.
[0003] First, conventional genetics attempts to understand genetic modifications and resulting phenotypes by selecting specific gene sites.
[0004] Subsequently, biochemists have used radiation and chemical mutagens to increase the probability of genetic mutation in experimental organisms. These methods, while very useful, are expensive and do not allow for easy control of the modifications introduced into the genetic material.
[0005] Knowledge of molecular biology and molecular mechanisms for cellular repair or defense against the host organism has made it possible to develop numerous techniques, the so-called reverse genetics, the aim of which is the opposite of so-called classical genetic screening.
[0006] Reverse genetics aims to introduce mutations into genetic material for the purpose of measuring and analyzing the resulting phenotypic effects.
[0007] Genome editing strategies have evolved over the last 30 years, and the latest innovation and innovation in the context of targeted gene modification is the CRISPR / Cas9-associated CRISPR system (CRISPR / Cas-9). In this regard, two patents can be mentioned: European Patent No. 3144390 B1 and U.S. Patent No. 8697359 B1, both of which describe this technology.
[0008] Therefore, the CRISPR / Cas-9 system has established itself as a reference tool for gene modification. However, this CRISPR modification is similar to molecular scissors and is not sufficient to allow complete recombination of entire exons.
[0009] More recently, the transposase encoded by CRISPR (transposon-encoded CRISPR-Cas system) system is a half-answer to the shortcomings of CRISPR technology. In fact, this technology uses both targeted integration via transposon Tn7 and subsequent addition of sequences to the genome by CRISPR technology. However, in any case, it is not recombination.
[0010] Prime editing technology is another possibility for gene replacement that combines CRISPR tools with reverse transcriptase. This latest advance in genome editing allows for the replacement of one of the two DNA strands, depending on the size, but it is still inherent in cellular integration and repair systems and their associated strong limitations.
[0011] The present invention also aims, inter alia, to overcome these drawbacks of the prior art.
[0012] One of the aims of the present invention is to provide recombinant tools that allow genome editing.
[0013] Another object of the present invention is that this new tool is independent of the size of the sequence being manipulated, and that the system is user controlled and controllable.
[0014] It is yet another object of the present invention to provide a tool that allows for single- or double-stranded replacement of a molecule of interest at the user's discretion using a DNA repair system in a controlled and limited manner.
[0015] The present invention provides a first single-stranded nucleic acid molecule comprising or consisting essentially of an A sequence that allows for the insertion of a complementary sequence of a nucleic acid of interest, The present invention relates to a first single-stranded nucleic acid molecule, wherein the A sequence is linked at its 5' end to a first T-rich sequence of 40 to 60 nucleotides in length and at its 3' end to a second T-rich sequence of 40 to 60 nucleotides in length, the first and second T-rich sequences each comprising a first and a second domain of 6 to 12 G / C-rich nucleotides, the sequence of the first domain being complementary to the sequence of the second domain, and the first and second domains being located 15 to 52 nucleotides from the A sequence, and the first molecule comprises at its 5' end a first sequence oriented 5' to 3' for recognizing a transposase and at its 3' end at least one second sequence for recognizing a transposase.
[0016] The present invention also provides a first single-stranded nucleic acid molecule comprising or consisting essentially of an A sequence that allows for insertion of a complementary sequence of a nucleic acid of interest, a first single-stranded nucleic acid molecule, wherein the A sequence binds at its 5' end to a first A / T-rich, particularly T-rich, sequence of 40 to 60 nucleotides in length and at its 3' end to a second A / T-rich, particularly T-rich, sequence of 40 to 60 nucleotides in length, the first and second A / T-rich, particularly T-rich, sequences comprising first and second domains of 6 to 12 G / C-rich nucleotides, respectively, the sequence of the first domain being complementary to the sequence of the second domain, the first and second domains being located 15 to 52 nucleotides from the A sequence, and the first molecule comprising at its 5' end a first sequence oriented 5' to 3' for recognizing a transposase and at its 3' end at least one second sequence for recognizing a transposase; a second single-stranded nucleic acid molecule comprising, or consisting essentially of, at its 5' end at least one complementary sequence of a second sequence for recognizing a transposase, The present invention relates to a molecular complex in which a first and a second single-stranded nucleic acid molecule are paired according to base complementarity defined by Watson-Crick to define two double-stranded binding sites for a transposase.
[0017] This is because the present invention provides a molecular complex comprising a first single-stranded nucleic acid molecule and a second single-stranded nucleic acid molecule, wherein the second single-stranded nucleic acid molecule comprises, or essentially consists of, at its 5' end, a complementary sequence to at least one of a second sequence for recognizing a transposase; means relating to a molecular complex in which a first and a second single-stranded nucleic acid molecule are paired according to base complementarity as defined by Watson-Crick to define two double-stranded binding sites for a transposase.
[0018] The present invention is based on the unexpected observation made by the inventors that the use of specific single-stranded guides that can target regions of a nucleic acid of interest allows for the recruitment of transposases in a controlled, "site-specific" manner, thus making it possible to use the recombination properties of transposases to replace sequences in the molecule of interest.
[0019] The molecular complex described above is actually the basic unit of the technology defined in the present invention.This basic unit is useful for directing recombinase to the specific site where recombination, and therefore sequence replacement, must occur.Unlike the CRISPR / Cas9 system, which requires the presence of PAM type (NGG) sequence, the molecular tool defined herein can be used for any target sequence, regardless of its sequence.
[0020] The aforementioned molecular complex is therefore a basic unit completed by: a region of homology to the target sequence, and Replacement region of target sequence.
[0021] It is therefore an intermediate product of the tools described below.
[0022] The molecular complex consists of two single-stranded nucleic acid molecules, which can be DNA molecules, RNA molecules, or mixed RNA and DNA molecules.
[0023] These two molecules are partially complementary to each other according to the Watson-Crick rule for nucleic acid base complementarity, i.e., adenine pairs with thymidine or uracil, cytosine pairs with guanine, and vice versa.
[0024] More specifically, each of the two molecules forming the complex comprises one of the sequences of the double-stranded molecule strands corresponding to the binding sequence of transposase.In addition, each single-stranded molecule thus comprises a transposase binding "half sequence", and therefore cannot interact with the corresponding transposase.On the other hand, when the two molecules of the complex interact together, a double-stranded molecule is thus formed by base pairing as defined herein above, and reconstitutes the double-stranded binding site of transposase, so that the latter can interact with the formed molecule.
[0025] The first molecule.
[0026] The first molecule of the complex is a molecule containing a nucleic acid sequence that, when modified, allows specific targeting of a desired region of a nucleic acid molecule of interest. This desired sequence is selected by the user of the system according to the selected target. This desired sequence is inserted into the first molecule of the complex in region A. This region A corresponds to at least two nucleic acids between which a sequence that allows targeting of the target molecule is inserted. Considering the oriented structure of nucleic acids (5' to 3'), it is important that the sequence that allows targeting of the desired region is positioned in the correct direction to enable pairing with the target sequence.
[0027] Advantageously, the A region also contains one or more sites for recognizing a restriction enzyme to facilitate oriented insertion. One or more of the following sites may be present in the A region:
[0028] [Table 1 (01)]
[0029] [Table 1 (02)]
[0030] [Table 1(03)]
[0031] [Table 1(04)]
[0032] [Table 1 (05)]
[0033] [Table 1 (06)]
[0034] Obviously, in the context of chemical synthesis of the first molecule, it is not necessary to have a cloning (or insertion) site for the sequence that allows targeting of the target region, but rather great care must be taken to provide a correctly oriented sequence, although this is of course possible.
[0035] The first molecule further comprises A / T-rich sequences, or in the case of RNA, A / U-rich sequences, on either side of the A region to allow for a certain flexibility of structure. A / T-rich or A / U-rich is understood in the present invention to mean a sequence containing more than 50% A or T or U, preferably more than 50% T or U, relative to the total number of nucleotides constituting the sequence. These sequences on either side of the A region have a nucleotide size ranging from 10 nucleotides to 60 nucleotides.
[0036] The flexibility of these sequences flanking the A region, due to the presence of numerous A, T, or U bases, may have the effect of allowing poorly regulated recombinase-mediated recombination, perhaps even when the complex has not yet recognized the target molecule.
[0037] To overcome this problem, a GC-rich sequence is introduced into each of the A / T-rich sequences, particularly the T-rich sequence or the A / U-rich sequence, adjacent to the A region. These G / C-rich regions consist of 6 to 12 nucleotides, and the amount of C or G bases exceeds 50% of the nucleotides contained in the G / C-rich sequence.
[0038] To stabilize the structure of the first molecule and, as described earlier herein, to prevent inadvertent recombination, a G / C-rich region is located 15 to 52 nucleotides from the end of the A region.
[0039] For clarity, if the A region consists of three nucleotides (the central nucleotide corresponds to position 0), the A / T-rich region or the A / U-rich region starts at position −2 on the left and +2 on the right. Thus, the G / C-rich region is located from positions −17 to −54 on the left and from positions +17 to +54 on the right.
[0040] Another important factor is that the G / C-rich sequence to the right (or 5') of the A region is necessarily complementary to the G / C-rich region to the right (or 3') of the A region (according to the Watson-Crick pairing rules). Also, the first single-stranded molecule pairs with itself in the G / C-rich region, preventing any recombination by the transposase unless there is an interaction with the complementary target sequence of the region to be inserted into the A region of the first molecule.
[0041] Finally, the first molecule contains at its 5' end a sequence corresponding to a first site for binding to a transposase and at its 3' end a second site for binding to a transposase.
[0042] The first and second binding sites are advantageously the same, in particular both corresponding to the same strand of the double-stranded binding site of the transposase, which means that only the first transposase binding site present in the 5' region of the first molecule can pair integrally, and therefore stably, with the transposase binding site present in the 3' region.
[0043] The first and second binding sites are preferably the same, but each corresponds to a different strand of the double-stranded transposase binding site. For example, if the first transposase binding site corresponds to the sense strand, the second transposase binding site corresponds to the sequence of the complementary strand. This can then have one of two configurations: i) the second binding site corresponding to the complementary strand is oriented in a 3' to 5' direction (in which case it can pair with the first transposase binding site to form a double-stranded site), or ii) the second binding site corresponding to the complementary strand is oriented in a 5' to 3' direction (in which case it cannot pair with the first transposase binding sequence because their orientations are not complementary). In case i), if the first single-stranded molecule pairs with itself at the first and second binding sites, the aforementioned complex cannot be formed because there are no longer any single-stranded complementary regions available to pair with the second molecule to form two double-stranded transposase binding sites.
[0044] Furthermore, if the first molecule lacks a complementary sequence to the target region in portion A, or if it contains such a target sequence but does not interact (pair) with this target sequence, it forms a three-dimensional structure in which the entire molecule is single-stranded except for the regions corresponding to the G / C-rich regions that pair with each other.
[0045] A linear schematic of the first molecule is shown in [Figure 1], and a schematic of its paired form is shown in [Figure 2].
[0046] The second molecule.
[0047] The second molecule of the complex is simpler than the first molecule. The second molecule contains a transposase binding site in its 5' portion that is complementary to the site for binding to a transposase present in the 3' portion of the first molecule. Thus, when the complex is formed, the (single-stranded) transposase binding half-site located at 3' of the first molecule can pair with the (single-stranded) transposase binding half-site located at 5' of the second molecule to form a double-stranded transposase binding site, a double-stranded site to which a transposase can bind.
[0048] The 3' portion of the second molecule has a region similar to region A of the first molecule, which allows it to accept a specific sequence corresponding to the sequence to be inserted in place of the target molecule of interest. Below is a more detailed description of how to prepare the second molecule to allow this substitution.
[0049] Complex The complex formed from the first and second molecules is shown schematically in Figure 3.
[0050] The complex is such that when the first and second molecules are paired via the transposase binding half-sites, the complex is able to bind a transposase dimer (a functional dimer that allows recombination).
[0051] In addition, either the first or second molecule further comprises a complementary sequence of the transposase binding half-site located 5' of the first molecule. This complementary region of the transposase binding half-site located 5' of the first molecule can be located 5' or 3' of the first molecule, or even 5' of the second molecule, preferably 5' of the complementary sequence of the transposase binding half-site located 3' of the first molecule.
[0052] In the present invention, "first and second single-stranded nucleic acid molecules are such that they pair according to base complementarity defined by Watson-Crick to define two double-stranded binding sites for a transposase." Since the first molecule contains at least one transposase binding half-site in its 5' region and at least one transposase binding half-site in its 3' region, and the second molecule also contains at least one transposase binding half-site, this means that * the first molecule comprises, in its 5' region, a first sequence of a first transposase binding site and a complementary sequence of the first sequence of the first transposase binding site, and, in its 3' portion, a second sequence of a second transposase binding site; * the second molecule comprises a sequence complementary to the second sequence of the second transposase binding site; or * a first molecule comprising in its 5' region a first sequence of a first transposase binding site and in its 3' portion a second sequence of a second transposase binding site and a complementary sequence of the first sequence of the first transposase binding site; * the second molecule comprises a sequence complementary to the second sequence of the second transposase binding site; or * the first molecule comprises in its 5' region a first sequence of a first transposase binding site and in its 3' portion a second sequence of a second transposase binding site; * where the second molecule comprises a complement of the first sequence of the first transposase binding site and a complement of the second sequence of the second transposase binding site, It means that during pairing between the first and second molecules, two complete sites are formed.
[0053] The use of the term "at least" and the fact that the molecule "comprises" a sequence that forms a transposase recognition site allows one skilled in the art to select the position of the half-sequence that forms the binding site, so that ultimately, when the complex is formed, two complete sites are reconstituted.
[0054] Three options (two of which are detailed below) are shown schematically in Figure 4.
[0055] Advantageously, the invention relates to a complex as described above, in which the A sequence comprises the complementary sequence of the nucleic acid of interest.
[0056] As described above, the A region can contain a complementary sequence of a nucleic acid of interest. More specifically, the sequence contained in the A region of the first molecule of the complex is complementary to the 5' or 3' sequence of the sequence of the molecule of interest, such that the complex specifically recognizes this region of the nucleic acid molecule and allows the complex to replace the adjacent region with a region complementary to the region complementary to the sequence contained in the A region.
[0057] In other words, the present invention advantageously provides: a first single-stranded nucleic acid molecule comprising or consisting essentially of an A sequence that allows insertion of a complementary sequence of a nucleic acid of interest, wherein the complementary A sequence is bound at its 5' to a first A / T-rich, particularly T-rich, sequence of 40-60 nucleotides in length and at its 3' to a second A / T-rich, particularly T-rich, sequence of 40-60 nucleotides in length, the first and second A / T-rich, particularly T-rich, sequences comprising first and second domains of 6-12 G / C-rich nucleotides, respectively, the sequence of the first domain being complementary to the sequence of the second domain, the first and second domains being located 15-52 nucleotides from the A sequence, and the first molecule comprising at its 5' end a first sequence oriented 5' to 3' for recognizing a transposase and at its 3' end at least one second sequence for recognizing a transposase; a second single-stranded nucleic acid molecule comprising, or consisting essentially of, at its 5' end at least one complementary sequence of a second sequence for recognizing a transposase, The aforementioned complex, wherein the first and second single-stranded nucleic acid molecules are paired according to Watson-Crick defined base complementarity to define two double-stranded binding sites for the transposase.
[0058] In one advantageous embodiment, the invention relates to a method for the preparation of a nucleic acid sequence comprising a first molecule comprising, at its 5' end, a first sequence oriented 5' to 3' for recognizing a transposase, and at its 3' end, a second sequence oriented 5' to 3' for recognizing a transposase, The second molecule comprises, at its 5' end, a first complementary sequence of the first sequence for recognizing the transposase, followed by a second complementary sequence of the second sequence for recognizing the transposase.
[0059] In this advantageous embodiment of the complex of the present invention, the first molecule comprises a first transposase recognition sequence at the 5' position and a second transposase recognition sequence at the 3' position. The second molecule then comprises, at the 5' position, a first complementary sequence to the first transposase recognition sequence of the first molecule, followed by a second complementary sequence to the second transposase recognition sequence of the first molecule. Each molecule of the complex also comprises two transposase binding half-sites, such that upon complex formation, i.e., upon pairing of the first molecule with the second molecule, two adjacent double-stranded transposase binding sites are formed, to which a transposase dimer can then bind.
[0060] [Figure 5] shows a schematic diagram of this embodiment.
[0061] Advantageously, the present invention provides a method for the preparation of a nucleic acid sequence comprising a first molecule comprising, at its 5' end, a first sequence oriented from 5' to 3' for recognizing a transposase, and at its 3' end, a second sequence for recognizing a transposase, followed by a first complementary sequence of the first sequence for recognizing a transposase; The second molecule comprises at its 5' end a complementary sequence of the second sequence for recognizing the transposase.
[0062] In this advantageous embodiment of the complex of the invention, the first molecule comprises a first transposase recognition sequence at the 5' position and a second transposase recognition sequence at the 3' position, the latter immediately followed by a first complement of the transposase recognition sequence located 5' of the first molecule, and the second molecule then comprises a second complement of the second transposase recognition sequence of the first molecule at the 5' position.
[0063] Also in this embodiment, a first molecule can reform a transposase recognition double-stranded binding site by pairing a first recognition sequence 5' of the first molecule with a first recognition sequence 3' of the molecule. A second molecule must then pair with the first molecule to reconstitute a second double-stranded transposase binding site using a second recognition sequence 3' of the first molecule and a second complementary transposase recognition sequence located 5' of the second molecule.
[0064] FIG. 6B illustrates this embodiment schematically.
[0065] Another advantageous embodiment of the complex according to the invention can be envisaged, in which the first molecule comprises a first complementary sequence of the first transposase recognition site at the 5' position, followed by the first transposase recognition site. Furthermore, the first molecule comprises a second transposase recognition site at the 3' position. The second molecule then remains unchanged with respect to the previous embodiment.
[0066] Here, the 5' portion of the first molecule folds back on itself to reconstitute a double-stranded transposase recognition site by pairing with the first recognition site and the immediately adjacent complementary sequence. This embodiment is shown in Figure 7.
[0067] Additional similar embodiments exist in which the first molecule comprises a first complementary sequence 5' to the first transposase site, immediately followed by a first site for recognizing the first transposase.
[0068] Advantageously, said transposase is a bacterial transposase chosen from the transposase of transposon Tn5, the transposase of transposon Tn9, the transposase of transposon Tn10, Tn903, Tn602, or even the transposase of transposon Tc1, or more generally the transposases of the mariner transposon superfamily.
[0069] Other examples of transposases that can be used in the context of the present invention are the Vibrio harveyi transposase (the transposase characterized by Agilent and used in the product SureSelect QXT), the MutA transposase and Mu transposase recognition site including the terminal sequences R1 and R2, the transposase of Staphylococcus aureus transposon Tn552, the transposase of transposon Tn7, the Tn / O and IS10 transposases, and the transposase of transposon Tn3.
[0070] The Tn5 transposase is the best known. It is encoded by the Tnp gene of the transposon Tn5. The transposase initiates transposition by forming a transposase dimer that binds to its target sequence. In association with this complex, the transposase then catalyzes four phosphoryl transfer reactions (DNA cleavage, DNA hairpin formation, hairpin resolution, and strand transfer to the target DNA), resulting in the integration of the transposon into its new DNA site, a process known as "tagmentation."
[0071] The present invention is based on this tagmentation principle: by using the tagmentation properties of transposase, the complex allows for the targeted insertion of one sequence into another sequence.
[0072] Also, in the context of the present invention, when a transposase is mentioned, it refers to one of the aforementioned transposases, namely the transposase of transposon Tn5, Tn9, Tn10 or Tc1 / mariner (or a transposase mutated to increase their transposition or tagmentation activity).
[0073] In the present invention, when a transposase that binds to a complex formed by a first molecule and a third molecule and a transposase that binds to a complex formed by a second molecule and a third molecule are used simultaneously, a pair of transposases derived from transposons Tn5 and Tn10 is preferred.
[0074] In one advantageous embodiment, the invention relates to a kit comprising a vector allowing the expression of a first molecule of said complex and a vector allowing the expression of said second molecule.
[0075] In the context of this kit, the vector is preferably a circular molecule of double-stranded DNA that has all the elements to allow its replication in a host cell (prokaryotic and / or eukaryotic cell) and that has elements to allow the expression of the first or second molecule of the complex.
[0076] If the first and second molecules must be in the form of single-stranded DNA molecules, the sequence of each of the first and second molecules is under the control of a sequence that allows the synthesis of single-stranded DNA from double-stranded DNA. This is the case, for example, with the f1 bacteriophage origin of replication sequence contained in a phagemid-type vector. Thus, in the presence of the auxiliary phase M13, which contains all the genes necessary to activate the f1 sequence, the vector produces single-stranded DNA from double-stranded plasmid DNA.
[0077] Thus, the kit may contain either two independent vectors each containing the sequence of one or the other of the first and second molecules that form the complex described above, or a single vector containing the two sequences but genetically isolated from each other.
[0078] The aforementioned kits may also contain other elements such as a transposase to enable transposition.
[0079] Advantageously, said complex is such that the first and second transposase recognition sequences are sequences for recognizing a Tn5 transposase having one of the following sequences: -CTGtCTCTTataCAcAtcT (SEQ ID NO: 1), -CTGACTCTTataCACAagT (SEQ ID NO: 3), and -CTGtCTCTTgatCAgATCT (SEQ ID NO: 5).
[0080] As a result, the corresponding complementary sequence is: - AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), - ActTGTGtatAAGAGTCAG (SEQ ID NO: 4), and -AGATcTGatcAAGAGaCAG (SEQ ID NO: 6).
[0081] Other transposase recognition sequences are as follows: Tn5MErev, 5'-[phos]CTGTCTCTTATACACATCT-3' (SEQ ID NO: 11) Tn5ME-A (Illumina FC-121-1030), 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 12) and Tn5ME-B (Illumina FC-121-1031); 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 13)
[0082] The further sequence is as follows: Sense sequence of SEQ ID NO: i antisense sequence of SEQ ID NO: i+1, i ranges from 269 to 424.
[0083] This means, for example, that the following respective sense and antisense sequence pairs are considered: SEQ ID NO:269 and SEQ ID NO:270, SEQ ID NO:271 and SEQ ID NO:272, SEQ ID NO:273 and SEQ ID NO:274, SEQ ID NO:275 and SEQ ID NO:276, SEQ ID NO:277 and SEQ ID NO:278, SEQ ID NO:279 and SEQ ID NO:280, SEQ ID NO:281 and SEQ ID NO:282, SEQ ID NO:283 and SEQ ID NO:284, SEQ ID NO:285 and SEQ ID NO:286, SEQ ID NO:287 and SEQ ID NO:288, SEQ ID NO:289 and SEQ ID NO: :290, SEQ ID NO:291 and SEQ ID NO:292, SEQ ID NO:293 and SEQ ID NO:294, SEQ ID NO:295 and SEQ ID NO:296, SEQ ID NO:297 and SEQ ID NO:298, SEQ ID NO:299 and SEQ ID NO:300, SEQ ID NO:301 and SEQ ID NO:302, SEQ ID NO:303 and SEQ ID NO:304, SEQ ID NO:305 and SEQ ID NO:306, SEQ ID NO:307 and SEQ ID NO:308, SEQ ID NO:309 and SEQ ID NO:310, SEQ ID NO:311 and SEQ ID NO:312, SEQ ID NO:313 and SEQ ID NO:314, SEQ ID NO:315 and SEQ ID NO:3 16, SEQ ID NO: 317 and SEQ ID NO: 318, SEQ ID NO: 319 and SEQ ID NO: 320, SEQ ID NO: 321 and SEQ ID NO: 322, SEQ ID NO: 323 and SEQ ID NO: 324, SEQ ID NO: 325 and SEQ ID NO: 326, SEQ ID NO: 327 and SEQ ID NO: 328, SEQ ID NO: 329 and SEQ ID NO: 330, SEQ ID NO: 331 and SEQ ID NO: 332, SEQ ID NO: 333 and SEQ ID NO: 334, SEQ ID NO: 335 and SEQ ID NO: 336, SEQ ID NO: 337 and SEQ ID NO: 338, SEQ ID NO: 339 and SEQ ID NO: 340, SEQ ID NO: 341 and SEQ ID NO: 34 2. SEQ ID NO: 343 and SEQ ID NO: 344, SEQ ID NO: 345 and SEQ ID NO: 346, SEQ ID NO: 347 and SEQ ID NO: 348, SEQ ID NO: 349 and SEQ ID NO: 350, SEQ ID NO: 351 and SEQ ID NO: 352, SEQ ID NO: 353 and SEQ ID NO: 354, SEQ ID NO: 355 and SEQ ID NO: 356, SEQ ID NO: 357 and SEQ ID NO: 358, SEQ ID NO: 359 and SEQ ID NO: 360, SEQ ID NO: 361 and SEQ ID NO: 362, SEQ ID NO: 363 and SEQ ID NO: 364, SEQ ID NO: 365 and SEQ ID NO: 366, SEQ ID NO: 367 and SEQ ID NO: 368,SEQ ID NO: 369 and SEQ ID NO: 370, SEQ ID NO: 371 and SEQ ID NO: 372, SEQ ID NO: 373 and SEQ ID NO: 374, SEQ ID NO: 375 and SEQ ID NO: 376, SEQ ID NO: 377 and SEQ ID NO: 378, SEQ ID NO: 379 and SEQ ID NO: 380, SEQ ID NO: 381 and SEQ ID NO: 382, SEQ ID NO: 383 and SEQ ID NO: 384, SEQ ID NO: 385 and SEQ ID NO: 386, SEQ ID NO: 387 and SEQ ID NO: 388, SEQ ID NO: 389 and SEQ ID NO: 390, SEQ ID NO: 391 and SEQ ID NO: 392, SEQ ID NO: 393 and SEQ ID NO: 394, SEQ ID NO: 395 and SEQ ID NO: 396, SEQ ID NO: 397 and SEQ ID NO: 398, SEQ ID NO: 399 and SEQ ID NO: 400, SEQ ID NO: 401 and SEQ ID NO: 402, SEQ ID NO: 403 and SEQ ID NO: 404, SEQ ID NO: 405 and SEQ ID NO: 406, SEQ ID NO: 407 and SEQ ID NO: 408, SEQ ID NO: 409 and SEQ ID NO: 410, SEQ ID NO: 411 and SEQ ID NO: 412, SEQ ID NO: 413 and SEQ ID NO: 414, SEQ ID NO: 415 and SEQ ID NO: 416, SEQ ID NO: 417 and SEQ ID NO: 418, SEQ ID NO: 419 and SEQ ID NO: 420, SEQ ID NO: 421 and SEQ ID NO: 422, SEQ ID NO: 423 and SEQ ID NO: 424,
[0084] Advantageously, the first G / C rich domain of the first molecule has the sequence
[0085] [Table 2] corresponds to the first G / C-rich domain, and so the second G / C-rich domain is the same. In fact, as the molecule folds back on itself, the second G / C-rich domain is in a complementary and antiparallel orientation to the first G / C-rich domain, with the interaction occurring in the palindromic region (underlined in the sequence herein above).
[0086] The first and second G / C rich domains also have the sequences
[0087] [Table 3] The foregoing description applies mutatis mutandis.
[0088] Other G / C rich domain sequences may be as follows: A first G / C-rich domain of the sequence GGTCGC (SEQ ID NO: 427) and a second C / C-rich domain of the sequence GCGACC (SEQ ID NO: 428).
[0089] These examples are given for illustrative purposes only and are not intended to limit the scope of the invention.
[0090] In an advantageous embodiment, the A / T-rich sequence of the first molecule of the complex consists essentially of A or T or consists of A or T.
[0091] Even more advantageously, the A / T rich sequence of the first molecule of the complex consists of Ts.
[0092] Even more advantageously, said complex is such that it comprises the following sequence corresponding to the first molecule:
[0093] [Table 4] where X represents zero nucleotides, two nucleotides, or at least one restriction site.
[0094] The first transposase binding site is shown adjacent and the second transposase binding site is underlined.
[0095] Even more advantageously, said conjugate is such that it comprises the following sequence corresponding to the second molecule:
[0096] [Table 5] where Y represents zero nucleotides, two nucleotides, or at least one restriction site.
[0097] The first transposase binding site is shown adjacent and the second transposase binding site is underlined.
[0098] Advantageously, said complex is such that it comprises the following sequence corresponding to the first molecule:
[0099] [Table 6] where X represents zero nucleotides, two nucleotides, or at least one restriction site.
[0100] The complementary sequences of the first transposase binding site and the second transposase binding site are underlined, and the first transposase binding site is italicized and underlined.
[0101] In this embodiment, the conjugate is such that it comprises the following sequence corresponding to the second molecule:
[0102] [Table 7] where Y represents zero nucleotides, two nucleotides, or at least one restriction site.
[0103] The complementary sequence of the second transposase binding site is underlined.
[0104] Advantageously, the invention relates to the following conjugates: - the first molecule of the following sequence: 5'-TGCAGCTGR1TTTTTTTTTTTTTTTTTTTTTTTTTGGCGATCGCTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTGCGATCGCCTTTTTTTTTTTTTTTTGATACATGTTTR2CTGTAAGC-3' (SEQ ID NO: 436) wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), X corresponds to a sequence that allows recognition of the target region and, a second molecule comprising the sequence: 5'-R2CAGCTGCAGACAAAGCTTACAGR1TTTTTTTTTTTTTTTTTTTTTTcatatgccaagtY-3' SEQ ID NO: 437 wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), Y corresponds to no nucleotide or to a replacement sequence for the target region.
[0105] This means that the complex consists of molecules with the following sequence: 5'-TGCAGCTGCTGtCTCTTataCAcAtcTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTGGCGATCGCTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTGATACATGTTTAgaTgTGtatAAGAGaCAG CTGTAAGC-3' (SEQ ID NO: 438) and 5'-AgaTgTGtatAAGAGaCAGCAGCTGCAGACAAAGCTTACAGCTGtCTCTTataCAcAtcTTTTTTTTTTTTTTTTTTTTTTTcatatgccaagtY-3' (SEQ ID NO: 439) Or, X and Y are as defined herein above.
[0106] Advantageously, the invention relates to the following conjugates: - the first molecule of the following sequence:
[0107] [Table 8] wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), X corresponds to a sequence that allows recognition of the target region and, a second molecule comprising the sequence:
[0108] [Table 9] wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), Y corresponds to no nucleotide or to a replacement sequence for the target region.
[0109] Advantageously, the invention relates to the following conjugates: - the first molecule of the following sequence: 5'-TGCAGCTGR2TTTTTTTTTTTTTTTTTTTTTTTTTGGCGATCGCTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTCGCGATCGCCTTTTTTTTTTTTTTTTGATACATGTTTR2CTGTAAGC-3' (SEQ ID NO: 643) wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), X corresponds to a sequence that allows recognition of the target region and, a second molecule comprising the sequence: 5'-R1CAGCTGCAGACAAAGCTTACAGR1TTTTTTTTTTTTTTTTTTTTTTcatatgccaagtY-3' SEQ ID NO: 644) wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), Y corresponds to no nucleotide or to a replacement sequence for the target region.
[0110] Advantageously, the invention relates to the following conjugates: - the first molecule of the following sequence:
[0111] [Table 10] wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), X corresponds to a sequence that allows recognition of the target region and, a second molecule comprising the sequence:
[0112] [Table 11] wherein R1 is 5'-CTGtCTCTTataCAcAtcT (SEQ ID NO: 1); R2 is 5'-AgaTgTGtatAAGAGaCAG (SEQ ID NO: 2), Y corresponds to no nucleotide or to a replacement sequence for the target region.
[0113] Advantageously, said complex consists of the following pair of sequences:
[0114] [Table 12(01)]
[0115] [Table 12(02)]
[0116] [Table 12(03)]
[0117]
Table 12(04)
[0118]
Table 12(05)
[0119]
Table 12(06)
[0120]
Table 12(07)
[0121]
Table 12(08)
[0122]
Table 12(09)
[0123]
Table 12(10)
[0124]
Table 12(11)
[0125]
Table 12(12)
[0126]
Table 12(13)
[0127] [Table 12(14)]
[0128] [Table 12(15)]
[0129] [Table 12(16)]
[0130] [Table 12(17)]
[0131] In other words, the present invention advantageously relates to a complex as described above comprising a pair of first and second molecules, wherein the first and second molecules comprise the following respective sequences: SEQ ID NO: 440 and SEQ ID NO: 441 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 440 and SEQ ID NO: 441 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 440 and SEQ ID NO: 442 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 440 and SEQ ID NO: 442 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 440 and SEQ ID NO: 443 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 440 and SEQ ID NO: 443 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 444 and SEQ ID NO: 445 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 444 and SEQ ID NO: 445 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 444 and SEQ ID NO: 446 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 444 and SEQ ID NO: 446 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 444 and SEQ ID NO: 447 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 444 and SEQ ID NO: 447 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 448 and SEQ ID NO: 445 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number ranging from 269 to 424, the same as R1), SEQ ID NO: 448 and SEQ ID NO: 445 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 448 and SEQ ID NO: 446 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 448 and SEQ ID NO: 446 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 448 and SEQ ID NO: 447 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 448 and SEQ ID NO: 447 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 449 and SEQ ID NO: 445 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 449 and SEQ ID NO: 445 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 449 and SEQ ID NO: 446 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 449 and SEQ ID NO: 446 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 449 and SEQ ID NO: 447 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 449 and SEQ ID NO: 447 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 450 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 450 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 450 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 450 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 450 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 450 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424); SEQ ID NO: 454 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 454 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 454 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 454 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 454 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number ranging from 269 to 424, the same as R1), SEQ ID NO: 454 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 455 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 455 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 455 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 455 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 455 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 455 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424); SEQ ID NO: 456 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 456 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 456 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 456 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 456 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number ranging from 269 to 424, the same as R1), SEQ ID NO: 456 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 457 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 457 and SEQ ID NO: 451 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 457 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number ranging from 269 to 424, the same as R1), SEQ ID NO: 457 and SEQ ID NO: 452 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 457 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number ranging from 269 to 424, the same as R1), SEQ ID NO: 457 and SEQ ID NO: 453 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 458 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 458 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 458 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424); SEQ ID NO: 458 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 458 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 458 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 462 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 462 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 462 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 462 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424); SEQ ID NO: 462 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 462 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 463 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424); SEQ ID NO: 463 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 463 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424); SEQ ID NO: 463 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 463 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 463 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 464 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424); SEQ ID NO: 464 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 464 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 464 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424); SEQ ID NO: 464 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 464 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 465 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424); SEQ ID NO: 465 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 465 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424); SEQ ID NO: 465 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424); SEQ ID NO: 465 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424), SEQ ID NO: 465 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 466 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is the same even number as R1, ranging from 269 to 424); SEQ ID NO: 466 and SEQ ID NO: 459 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424), SEQ ID NO: 466 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 466 and SEQ ID NO: 460 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424); SEQ ID NO: 466 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an even number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n+1, where n is an even number, the same as R1, ranging from 269 to 424), SEQ ID NO: 466 and SEQ ID NO: 461 (wherein R1 is any one of the sequences of SEQ ID NO: n, where n is an odd number ranging from 269 to 424, and R2 is any one of the sequences of SEQ ID NO: n-1, where n is the same odd number as R1, ranging from 269 to 424).
[0132] Advantageously, the invention relates to a complex as described above, comprising one of the 300 pairs of first and second molecules of Table 4 below.
[0133] [Table 13(01)]
[0134] [Table 13(02)]
[0135] [Table 13(03)]
[0136] [Table 13(04)]
[0137] [Table 13(05)]
[0138] [Table 13(06)]
[0139] [Table 13(07)]
[0140] In another aspect, the present invention provides a method for producing a pharmaceutical composition comprising: a first single-stranded nucleic acid molecule comprising or consisting essentially of an A sequence that allows insertion of a complementary sequence of a nucleic acid of interest, or comprising the complementary sequence of a nucleic acid of interest, wherein the complementary A sequence is bound at its 5' to a first A / T-rich, particularly T-rich, sequence of 40-60 nucleotides in length and at its 3' to a second A / T-rich, particularly T-rich, sequence of 40-60 nucleotides in length, the first and second A / T-rich, particularly T-rich, sequences comprising first and second domains of 6-12 G / C-rich nucleotides, respectively, the sequence of the first domain being complementary to the sequence of the second domain, the first and second domains being located 15-52 nucleotides from the A sequence, and the first molecule comprising at its 5' end a first sequence oriented 5' to 3' for recognizing a transposase and at its 3' end at least one second sequence for recognizing a transposase; a second single-stranded nucleic acid molecule comprising or consisting essentially of a B sequence that allows insertion of a complementary sequence of a nucleic acid of interest, or comprising the complementary sequence of a nucleic acid of interest, wherein the complementary B sequence is bound at its 5' to a third T-rich sequence of 40 to 60 nucleotides in length and at its 3' to a fourth T-rich sequence of 40 to 60 nucleotides in length, the third and fourth T-rich sequences comprising third and fourth domains of 6 to 12 G / C-rich nucleotides, respectively, the sequence of the third domain being complementary to the sequence of the fourth domain, and the third and fourth domains being located 15 to 52 nucleotides from the B sequence; and the second molecule comprising at its 5' end at least one first sequence oriented 5' to 3' for recognizing a transposase and at its 3' end a second sequence for recognizing a transposase; a second single-stranded nucleic acid molecule, wherein sequence B is the complement of the nucleic acid of interest, sequence A is located 5' of the region of interest in the nucleic acid of interest, and sequence B is located 3' of the region of interest in the nucleic acid of interest; a third single-stranded molecule, in its 5' portion, at least one complementary sequence of the second sequence for recognizing the transposase of the first molecule; a third single-stranded molecule comprising, in its 3' portion, at least one complementary sequence of the first sequence for recognizing a transposase (of the second molecule); a region located between the complement of the second sequence for recognizing the transposase (of the first molecule) and the complement of the first sequence for recognizing the transposase (of the second molecule), which allows insertion of a replacement nucleic acid molecule (single stranded), The present invention relates to an ensemble, wherein first and third single-stranded nucleic acid molecules are paired according to base complementarity defined by Watson and Crick to define two double-stranded binding sites for a transposase, and wherein second and third single-stranded nucleic acid molecules are paired according to base complementarity defined by Watson and Crick to define two double-stranded binding sites for a transposase.
[0141] Said assemblies therefore include the molecular complexes described hereinabove, and all embodiments and technical details described hereinabove for molecular complexes apply mutatis mutandis to said assemblies.
[0142] In this embodiment of the invention, a collection of three molecules is described that allows for precise replacement of a sequence contained in a nucleic acid of interest with a sequence of choice.
[0143] The assembly according to the present invention is based on the above-described complex, to which a third molecule similar to the first molecule is added. [Figure 8] shows a schematic diagram of the assembly according to the present invention.
[0144] The first molecule of the complex corresponds to the first molecule of the assembly, the second molecule of the complex corresponds to the third molecule of the assembly, and the third molecule of the assembly is additionally structurally similar to the first molecule of the assembly or complex.
[0145] When the three molecules are correctly paired, the assembly according to the present invention contains two pairs of transposase-recognized double-stranded binding sites: a first pair obtained by hybridization of the first molecule with a third molecule; and A second pair obtained by hybridization of the second molecule with a third molecule. Includes.
[0146] Sequence A contained in a first molecule of the collection is complementary to the same strand of nucleic acid that is complementary to sequence B contained in a second molecule. In other words, sequence A contained in the first molecule and sequence B contained in the second molecule can simultaneously hybridize to the same nucleic acid because the two sequences A and B do not recognize the same sequence.
[0147] To clarify the observation further, the interest of the collection defined in the present invention is to propose a first molecule and a second molecule (both as defined herein above), each of which has sequences A and B capable of recognizing, on the one hand, a sequence located 5' of a target sequence of a nucleic acid of interest, and, on the other hand, a sequence located 3' of the same target sequence of a nucleic acid of interest. Thus, sequences A and B are complementary to regions of the nucleic acid molecule of interest that flank the sequence of interest to be replaced.
[0148] Thus, the first and second molecules of interest are required to flank the sequence to be replaced, thereby specifically targeting the molecule of interest.
[0149] The third molecule of the collection is then a molecule that provides the nucleic acid molecule containing the replacement sequence.
[0150] From a mechanistic point of view, the assembly according to the present invention is such that it consists of three molecules, the first and second molecules being spatially and structurally organized such that their G / C-rich regions are paired.
[0151] On either side of the third molecule, i.e., on either the 5' and 3' sides, two pairs of sites for binding to the transposase by hybridization with the first and second molecules allow the transposase dimer to bind to the assembly when the transposase is present.
[0152] Note that the transposase binding site of the first molecule may be the same as or different from the transposase binding site of the second molecule. If the binding sequences are the same, the transposase at the 5' end of the third molecule (due to hybridization of the 5' portion of the third molecule with the first molecule) and the transposase at the 3' end of the third molecule (due to hybridization of the 3' portion of the third molecule with the second molecule) are the same. Also, by way of example, if the binding sites are all binding sites for Tn5 transposase, the assembly will associate with two Tn5 transposase dimers.
[0153] It is also possible that the binding sites of the first and second molecules do not recognize the same transposase, in which case, according to the definition of ensemble set forth herein above, the 5' portion of the third molecule, upon hybridization with the first molecule, forms two double-stranded sites for binding to the first transposase, and the 3' portion of the third molecule, upon hybridization with the second molecule, forms two double-stranded sites for binding to the second transposase.
[0154] Next, when the aforementioned assembly linked to two transposase dimers is brought into the presence of a nucleic acid molecule of interest, whose 5' portion is complementary to sequence A of a first molecule of the assembly and whose 3' portion is complementary to sequence B of a second molecule of the assembly, the nucleic acid molecule of interest pairs with the assembly at the aforementioned regions A and B. This interaction has the consequence of disrupting the interaction of the two G / C-rich sequences of each of the first and second molecules of the assembly. Thus, the third molecule and the nucleic acid molecule of interest move close to each other in space, allowing the transposases to exert their tagmentation activity, so that the sequence of the molecule of interest flanked by the complementary sequences of sequences A and B is replaced at 5' and 3' by the sequence of the third molecule of the assembly located between the half-sites for binding to the transposase.
[0155] In the present invention, to form a single-stranded DNA molecule, the first molecule, the second molecule and the third molecule of the assembly are advantageously molecules made of deoxyribonucleotides.
[0156] Even more advantageously, the first, second and third molecules of the collection are hybrid DNA / RNA molecules, the "backbone" of the molecules being DNA, and the recognition sequences A and B of the nucleic acid molecule of interest and the central region of the third molecule being RNA. This is particularly advantageous when the sequence substitutions that enable the present invention are made directly to RNA molecules.
[0157] [FIG. 9]-A shows the interaction between a target molecule and the assembly according to the present invention.
[0158] Advantageously, the invention relates to a complex as described above, in which the first molecule or the second molecule, or both molecules, are coupled, in particular via a modified nucleotide, to an enzyme intended to promote the displacement of the molecule of interest. helicases (enzymes capable of opening the two strands of double-stranded molecules that are supercoiled or even associated with proteins such as histones), topoisomerases (enzymes that affect the topological structure of DNA by generating temporary breaks), a ligase that allows the formation of a phosphodiester linkage between the phosphate 5' end of a nucleotide and the OH 3' end of another nucleotide, Polymerases that synthesize nucleic acid molecules in a 5' to 3' direction, specifically from a free 3' OH initiation site, by copying the antiparallel complementary strand according to the Watson-Crick model. It could be.
[0159] It is also possible to combine two or more of the enzymes in order to have all the enzyme material necessary to enable the sequence substitutions contemplated within the scope of the present invention.
[0160] In a particular embodiment, the first molecule of the collection can contain one or more modified nucleotides in its 5' portion, more precisely between the first transposase binding site and region A. Similarly, the third molecule of the collection can contain one or more modified nucleotides in its 3' portion, more precisely between the first transposase binding site and region A.
[0161] This modified nucleotide is in particular modified by grafting the substituted carbon chain with a protein tag or a molecule allowing a specific interaction, such as streptavidin or biotin.
[0162] This modification then allows the specific binding of enzymes that may be useful for promoting tagmentation and sequence replacement to the first molecule of this assembly. For example, it is particularly advantageous to have a streptavidin graft, which allows the grafting of a biotinylated helicase (or conversely, a helicase grafted to streptavidin and biotinylated nucleotides), which is useful for dissociating the two strands of a double-stranded molecule. It is also possible to attempt grafting with a biotinylated ligase (or a ligase coupled to streptavidin) to connect the recombinant strand to the 3' side.
[0163] The oligonucleotide may be associated with a first or second molecule of the collection such that the oligonucleotide pairs with a predetermined region of the first or second molecule, which oligonucleotide is then advantageously coupled to a grafting molecule as described herein above.
[0164] The collection is suitable for single-stranded sequence replacement, for example, replacement of sequences on single-stranded DNA molecules or RNA.
[0165] To replace the two strands of a double-stranded molecule, it is useful to use two of the aforementioned assemblies, and these two assemblies are the first collection comprises first, second and third molecules as defined herein above; the second collection comprises a fourth, fifth, and sixth sequence; the result: the third and sixth sequences each comprise a strand of a replacement sequence, and these sequences are complementary according to Watson-Crick pairing; the first and fourth sequences comprise sequences that recognize the 5' portion of the region to be replaced and the 3' portion of the complementary strand of the region to be replaced, respectively, in region A and region A' (region A' is the fourth molecule's equivalent of region A of the first molecule); The second and fifth sequences comprise sequences that recognize the 3' portion of the region to be replaced and the 5' portion of the complementary strand of the region to be replaced, respectively, in the B region and B' region (the B' region is the fifth molecule's equivalent of the B region of the second molecule). It's like this.
[0166] Next, this double complex or this double assembly is shown in [Figure 10]-A.
[0167] Advantageously, the invention relates to an assembly as described above, comprising one of the 300 pairs of first and third molecules of Table 5 below.
[0168] [Table 14(01)]
[0169] [Table 14(02)]
[0170] [Table 14(03)]
[0171] [Table 14(04)]
[0172] [Table 14(05)]
[0173] [Table 14(06)]
[0174] [Table 14(07)]
[0175] In yet another aspect, the present invention relates to a kit or set comprising at least one vector allowing the expression of a recombinase and the first, second and third molecules of the collection as defined herein above.
[0176] The kit consists of either one container containing the three molecules of the assembly, or several separate containers.
[0177] The transposase included in the kit is a transposase that can recognize a binding site formed by the interaction between the first molecule and the third molecule of the complex, and / or the second molecule and the third molecule of the complex.
[0178] If the complex is such that it contains two binding sites for two different transposases, the kit contains two different vectors, each encoding one of the transposases.
[0179] The kit may also further comprise other enzymes (eg, a helicase or a ligase), as discussed earlier herein.
[0180] The kit according to the invention advantageously comprises what is necessary to form the two assemblies defined herein above.
[0181] The present invention also relates to the use of a collection as defined herein above for nucleic acid manipulation, in particular for replacing a target sequence with a sequence of interest, with the proviso that said use does not include a method for modifying the germline genetic identity of a human being and / or is not a method for treating the human or animal body by surgery or therapy.
[0182] As previously described herein, the aforementioned collections utilize the tagmentation properties of transposases to allow for specific targeting of target regions and replacement of them with sequences of interest.
[0183] The collections according to the invention are particularly advantageous for carrying out targeted genetic modifications (and in particular replacement of genes or non-coding sequences) in different organisms, for example in plants to obtain genetically modified plants, or in animals to create models of human diseases, or in living therapeutic tools to allow drugs to be tested.
[0184] Furthermore, the present invention relates to the aforementioned collection for use as a drug, more particularly as a gene therapy agent, in particular for treating or preventing diseases associated with nucleic acid modifications.
[0185] Since the above collection allows for the substitution of specific sequences, it is possible to design first and second molecules that specifically recognize a mutated gene by modification such as substitution, deletion or insertion, and to design a third molecule that contains the reference wild-type sequence of the mutated gene.
[0186] Also, in the presence of the appropriate transposase, it is possible to replace the mutant gene sequence with the wild-type sequence, thus treating individuals suffering from a disease caused by a genetic mutation.
[0187] The present invention also relates to use of the above-mentioned complex for producing a medicament for treating or preventing a disease associated with nucleic acid modification.
[0188] In some embodiments, these drugs are not intended to alter human germline identity, nor are they intended to cause undue suffering in animals.
[0189] The present invention further provides a method, particularly an in vitro method, for replacing a target region of a nucleic acid molecule with a region of interest of another nucleic acid molecule so as to obtain a hybrid nucleic acid molecule, comprising: contacting a population as defined herein with a nucleic acid comprising a target region, The collection comprises a first molecule, wherein the A sequence of the first molecule comprises a complementary sequence to a region immediately 5' of the target region; the B sequence of the first molecule comprises a complementary sequence to the region immediately 3' of the target region; contacting the third molecule, wherein the third molecule comprises a region of interest located between the complementary sequence of the second sequence for recognizing the transposase of the first molecule and the complementary sequence of the first sequence for recognizing the transposase of the second molecule, to obtain a displacement complex; placing the displacement complex in the presence of a transposase that recognizes the double-stranded binding site of the transposase contained in the assembly to obtain a recombination complex; recombining the combination complex to obtain a hybrid nucleic acid molecule containing a region of interest in place of the target region.
[0190] Note that during replacement, i.e., tagmentation by the transposase, the 3' end of the replacement fragment is joined by a phosphodiester bond to the 5' end of the sequence immediately following the replaced sequence. This ligation occurs during tagmentation.
[0191] Conversely, the 3' end of the sequence immediately preceding the displaced sequence and the 5' end of the displaced strand are not ligated to each other, and to terminate the displacement, it is necessary to use a ligase to join these two ends.
[0192] Note also that at the substitution site, the substitution molecule is flanked 3' and 5' by transposase recognition sequences.
[0193] The mechanism of displacement of the single-stranded target is shown in [Figure 9].
[0194] In the first step, an assembly consisting of first, second, and third molecules, each containing a complementary region of the 5' portion of the target sequence, a complementary region of the 3' portion of the target sequence, and a sequence of interest, is organized in space so that the G / C-rich complementary regions of the first molecules pair with each other and the G / C-rich complementary regions of the second molecules pair together (Figure 9)-A).
[0195] In the presence of a target sequence, the complementary regions of the 5' and 3' portions of the target sequence contained in the first and second molecules, respectively, pair with their respective complementary regions to form a tetramolecular complex containing the target molecule and three molecules of the assembly. The interaction (or pairing) between the first and second molecules and their targets has the effect of disrupting the interaction of the G / C-rich regions, so that the transposase binding site, where the transposase dimer is located, is positioned spatially close to the target molecule. The transposase can then exert its activity and cleave both the target molecule and the molecule of interest (between the two binding sites), as shown in Figure 9-B.
[0196] Therefore, tagmentation is The 3' end of the third molecule binds to the free 5' end generated by the transposase in the region 3' of the target sequence. This binding is a covalent bond by ligation of the two ends, The 5' end of the third molecule is placed opposite the free 3' end generated by the transposase in the 5' region of the target sequence. The transposase creates a 9-base pair deletion in the 5' region of the target sequence (shown by the dotted circle in Figure 9-C). Ligation is required to recover the entire molecule with the desired sequence located between the 5' and 3' regions of the original target sequence.
[0197] The final molecule resulting from tagmentation consists, in the 5' to 3' direction, of a portion of the 5' region of the initial target sequence with a partial deletion of 9 base pairs, followed by a second transposase binding sequence, followed by the sequence of interest itself, followed by the first transposase binding sequence, and finally a portion of the 5' region of the initial target sequence.
[0198] In yet another aspect, the present invention relates to a method, in particular an in vitro or ex vivo method, for editing the genome of a cell, which allows replacing a specific fragment of double-stranded DNA of the genome of a cell with another double-stranded DNA fragment of interest, to obtain a recombinant hybrid genome comprising the other double-stranded DNA fragment of interest in place of the specific fragment of double-stranded DNA, comprising: Providing a first assembly as defined hereinabove, the A sequence of the first molecule comprises the complementary sequence of the 5' flanking region of the particular fragment; the B sequence of the second molecule comprises the complementary sequence of the 3' flanking region of the particular fragment; preparing a first collection, in which a third molecule includes a sequence of one of the strands of a specific fragment between a region located between a complementary sequence of a second sequence for recognizing a transposase of the first molecule and a complementary sequence of the first sequence for recognizing a transposase of the second molecule; Optionally, providing a second assembly as defined herein above, the A sequence of the complementary region of the first molecule of the second collection comprises the complementary sequence of the 5' flanking region of the specific fragment; the B sequence of the complementary region of the second molecule of the second collection comprises the complementary sequence of the 3' flanking region of the specific fragment; the third molecule of the second collection comprises a sequence of a complementary strand of a specific fragment contained in the third sequence of the first collection between a region located between a complementary sequence of the second sequence for recognizing the transposase of the first molecule of the second collection and a complementary sequence of the first sequence for recognizing the transposase of the second molecule of the second collection; the complementary sequence of the 5' flanking region of the specific fragment contained in the A region of the first molecule of the first collection is at most 95% complementary to the complementary sequence of the 5' flanking region of the specific fragment contained in the A region of the first molecule of the second collection; providing a second collection, in which the complementary sequences of the 3' flanking regions of the specific fragments contained in the B regions of the first molecules of the first collection are at most 95% complementary to the complementary sequences of the 3' flanking regions of the specific fragments contained in the B regions of the first molecules of the second collection; To obtain the recombinant complex, contacting the cell with a recombination complex to obtain a cell ready to be edited; expressing a transposase in the cell ready to be edited to obtain an edited cell; selecting the edited cells, wherein the genome of the edited cells comprises, instead of the particular double-stranded DNA fragment, another double-stranded DNA fragment of interest; The present invention relates to a method, provided that the method is not for altering the germline genetic identity of a human being and is not for treating the human or animal body by surgery or therapy.
[0199] Advantageously, the present invention provides a method for editing the genome of a cell, which allows replacing a specific fragment of double-stranded DNA of the genome of a cell with another double-stranded DNA fragment of interest, to obtain a recombinant hybrid genome comprising the other double-stranded DNA fragment of interest in place of the specific fragment of double-stranded DNA, Providing a first assembly as defined hereinabove, the A sequence of the first molecule comprises the complementary sequence of the 5' flanking region of the particular fragment; the B sequence of the second molecule comprises the complementary sequence of the 3' flanking region of the particular fragment; preparing a first collection, in which a third molecule includes a sequence of one of the strands of a specific fragment between a region located between a complementary sequence of a second sequence for recognizing a transposase of the first molecule and a complementary sequence of the first sequence for recognizing a transposase of the second molecule; Providing a second assembly as defined hereinabove, the A sequence of the complementary region of the first molecule of the second collection comprises the complementary sequence of the 5' flanking region of the specific fragment; the B sequence of the complementary region of the second molecule of the second collection comprises the complementary sequence of the 3' flanking region of the specific fragment; the third molecule of the second collection comprises a sequence of a complementary strand of a specific fragment contained in the third sequence of the first collection between a region located between a complementary sequence of the second sequence for recognizing the transposase of the first molecule of the second collection and a complementary sequence of the first sequence for recognizing the transposase of the second molecule of the second collection; the complementary sequence of the 5' flanking region of the specific fragment contained in the A region of the first molecule of the first collection is at most 95% complementary to the complementary sequence of the 5' flanking region of the specific fragment contained in the A region of the first molecule of the second collection; providing a second collection, in which the complementary sequences of the 3' flanking regions of the specific fragments contained in the B regions of the first molecules of the first collection are at most 95% complementary to the complementary sequences of the 3' flanking regions of the specific fragments contained in the B regions of the first molecules of the second collection; To obtain the recombinant complex, contacting the cell with a recombination complex to obtain a cell ready to be edited; expressing a transposase in the cell ready to be edited to obtain an edited cell; selecting the edited cells, wherein the genome of the edited cells comprises, instead of the particular double-stranded DNA fragment, another double-stranded DNA fragment of interest; The present invention relates to a method, provided that the method is not for altering the germline genetic identity of a human being and is not for treating the human or animal body by surgery or therapy.
[0200] Insofar as this genome editing method is based on the methods for replacing single molecules and using assemblies described previously herein, all previously described features and all variants apply mutatis mutandis herein.
[0201] It should be noted that in this aspect of the invention, the first molecule of the considered assembly (first or second) is always located 5' of the third molecule of the considered assembly.
[0202] Also, a first molecule of a first population is positioned "on top" of a second molecule of a second population, and vice versa, and the 5' to 3' orientation allows for precise positioning of the sequence of interest relative to the target sequence.
[0203] Genome editing consists of modifying the genome of a cell with high precision. It is possible to inactivate genes, introduce targeted mutations, correct specific mutations, or insert new genes. This genetic engineering technique involves nucleases (herein referred to as transposases) that are able to cleave nucleic acids at phosphodiester bonds.
[0204] As explained above in the context of the present invention, the first and second molecules of the assembly according to the present invention make it possible to position the assembly adjacent to the target region in order to replace it with a sequence of interest contained in a third sequence.
[0205] In the context of editing double-stranded molecules, in order to simultaneously replace the two strands of the target molecule and have two ensembles according to the invention, it is necessary for each ensemble to specifically target one of the two strands of the target molecule.
[0206] Note that the complementary sequence contained in the A sequence of a first molecule in the first collection is not complementary to the complementary sequence contained in the A sequence of a first molecule in the second collection. Similarly, the B sequence of a second molecule in the first collection is not complementary to the complementary sequence contained in the B sequence of a second molecule in the second collection.
[0207] Furthermore, it is advantageous that the complementary sequence contained in the A sequence of a first molecule of the first collection is offset with respect to the complementary sequence contained in the B sequence of a second molecule of the first collection (these two molecules are one "on" the other). This means that the area for recognizing these complementary sequences is at most 95%. In other words, if the complementary sequences contained in the A sequence and the B sequence contain 20 nucleotides, the complementary sequences will only be complementary to each other over a maximum of 19 nucleotides.
[0208] However, it is preferred that these complementary regions have as little complementarity as possible to each other, or even that they are not complementary to each other.
[0209] To clarify these observations, if the 5' portion of the target sequence is considered to comprise 40 nucleotides, then it is appropriate that the complementary region contained in the A sequence of the first molecule of the first collection is complementary to the first 20 nucleotides, whereas the complementary region contained in the B sequence of the second molecule of the second collection is complementary to the last 20 nucleotides of the complementary sequence of the 5' portion of the target sequence.
[0210] This offset avoids any incorrect orientation of the target sequence even if the sequence orientation does not allow it.
[0211] The sequence and results of the tagmentation process are shown in Figure 10.
[0212] To obtain the final molecule, and due to the 9 base pair deletion at 5' after tagmentation, it is necessary for the cell to mobilize repair systems to fill the hole, in particular by using a DNA polymerase that copies the complementary strand whose 3' end was joined during tagmentation. [Brief explanation of the drawings]
[0213] The invention will be better understood on reading the following examples and figures. [Figure 1] [Figure 1] is a schematic diagram of a linear first nucleic acid molecule. The rectangle with an arrow indicates a site for transposase binding, and the rectangle with a diagonal line indicates a G / C-rich region. [Figure 2] [Figure 2] is a schematic diagram of a first nucleic acid molecule in a structured form. The explanation is the same as in [Figure 1]. [Figure 3] [Figure 3] is a schematic diagram of a complex according to the present invention. The explanation is the same as in [Figure 1]. [Figure 4] [Figure 4] is a schematic diagram of two versions of the complex according to the invention. The description is the same as in [Figure 1]. The different options are indicated by dotted lines. [Figure 5] [Fig. 5] is a schematic diagram of a first embodiment of the composite according to the present invention. The explanation is the same as in [Fig. 1]. [Figure 6] [Figure 6] is a schematic diagram of a second (6A) and third (6B) embodiment of a complex according to the invention, or in which the 3' region of the first sequence contains two transposase half-sites. Description is the same as in [Figure 1]. [Figure 7] [Figure 7] is a schematic diagram of a third embodiment of the composite according to the present invention. The explanation is the same as in [Figure 1]. [Figure 8] [Fig. 8] is a schematic diagram of an assembly according to the present invention. The explanation is the same as in [Fig. 1]. [Figure 9] [Figure 9] is a schematic diagram of a particular sequence of steps for replacing a target sequence in a single-stranded molecule with a sequence of interest. A: Represents unpaired target molecules and the assembly. B: Represents paired target molecules and the assembly. The transposase is shown as a dotted line, and the cut is shown as a pair of scissors. Note that the cut on the third molecule of the assembly occurs between the two transposase binding sites on either side of the molecule (two cuts). C: Represents the molecule resulting from tagmentation. The 9-base deletion 5' of the replaced region is shown as a dotted line. [Figure 10] [Figure 10] is a schematic diagram of a particular sequence of steps for replacing a target sequence in a double-stranded molecule with a sequence of interest that is itself double-stranded. A: Represents unpaired target molecules and the assembly. B: Represents paired target molecules and the assembly. The transposase is shown as a dotted line, and the cut is shown as a pair of scissors. Note that the cut on the third molecule of the assembly occurs between the two transposase binding sites on either side of the molecule (two cuts). C: Represents the molecule resulting from tagmentation. The 9-base deletion 5' of the replaced region is shown as a dotted line. [Figure 11] [Figure 11] shows an agarose gel demonstrating tagmentation according to the present invention. a) Agarose gel containing three sample groups: a test of different transposase complexes (negative control, Tn5 WT, Tn5 Me, and Tn5 DREAMT, i.e., according to the present invention) with mCherry-CD9 plasmid, PCR amplification of HEK 293T total mRNA, and PCR amplification of HEK 293T total mRNA transfected with mCherry-CD9 plasmid (with PCR + / - control). b) Agarose gel of different transposase mixes (see a) concentrated 10-fold on the mCherry-CD9 plasmid. [Figure 12][Figure 12] shows the results of Sanger sequencing of the positive amplified band (gel [Figure 10]a) for the Tn5 DREAMT mix. The obtained sequence (SEQ ID NO: 429) and the corresponding chromatogram are presented. ** indicates the GFP insertion zone. The mCherry sequence adjacent to the GFP sequence is underlined. [Figure 13] [Figure 13] shows an agarose gel containing three sample groups: testing different transposase complexes (negative control, Tn5 WT, Tn5 Me, and Tn5 DREAMT) with the mCherry-CD9 plasmid, PCR amplification of HEK 293T total cDNA, and PCR amplification of HEK 293T total cDNA transfected with the mCherry-CD9 plasmid (with PCR + / - controls). [Figure 14] [Figure 14] shows the results of Sanger sequencing of the positive amplified band (gel [Figure 13]) for the Tn5 DREAMT mix. The obtained sequence (SEQ ID NO: 430) and the corresponding chromatogram are presented. ** indicates the GFP insertion zone. It is flanked by CD9 sequences. [Figure 15] [Figure 15] shows an agarose gel containing three sample groups: testing different transposase complexes (negative control, Tn5 WT, Tn5 Me, and Tn5 DREAMT) with the mCherry-CD9 plasmid, PCR amplification of HEK 293T total cDNA, and PCR amplification of HEK 293T total cDNA transfected with the mCherry-CD9 plasmid (with PCR + / - controls). [Figure 16] [Figure 16] shows the results of Sanger sequencing of the positive amplified band (gel [Figure 15]) for the Tn5 DREAMT mix. The obtained sequence (SEQ ID NO: 431) and the corresponding chromatogram are presented. ** indicates the GFP insertion zone. It is flanked by CD9 sequences. [Figure 17]Figure 17 shows the testing of the "newly designed" UVRD-mSA / Tn5 / DREAMT complex on the mCherry-CD9 plasmid. HEK 293T cells were transfected with the mCherry-CD9 plasmid and then examined with the DREAMT technology to detect a visible color change (from red to green, as shown in the figure). [Figure 18] [Figure 18] shows transfection of a HEK 293T cell line stably expressing mCherry-CD9 using the technique of the present invention, followed by detection of a visible color change (from red to green, as shown in the figure). [Figure 19] [Figure 19] shows the Sanger sequencing results of the amplified GFP fragment from a cDNA library derived from the mCherry-CD9+ cell line of HEK 293T cells (cells used in D18) transfected with the "newly designed" DREAMT technology (SEQ ID NO: 432). GFP substitutions are shown adjacently. The sequence of SEQ ID NO: 433 shows the theoretical GFP sequence (adjacent portion). [Example]
[0214] Example 1 - Implementation of the invention in vitro. The purpose of this example is to demonstrate that the collection according to the invention allows for the easy and specific substitution of the sequence of a single-stranded nucleic acid molecule (RNA or cDNA) with a sequence of interest.
[0215] In this example, the goal is to replace the sequence of mCherry-CD9 with the sequence of GFP. -Preparing recombinant assemblies targeting mCherry-CD9.
[0216] Prepare two separate tubes containing: For tube 1 (10 μL): + 10 μM Loop A oligonucleotide (first molecule) containing a sequence for recognizing mCherry in its A region; + 10 μM of oligonucleotide A (third molecule), containing a restriction half-site 3′ + reverse BSPEi Frag oligo (10 μM) For tube 2 (10 μL), + 10 μM Loop B oligonucleotide (second molecule) containing a sequence in the B region for recognizing the CD9 sequence; + 10 μM of oligonucleotide A (third molecule), containing a restriction half-site 3′ + NdeI Frag oligo (10 μM).
[0217] The two tubes are then heated to 95°C for 5 minutes and then left at room temperature for 1 hour.
[0218] Once at room temperature, the contents of the tube are placed in the presence of either the enzyme NdeI or the enzyme BSPEi to allow digestion of the restriction sites.
[0219] The digestion product is then purified using a PCR purification kit (elution volume 20 μL).
[0220] In parallel, a sequence encoding GFP is amplified to contain NdeI and BSPEi restriction sites at the 5' and 3' ends. The PCR product is then digested with the two restriction enzymes and purified using a PCR purification kit (elution volume 20 μL).
[0221] The contents of tubes 1 and 2 are then combined with the (amplified and digested) GFP fragment in the presence of T4 phage ligase (4 μL, i.e., 100 U), together with 5 μL of T4 buffer (10x) and 1 μL of ddH2O, to a total reaction volume of 50 μL.
[0222] The solution is then left at room temperature for 1 hour, and then gel-purified to isolate the largest GFP fragment, which is larger than 1 kilobase (kb) and contains two single-stranded molecules at its 5' and 3' ends (the first and second molecules of the assembly), forming the recombinant complex.
[0223] The recombination complex is then ready to be incubated with the transposase dimer.
[0224] In controls, 10 μM each of MeA and MeB oligonucleotides were used in combination with the MERev fragment (previously described in S. Picelli et al. Genome Res 2014): Tn5MErev, 5'-[phos]CTGTCTCTTATACACATCT-3' (SEQ ID NO: 11) Tn5ME-A (Illumina FC-121-1030), 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 12) and Tn5ME-B (Illumina FC-121-1031); 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 13) These oligos were prepared using the following primers: 1) 1000 nucleotides (A, B, C, D, E ...
[0225] Generate Tn5 transposase using the recommendations of the following publication: S. Picelli et al. Genome res 2014.
[0226] The two previous preparations, the recombination complex and the positive control oligo, were then mixed with a preparation of Tn5 transposase via the following protocol: 0.125 volumes of a 10 μM solution of the recombinant complex or control oligo (or only dialysis buffer (see S. Picelli et al. Genome, 2014) for so-called WT wild-type Tn5) + 0.4 volumes of 100% glycerol solution + 0.24 volumes of dialysis buffer (see S. Picelli et al. Genome res 2014) + 0.36 volumes of Tn5 solution.
[0227] The reaction is then brought from 45°C to 37°C over 30 minutes, decreasing by 1°C every 4 minutes in a thermocycler.
[0228] At the same time, several targets were created to test the technology according to the present invention.
[0229] The mCherry-CD9 plasmid was used as a negative control, and the mCherry-CD9 vector was used to generate mRNA and cDNA libraries from HEK 293T cells pre-transfected with Lipofectamine 3000.
[0230] 48 hours after transfection and visual inspection of HEK cells (red membrane due to mCherry expression), the total population of mRNA is extracted and reverse transcribed into cDNA.
[0231] These mRNA or cDNA libraries are then used for our "in vitro" testing of the DREAMT technology.
[0232] To characterize and test the technology according to the invention, different mRNA / cDNA and plasmid targets were contacted with: Tn5 WT dimer: no oligos, containing Tn5 Me:Me oligo (positive control), and Tn5 complex: includes the assembly according to the present invention.
[0233] Three types of targets were tested: mCherry-CD9 plasmid and mRNA or cDNA libraries with or without transfection of mCherry-CD9.
[0234] The double-stranded or single-stranded DNA / RNA solutions were contacted with different solutions of activated Tn5 by the following reaction: 1x or 10x concentration of Tn5 WT, Tn5 Me, Tn5 complex solution or the same volume of ddH20 + 500ng of mCherry-CD9 plasmid or 2µL of cDNA or mRNA + 5x tagmentation buffer (100mM HEPES, 50mM MgCl2, 40% PEG 3500) +ddH20 (appropriate amount 20 μL).
[0235] Each tagmentation reaction is carried out in a thermocycler at 55°C for 7 minutes, then 0.5 μL of proteinase K solution (20 μg / μL) is added before a second incubation temperature of 55°C for 7 minutes to inactivate the transposase.
[0236] Therefore, as shown in Figure 11, the 3' end of GFP was ligated to the mCherry target sequence.
[0237] On the agarose gel in Figure 11, the mCherry-CD9 plasmid (Davidson Lab) was contacted with two concentrations of Tn5 solution (1x and 10x). As expected, at the higher concentration (10x), the Tn5 WT mix and Tn5 Me positive control underwent tagmentation, as indicated by degradation of the plasmid (smear, disappearance of the plasmid band).
[0238] Furthermore, as expected, the assembly alone was unable to perform tagmentation (plasmid degradation), as indicated by the preservation of DNA bands on an agarose gel (Figure 11) [because it does not contain a molecule (helicase) for separating the two strands of the plasmid's double-stranded DNA to release the Tn5 molecule and pair with the mCherry-CD9 target, enabling tagmentation]. This data allows us to conclude that by limiting Tn5 complex dimers, they can exert their tagmentation activity only by pairing the recombining strand with the target single strand (here, the mCherry-CD9 mRNA fragment).
[0239] The same concentration of Tn5 solution (1x) was used with 2 μL of mRNA library (derived from normal or mCherry-CD9 transfected HEK 293T cells), followed by PCR using mCherry forward primer and GFP forward primer (see Table 3).
[0240] Table 3: Oligonucleotides used in PCR reactions and plasmid construction
[0241] [Table 15]
[0242] As expected, no signal was detected for the mRNA library obtained from cells without transfection of the mCherry-CD9 vector. Furthermore, no amplification was detected for the negative controls (mRNA fragment alone, mRNA + Tn5 WT, mRNA + Tn5 Me) or the mRNA library samples transfected with mCherry-CD9. Nevertheless, a clear band at the correct molecular weight of approximately 1 kb was detected for the sample corresponding to the replacement of mCherry-CD9 with the GFP sequence (Tn5 complex) except for the samples from this same batch (Figure 11). This band highlights the cleavage of the mCherry 3' end (antisense mRNA) with the insertion of the GFP 3' end (sense strand), or, in other words, the ligation of the mCherry mRNA antisense strand to the GFP sense strand in the 3' to 5' direction.
[0243] The amplified product was sequenced, and the resulting sequence is shown in Figure 12. As expected, amplification of mCherry to the 3' end of GFP was detected, along with the detection of the two sequences connected by ligation. The chromatogram from this sequence clearly shows a 10-bp precipitation zone (tagmentation) from the target sequence (Figure 12).
[0244] The same analysis was repeated, this time with a modified design for ligation of the 3' end of the cDNA sense strand to CD9 (Fig. 12). Therefore, the same results were obtained using an amplification gel showing the expected band of approximately 1 kb for a cDNA library contacted with the assembly of the present invention (Tn5 complex) using mCherry-CD9 transfection (Fig. 13). The resulting fragments were then sequenced (Fig. 14). As expected, this sequencing allowed us to identify the sequence of the CD9 sense strand in the precipitated cDNA (ligated at the 3' end of the GFP sense sequence, Fig. 14), which contained multiple sequences of the GFP 5' sense strand located 5 bp from the target sequence.
[0245] Example 2 - Implementation of the invention in cellulo. Encouraged by these promising results, the design was modified to obtain a version compatible with direct transfection of an assembly consisting of a Tn5 plasmid and a plasmid of the UVRD bacterial helicase fused to a double-strand opening protein, monomeric streptavidin (UVRD-mSA plasmid), so that the latter is in situ bound to one of the assembly molecules via a pre-biotinylated oligo. To do this, a new version of the assembly was developed to facilitate easier placement and allow for near-100% control of the design (reducing Tn5 dimers not restricted by the Beacon sarcophagus system and thus reducing the associated "off-target" potential). Therefore, for this "newly designed assembly," the same type of protocol as before was performed, with certain key modifications. Regarding the actual design, the looped nDREAMT A, B, C, and D oligos (one per Tn5 dimer and one per target for tagmentation of four target regions—perfect double-stranded displacement) are longer, containing three sequences for transposase binding. This allows for rapid and controlled pairing with the simple addition of a final sequence for transposase binding linked to a GFP fragment to replace the mCherry-CD9 target fragment (nDREAMT A, B, C, and D oligos). Furthermore, other nDREAMT Bio A and B oligos (biotinylated oligos) can be added to allow in cell assembly of the assemblies with the UVRD-mSA fusion protein.
[0246] The following oligomers were mixed in two separate tubes 1 and 2: nDREAMT A and C loop (10 μM) + nDREAMT A and C oligos (10 μM) + nDREAMT Bio A oligo (10 μM) for tube 1 (10 μL), and nDREAMT B and D loop (10 μM) + nDREAMT B and D oligos (10 μM) + nDREAMT Bio B oligo (10 μM) for tube 2 (10 μL). The two tubes were then heated to 95°C for 5 minutes, then left at room temperature for 1 hour, and then each tube was digested with the corresponding restriction enzyme and purified using a PCR purification kit (elution volume 20 μL).
[0247] At the same time, the amplified sequence of GFP containing the restriction enzyme NdeI at the 5' and the restriction enzyme BSPEi at the 3' flanking ends (an oligo with a restriction site structure digested after PCR) was amplified, then digested with the corresponding restriction enzymes and purified by a PCR purification kit eluting in a volume of 20 μL.
[0248] The contents of tubes 1 and 2 are then mixed with the amplified and digested GFP neo fragment, a solution of T4 ligase (4 μL or 100 U), 5 μL of T4 buffer (10x), and 1 μL of ddH2O to a total reaction volume of 50 μL. The solution is left at room temperature for 1 hour, and then gel-purified to isolate the largest GFP fragment, which is larger than 1 kb and contains four sarcophagi A, B, C, and D at its 5' and 3' ends. The GFP-replaced fragment, with its four sarcophagi positioned at its 5' and 3' ends, is then ready to accept a transposase dimer (Figure 15).
[0249] Before testing this system at the cellular level, we tested the new design through the same series of experiments as before, using a GFP 3' ligation (sense strand) to the CD9 end of the mCherry-CD9 target cDNA (sense strand). Thus, after amplification with GFP forward and CD9 reverse PCR primers, we obtained the same type of results as before, with the expected band of approximately 0.5 kb (Figure 15). (All remaining negative controls were undetectable.) Sanger sequencing of this band revealed, as expected, detection of the CD9 fragment in the 5' to 3' direction, within two bases of the original target sequence (Figure 16), due to the ligation of the GFP multiplex sequence at its 3' end (Figure 16).
[0250] Following this positive validation of the new design, we performed in situ testing in HEK 293T cells expressing mCherry-CD9. To this end, we created two intracellular protein expression plasmids: one for the Tn5 transposase and the other for the UVRD-mSA protein. For the Tn5 plasmid, the Tn5 sequence was amplified by PCR using the PTXB1-Tn5 plasmid (S. Picelli et al., Genome Res., 2014). Note that a stop codon was substituted for the intein tail tag (S. Picelli et al., Genome Res., 2014), which allows for purification of the Tn5 transposase by binding to chitin beads and self-cleavage, leaving no tag element on the Tn5 protein. The entire construct was cloned into the mCherry-CD9 fragment of the mCherry-CD9 plasmid (Davidson Lab) using BMT1 and EcoR1 restriction enzymes, ultimately yielding the CMV-Tn5 promoter plasmid. The mSA fragment was PCR amplified using the pRSET-mSA plasmid (Sheldon Park Lab) (Table 1), which added a 5' BMT1 restriction site and a 3' EcoR1 restriction site. The first cloning step was performed in the mCherry-CD9 plasmid using the BM1 and EcoR1 enzymes, replacing the mCherry-CD9 fragment with the mSA fragment to obtain the CMV-mSA promoter plasmid. For the UVRD fragment, an Escherichia coli (E. coli) bacterial cDNA library was generated (DH5α). PCR amplification of the UVRD cDNA was performed using the correct oligonucleotides (Table 1). The 5' side contains the usual elements, such as the BMT1 and Kozac sequences for cloning, but also an NLS import sequence for importing the complete complex + Tn5 + UVRD-mSA complex into the cell nucleus. A flexible segment (GGGSx3-type polyglycine tail) was added at the 3' end, and one end of the mSA 5' sequence was added to the precise restriction enzyme PMiI contained in the mSA 5' sequence. A second cloning step was then performed using BMT1 and PMiI restriction enzymes to allow the addition of the UVRD sequence at the 5' end, resulting in the final plasmid of the CMV-UVRD-mSA promoter type.
[0251] HEK 293T cells were transfected with the mCherry-CD9 plasmid using Lipofectamine 3000. After waiting 24 hours and confirming the red signal on the cell membrane by confocal microscopy, the cells were transfected with three compounds: the CMV-Tn5 and CMV-UVRD-mSA plasmids and the newly designed complex. This complete complex allows double-stranded displacement of the sequence located between four targets (a quadruplet of targets: two mCherry-side and two CD9-side) with the CMV-GFP sequence. Therefore, a reduction in the red membrane signal or even its replacement by a green cytoplasmic signal was expected from this experiment.
[0252] As expected and presented in FIG. 17, green cells were detected 5 days after the second transfection by the technique according to the present invention.
[0253] These cells were isolated by cell cytometry sorting, and knowing that the CMV-GFP fragment had replaced the mCherry-CD9 portion of the CMV-mCherry-CD9 plasmid, this new CMV-GFP plasmid should contain an antibiotic selection sequence to generate the NeoR / KanR cell line. Six days after transfection, these still-green cells were placed in medium containing 2 mg / ml G418, which was changed daily for 14 days, and then maintained at a concentration of 0.5 mg / ml for 3 months.
[0254] As expected, a new GFP+ green cell line was obtained 19 days after transfection, which was highly stable despite repeated freezing and thawing cycles (see Figure 17).
[0255] To more reliably confirm this result in situ, we generated stable HEK 293T red membrane cell lines by transfection with the CMV-mCherry-CD9 plasmid and treatment with 2 mg / ml G418 as before. After 2 weeks of treatment, we obtained HEK 293T red membrane cell lines.
[0256] This cell line was transfected as before using the techniques described above. Three days after transfection, the first green cells began to appear (see Figure 18). Fourteen days after transfection, a specific population of green cells was identified and isolated for further analysis (see Figure 18).
[0257] After isolation of these cells, reverse transcription was performed using these cells 18 days after transfection to obtain a cDNA library.
[0258] The GFP sequence was then amplified by PCR, and the amplified fragment was then sequenced according to the Sanger method.
[0259] As expected, after analysis by comparison with a sequence database (NCBI BLAST), a close similarity was found with the theoretical GFP fragment replacing the mCherry-CD9 sequence (Figure 19). Additionally, after approximately one month, the resulting green cells remained stable with a strong green signal (Figure 18). These results confirm in cell culture the effectiveness of the present technology, which allows replacement of the mCherry-CD9 target sequence with the GFP replacement sequence, as indicated by a color change from red in the membrane to green in the cytoplasm. This color change allows verification of the nuclear import of the newly designed Tn5 UVRD-mSA complex (containing an NLS on the N-terminal sequence of UVRD), double-stranded DNA breakage performed by two 5' and 3' helicases, pairing at the four desired target points (two on the mCherry side and two on the CD9 side), and replacement of the mCherry-CD9 sequence with the CMV-GFP fragment by the present technology.
Claims
1. A first single-stranded nucleic acid molecule comprising or consisting essentially of an A sequence that enables insertion of a complementary sequence of a nucleic acid of interest, wherein the A sequence binds to a first A / T-rich, particularly T-rich sequence that is 40 to 60 nucleotides in length at the 5'-end and binds to a second A / T-rich, particularly T-rich sequence that is 40 to 60 nucleotides in length at the 3'-end, and the first and second A / T-rich, particularly T-rich sequences each contain first and second domains of 6 to 12 G / C-rich nucleotides, the sequence of the first domain is complementary to the sequence of the second domain, the first and second domains are positioned 15 to 52 nucleotides from the A sequence, and the first molecule contains, at its 5'-end, a first sequence oriented 5' to 3' for recognizing transposase and, at its 3'-end, at least one second sequence for recognizing the transposase, a first single-stranded nucleic acid molecule; a complex comprising the first single-stranded nucleic acid molecule and a second single-stranded nucleic acid molecule that contains or consists essentially of, at its 5'-end, at least one complementary sequence of the second sequence for recognizing the transposase; wherein the first and second single-stranded nucleic acid molecules are paired according to base complementarity defined by Watson-Crick so as to define two double-stranded binding sites of the transposase.
2. The complex according to claim 1, wherein the A sequence contains a complementary sequence of the nucleic acid of interest.
3. The first molecule contains, at its 5'-end, a first sequence oriented 5' to 3' for recognizing transposase and, at its 3'-end, a second sequence oriented 5' to 3' for recognizing the transposase, and the second molecule contains, at its 5'-end, a first complementary sequence of the first sequence for recognizing the transposase and, following that, a second complementary sequence of the second sequence for recognizing the transposase. The complex according to claim 1.
4. The first molecule contains, at its 5'-end, a first sequence oriented 5' to 3' for recognizing transposase, at its 3'-end, a second sequence for recognizing the transposase, and following that, a first complementary sequence of the first sequence for recognizing the transposase. The complex according to claim 1, wherein the second molecule comprises, at its 5' end, a complementary sequence of the second sequence for recognizing the transposase.
5. The complex according to claim 1, wherein the transposase is a bacterial transposase, particularly a transposase selected from Tn5, Tn9, Tn10 or Tc1 / mariner.
6. The complex according to claim 1, wherein the first molecule is coupled to the enzyme, particularly via a modified nucleotide.
7. The complex according to claim 1, wherein the first molecule comprises one of the following sequences: SEQ ID NO: 436, SEQ ID NO: 440, SEQ ID NO: 443, SEQ ID NO: 444, SEQ ID NO: 448, SEQ ID NO: 449, SEQ ID NO: 450, SEQ ID NO: 454, SEQ ID NO: 455, SEQ ID NO: 456, SEQ ID NO: 457, SEQ ID NO: 458, SEQ ID NO: 462, SEQ ID NO: 463, SEQ ID NO: 464, SEQ ID NO: 465, SEQ ID NO: 466 and SEQ ID NO:
641.
8. The complex according to claim 1, wherein the complex comprises a pair of first and second molecules, and the first and second molecules comprise the sequences defined in Table 2.
9. The complex according to claim 1, comprising one of the pairs of first and second molecules defined in rows 1 to 166, rows 171 to 172, rows 177 to 178, rows 183 to 184 and rows 201 to 288 of Table 4.
10. A first single-stranded nucleic acid molecule that contains or consists essentially of an A sequence that enables insertion of a complementary sequence of a nucleic acid of interest, or that contains a complementary sequence of a nucleic acid of interest, wherein the complementary sequence binds to a first T-rich sequence 40 to 60 nucleotides in length at the 5' end and to a second T-rich sequence 40 to 60 nucleotides in length at the 3' end, the first and second T-rich sequences each contain first and second domains of 6 to 12 G / C-rich nucleotides, the sequence of the first domain is complementary to the sequence of the second domain, the first and second domains are positioned 15 to 52 nucleotides from the A sequence, and the first molecule comprises, at its 5' end, at least one first sequence oriented 5' to 3' for recognizing the transposase and, at its 3' end, a second sequence for recognizing the transposase. A B sequence that contains or consists essentially of a B sequence that allows insertion of a complementary sequence of a nucleic acid of interest, or a second single-stranded nucleic acid molecule that contains a complementary sequence of the nucleic acid of interest, wherein the complementary B sequence binds to a third T-rich sequence that is 40 to 60 nucleotides long at the 5' end and binds to a fourth T-rich sequence that is 40 to 60 nucleotides long at the 3' end, and the third and fourth T-rich sequences each contain third and fourth domains of 6 to 12 G / C-rich nucleotides, the sequence of the third domain is complementary to the sequence of the fourth domain, the third and fourth domains are positioned 15 to 52 nucleotides from the B sequence, and the second molecule contains, at its 5' end, at least one first sequence oriented 5' to 3' for recognizing the transposase, and at its 3' end, the second sequence for recognizing the transposase. A second single-stranded nucleic acid molecule, wherein the B sequence is a complementary sequence of the nucleic acid of interest, the A sequence is positioned 5' of the target region of the nucleic acid of interest, and the B sequence is positioned 3' of the target region of the nucleic acid of interest. A third single-stranded molecule, at its 5' portion, at least one complementary sequence of the second sequence for recognizing the transposase of the first molecule; at its 3' portion, at least one complementary sequence of the first sequence for recognizing the transposase of the second molecule, a third single-stranded molecule. An aggregate comprising a complementary sequence of the second sequence for recognizing the transposase of the first molecule and a region located between the complementary sequence of the first sequence for recognizing the transposase of the second molecule that allows insertion of a single-stranded replacement nucleic acid molecule. An aggregate, wherein the first and third single-stranded nucleic acid molecules are paired according to the base complementarity defined by Watson-Crick so as to define two double-stranded binding sites of the transposase, and the second and third single-stranded nucleic acid molecules are paired according to the base complementarity defined by Watson-Crick so as to define two double-stranded binding sites of the transposase. [
11. ] The assembly according to claim 10, comprising one of the pairs of the first and third molecules defined in lines 1 to 166, 171 to 172, 177 to 178, 183 to 184 and 201 to 288 of Table 5.
12. A kit comprising at least one vector enabling the expression of recombinase and the first, second and third molecules of the assembly according to claim 10.
13. Use of the assembly according to claim 10 for manipulating nucleic acids, in particular for replacing a target sequence with a target sequence, provided that it does not include a method for modifying the human germline genetic identity and is not a method for treating the human or animal body by surgery or therapy.
14. A method for in vitro replacing a target region of a nucleic acid molecule with a target region of another nucleic acid molecule so as to obtain a hybrid nucleic acid molecule, comprising contacting the assembly according to claim 10 with the nucleic acid comprising the target region, wherein the assembly the A sequence of the first molecule comprises a complementary sequence of the region immediately 5' of the target region, the B sequence of the first molecule comprises a complementary sequence of the region immediately 3' of the target region, the third molecule comprises the target region in a region located between the complementary sequence of the second sequence for recognizing the transposase of the first molecule and the complementary sequence of the first sequence for recognizing the transposase of the second molecule to obtain a substitution complex, contacting, placing the substitution complex in the presence of a transposase that recognizes the double-stranded binding site of the transposase comprised in the assembly to obtain a recombination complex, and recombining the combination complex to obtain the hybrid nucleic acid molecule comprising the target region instead of the target region, provided that it is not a method for modifying the human germline genetic identity and is not a method for treating the human or animal body by surgery or therapy.
15. An in vitro or ex vivo method for editing the genome of a cell, which enables obtaining a recombinant hybrid genome containing a specific other double-stranded DNA fragment of interest instead of a specific fragment of the double-stranded DNA of the genome of the cell, by replacing the specific fragment of the double-stranded DNA with the other double-stranded DNA fragment of interest, Preparing the first assembly according to claim 10, wherein the A sequence of the first molecule includes a complementary sequence of an adjacent region 5' to the specific fragment, wherein the B sequence of the second molecule includes a complementary sequence of an adjacent region 3' to the specific fragment, Preparing the first assembly according to claim 10, wherein the third molecule includes a sequence of one strand of the specific fragment between regions located between the complementary sequence of the second sequence for recognizing the transposase of the first molecule and the complementary sequence of the first sequence for recognizing the transposase of the second molecule, Optionally, preparing the second assembly according to claim 10, wherein the A sequence of the complementary region of the first molecule of the second assembly includes a complementary sequence of an adjacent region 5' to the specific fragment, wherein the B sequence of the complementary region of the second molecule of the second assembly includes a complementary sequence of an adjacent region 3' to the specific fragment, Preparing the second assembly according to claim 10, wherein the third molecule of the second assembly includes a sequence of the complementary strand of the specific fragment included in the third sequence of the first assembly between regions located between the complementary sequence of the second sequence for recognizing the transposase of the first molecule of the second assembly and the complementary sequence of the first sequence for recognizing the transposase of the second molecule of the second assembly, wherein the complementary sequence of the adjacent region 5' of the specific fragment included in the A region of the first molecule of the first assembly is at most 95% complementary to the complementary sequence of the adjacent region 5' of the specific fragment included in the A region of the first molecule of the second assembly, Preparing the second assembly according to claim 10, wherein the complementary sequence of the adjacent region 3' of the specific fragment included in the B region of the first molecule of the first assembly is at most 95% complementary to the complementary sequence of the adjacent region 3' of the specific fragment included in the B region of the first molecule of the second assembly, To obtain a recombinant complex, contacting the cells with the recombinant complex to obtain cells ready to be edited; expressing the transposase in the cells ready to be edited to obtain edited cells; selecting the edited cells, wherein the genome of the edited cells comprises the other double-stranded DNA fragment of interest instead of the specific double-stranded DNA fragment; a method, provided that it is not a method for modifying the human germline genetic identity and is not a method for treating the human or animal body by surgery or therapy.