A dna prime editing system, recombinant vector, kit and method for inversion of a target dna sequence
Patent Information
- Application Number
- CN202510016521.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-01-06
AI Technical Summary
然而,这些方法对于哺乳动物细胞中的倒位仍然效率低下,并且需要耗时的连续转染来在倒位之前安装重组位点
[0070]与现有技术相比,本发明的有益之处在于:本发明提供了基于相同技术构思的四套DNA先导编辑系统。这四套DNA先导编辑系统都是通过特殊设计的先导编辑向导RNA(pegRNA),在目标DNA序列的同一条链的两端生成互补的配对片段,通过配对片段的互补作用实现目标DNA的倒位编辑。
Smart Images

Figure CN119899837B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing technology, and in particular to DNA lead editing systems, recombinant vectors, kits, and methods for inverting target DNA sequences. Background Technology
[0002] CRISPR-based genome editing tools have revolutionized our ability to manipulate genome sequences with unprecedented precision. CRISPR-Cas nucleases induce site-specific double-strand breaks (DSBs), enabling targeted genome modifications through repair mechanisms such as non-homologous end joining (NHEJ) or homology-directed repair (HDR). These pathways facilitate small insertions / deletions (indels) or precise modifications using donor templates.
[0003] In recent years, catalytically impaired Cas nucleases have been engineered to form base editors (BEs) in conjunction with deaminases or to develop lead editors (PEs) in conjunction with reverse transcriptases (RTs). Base editors allow for efficient base substitutions at specific genomic sites without relying on HDRs or donor DNA. Lead editors, on the other hand, use lead editing guide RNAs (pegRNAs) to guide precise sequence modifications via RT templates encoded within the pegRNA. By combining PEs with integrases, the range of precise DNA modifications can be extended from a few base pairs (bp) to approximately 40 kb, a range encompassing the average length of a single gene.
[0004] A significant portion of pathogenic mutations originate from structural variations in the genome. While these variations can be induced using the Cas9 nuclease in combination with a pair of guide RNAs (gRNAs), the primary editing outcome is typically a small insertion / deletion, with structural variations being relatively minor byproducts. Despite successes in precise editing at single gene loci, developing tools capable of precisely and efficiently designing large-scale structural variations in the mammalian genome remains a critical unmet need.
[0005] Unlike deletions and duplications, inversions typically do not alter genome copy number but can significantly impact gene expression and genome integrity. These rearrangements are crucial for understanding human evolution and are associated with a variety of genetic diseases, such as hemophilia A, mucopolysaccharidosis type II (Hunter's syndrome), and Emory-Dreyfus muscular dystrophy, which are often caused by inversions between reverse homologous sequences. Furthermore, inversions can lead to the formation of oncogene fusions, such as the ALK-EML4 fusion associated with lung cancer, highlighting their potential impact on gene function. Despite their association with genetic diseases and various cancers, efficient and convenient genetic tools for mimicking chromosomal inversions in cells and animals remain lacking. While the binding of Cas9 nucleases to gRNA pairs can induce inversions, these events are infrequent and are often masked by insertions / deletions. Recent methods, such as binding PE to integrase and using two pairs of pegRNAs to install recombination sites, have been shown to invert 40 kb fragments. However, these methods remain inefficient for inversions in mammalian cells and require time-consuming serial transfections to install recombination sites before inversion. Given the crucial role of inversions in genomic structure and disease, there is an urgent need for a more effective and programmable tool to manipulate and study these structural variations. Summary of the Invention
[0006] To address the shortcomings of the existing technologies, this invention provides a DNA inversion target DNA sequence editing system, recombinant vector, kit, and method. Utilizing a guide RNA (pegRNA), it enables more precise and efficient inversion editing of the target DNA sequence. Specifically, this is achieved through the following techniques.
[0007] In a first aspect, the present invention provides a first DNA leader editing system for inverting a target DNA sequence, the DNA leader editing system comprising a Cas protein and a reverse transcriptase, pegRNA-A and pegRNA-B;
[0008] The pegRNA-A includes a first guide RNA, a first primer binding site, and a first reverse transcription template, and the pegRNA-B includes a second guide RNA, a second primer binding site, and a second reverse transcription template;
[0009] The first reverse transcription template includes a first paired fragment, and the second reverse transcription template includes a second paired fragment; the sequences of the first paired fragment and the second paired fragment are complementary.
[0010] The pegRNA-A and pegRNA-B respectively guide the Cas protein to cleave the target fragment at the restriction sites on both sides of the same strand of the double-stranded DNA.
[0011] Furthermore, the sequence lengths of the first and second reverse transcription templates are 0 to 2,000 nucleotides; preferably 15 to 500 nucleotides.
[0012] Furthermore, the sequence lengths of the first and second paired fragments are 0 to 1,000 nucleotides; preferably 3 to 200 nucleotides; more preferably 30 to 100 nucleotides.
[0013] Furthermore, the first reverse transcription template also contains a first non-complementary template sequence, and the second reverse transcription template also contains a second non-complementary template sequence; the first non-complementary template sequence and the second non-complementary template sequence are not complementary to each other.
[0014] Furthermore, the lengths of the first non-complementary template sequence and the second non-complementary template sequence are 1 to 2,000 nucleotides; preferably 1 to 1,000 nucleotides; more preferably 1 to 500 nucleotides.
[0015] Further, the restriction sites located on both sides of the same strand of the double-stranded DNA are spaced 500 to 100,000,000 base pairs apart; preferably 10,000 to 30,000,000 base pairs apart.
[0016] Furthermore, the pegRNA-A and / or pegRNA-B further include a tail sequence; the tail sequence forms a hairpin or loop with itself, or the tail sequence includes a poly(A), poly(U), or poly(C) sequence, or the tail sequence includes an RNA-binding domain.
[0017] The DNA leader editing system provided by this invention introduces the Cas protein, reverse transcriptase RT, and two pegRNAs (A and B) into two complementary 3' flaps on the same strand of double-stranded DNA. P rime editing-based I nversionwith E Enhanced performance version 1, PIEv1). When these two complementary single-stranded DNA sequences combine, an inversion occurs under DNA repair. Figure 1 This method is referred to as the PIEv1 method / design in this invention.
[0018] A second aspect of the present invention provides a second DNA lead editing system for inverting target DNA sequences, characterized in that the DNA lead editing system comprises a Cas protein and reverse transcriptase, pegRNA-A, pegRNA-B, pegRNA-C, and pegRNA-D;
[0019] The pegRNA-A includes a first guide RNA and a first reverse transcription template; the pegRNA-B includes a second guide RNA and a second reverse transcription template; the pegRNA-C includes a third guide RNA, a third primer binding site, and a third reverse transcription template; and the pegRNA-D includes a fourth guide RNA, a fourth primer binding site, and a fourth reverse transcription template.
[0020] The first reverse transcription template includes a first paired fragment, the second reverse transcription template includes a second paired fragment, the third reverse transcription template includes a third paired fragment, and the fourth reverse transcription template includes a fourth paired fragment; the sequences of the first paired fragment and the second paired fragment are complementary, and the sequences of the third paired fragment and the fourth paired fragment are complementary.
[0021] The pegRNA-A and pegRNA-B respectively guide the Cas protein to cleave the target fragment at the enzyme sites on both sides of one strand of the double-stranded DNA; the pegRNA-C and pegRNA-D respectively guide the Cas protein to cleave the target fragment at the enzyme sites on both sides of the other strand of the double-stranded DNA.
[0022] The protospacer sequence adjacent to the cleavage site of the Cas protein cleavage site guided by pegRNA-A and the protospacer sequence adjacent to the cleavage site of the Cas protein cleavage site guided by pegRNA-C are positioned close to (PAM-in) or far from (PAM-out) relative to the Spacer sequence.
[0023] In this invention, the PAM-out form is as follows: Figure 7 As shown. Figure 7 In the diagram, the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-A is located to the right of the spacer sequence (ProtospacerA); the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-C is located to the left of the spacer sequence (ProtospacerB), and the two PAMs of pegRNA-A and pegRNA-C are positioned far apart from each other relative to the two spacer sequences. Similarly, the two PAMs of pegRNA-B and pegRNA-D are also positioned far apart from each other relative to the two spacer sequences.
[0024] In this invention, the PAM-in form is as follows: Figure 9 As shown. Figure 9In the diagram, the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-A is located to the right of the spacer sequence (ProtospacerA); the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-C is located to the left of the spacer sequence (ProtospacerB), and the two PAMs of pegRNA-A and pegRNA-C are positioned close to each other relative to the two spacer sequences. Similarly, the two PAMs of pegRNA-B and pegRNA-D are also positioned close to each other relative to the two spacer sequences.
[0025] Further, the sequence length between A and C is 2 to 10,000; preferably 500 to 2,000 nucleotides.
[0026] Further, the sequence length between B and D is 2 to 10,000; preferably 500 to 2,000 nucleotides.
[0027] Furthermore, the restriction sites located on both sides of the same strand of the double-stranded DNA are spaced 500 to 100,000,000 base pairs apart.
[0028] Furthermore, the distance between them is 10,000 to 30,000,000 base pairs.
[0029] To develop precise inversion editing tools, we attempted to introduce a second pair of complementary 3' single-stranded DNA molecules onto the other strand of the double-stranded DNA to ensure the precision of the interface at the other end. For two adjacent / nearby pegRNAs (… Figure 7 and 8 The pegRNAs pegRNA-A and pegRNA-C, and pegRNA-B and pegRNA-D, exist in two forms: PAM-out and PAM-in. Here, "adjacent / near" refers to two pegRNAs located on two DNA strands and adjacent to each other. For example, pegRNA-A and pegRNA-C are defined as a group of two "adjacent / near" pegRNAs, and pegRNA-B and pegRNA-D are defined as a group of two "adjacent / near" pegRNAs. Therefore, for the PAM-out and PAM-in forms, the second DNA lead editing system provided by this invention is actually divided into two more specific DNA lead editing systems, which are named PIEv2 method / design and PIEv3 method / design in this invention, respectively.
[0030] In the PAM-out case, the PIEv2 method / design targets four sites on the genomic double-stranded DNA, synthesizing two pairs of complementary 3' single-stranded DNA using reverse transcriptase RT. When these two pairs of complementary 3' single-stranded DNA bind together, inversions occur through DNA repair. After the inversion editing is complete, the four original target sites remain.
[0031] In the case of PAM-in, the PIEv3 method / design also targets four sites on the double-stranded DNA of the genome, synthesizing two pairs of complementary 3' single-stranded DNA via RT. When these two pairs of complementary 3' single-stranded DNA bind together, an inversion occurs under DNA repair, simultaneously deleting the sequence between the two nicks on the same side to prevent secondary editing.
[0032] Compared to the PIEv1 method / design, the PIEv2 method / design and the PIEv3 method / design can achieve more precise inversion editing.
[0033] A third aspect of the present invention also provides a DNA leader editing system for inverting target DNA sequences, the DNA leader editing system comprising a Cas protein and a reverse transcriptase, pegRNA-E and pegRNA-F;
[0034] The pegRNA-E includes a fifth guide RNA, a fifth primer binding site, and a fifth reverse transcription template; the pegRNA-F includes a sixth guide RNA, a sixth primer binding site, and a fifth reverse transcription template; the fifth and sixth guide RNAs target the target fragments respectively.
[0035] The fifth reverse transcription template includes a fifth paired fragment, and the sixth reverse transcription template includes a sixth paired fragment; the fifth paired fragment is complementary to the gene fragment sequence adjacent to the motif near the 3' end of the original spacer sequence of the cleavage site guided by the sixth guide RNA to cut the Cas protein, and the sixth paired fragment is complementary to the gene fragment sequence adjacent to the motif near the 3' end of the original spacer sequence of the cleavage site guided by the fifth guide RNA to cut the Cas protein.
[0036] The pegRNA-E and pegRNA-F respectively guide the Cas protein to cleave the target fragment at the restriction sites on both sides of the two opposite strands of the double-stranded DNA; the original spacer sequence adjacent motif of the restriction site cleaved by the Cas protein guided by pegRNA-E and the original spacer sequence adjacent motif of the restriction site cleaved by the Cas protein guided by pegRNA-F are set far away from the corresponding spacer sequence.
[0037] Further, the restriction sites located on both sides of the two opposite strands of the double-stranded DNA are spaced 500 to 50,000,000 base pairs apart; preferably 1,000 to 500,000 base pairs apart.
[0038] It is understood that in the various DNA lead editing systems provided by this invention, the first guide RNA, second guide RNA, third guide RNA, fourth guide RNA, fifth guide RNA, and sixth guide RNA each target their respective target fragments. The designations "first," "second," "third," "fourth," "fifth," and "sixth" are merely for distinction and do not imply any specific technical limitation. For different CRISPR / Cas gene editing systems, the guide RNA can be a specific sgRNA or crRNA.
[0039] For example, in the CRISPR / Cas9 gene editing system, the guide RNA is sgRNA; in the CRISPR / Cas12 or CRISPR / Cas13 gene editing systems, the guide RNA is crRNA.
[0040] A fourth aspect of the present invention provides a substance, which is any one of the following:
[0041] (1) A nucleic acid composition comprising the coding gene of any of the DNA lead editing systems provided above in this invention;
[0042] (2) A recombinant vector, wherein the vector comprises the coding gene of any of the DNA lead editing systems provided above in this invention;
[0043] (3) A recombinant microorganism, wherein the recombinant microorganism encodes the gene of any of the DNA lead editing systems provided above in this invention;
[0044] (4) A kit comprising any of the DNA lead editing systems provided above in this invention, or comprising the recombinant vector.
[0045] Furthermore, when the recombinant vector includes the coding gene of the DNA lead editing system provided in the second aspect of the present invention, the pegRNA-A and pegRNA-D are recombined on the same recombinant vector, and the pegRNA-B and pegRNA-C are recombined on the same recombinant vector.
[0046] In practical applications using the PIEv1, PIEv2, or PIEv3 methods / designs, pegRNAs are typically cloned into different vectors. When using this vector construction method, uneven plasmid distribution may occur after transfection. This is especially true when using the PIEv3 method / design, as it may lead to PIEv1-like editing results, producing inaccurate inversions.
[0047] To mitigate the aforementioned problems, for the PIEv3 method / design, this invention clones two pegRNAs (pegRNA-A and pegRNA-D) that are distal and target different DNA strands into the same vector, and clones two other pegRNAs (pegRNA-B and pegRNA-C) into the same vector. No editing occurs when only one vector is present in the cell; PIEv3 editing only occurs when both vectors are present simultaneously.
[0048] In a fifth aspect, the present invention provides the application of any of the above-described DNA leader editing systems in inverted target DNA sequences.
[0049] A sixth aspect of the present invention provides a method for reversing a target DNA sequence, comprising the following steps:
[0050] The target DNA sequence and any one of the DNA lead editing systems provided in the first aspect of the present invention are added to the reaction system. The reverse transcriptase uses the first reverse transcription template and the second reverse transcription template to reverse transcribe and generate two single-stranded DNA sequences, respectively. The two single-stranded DNA sequences complement each other to form a double-stranded region. Based on the DNA repair mechanism, the target DNA sequence is inverted, and the complementary DNA sequence of the first paired fragment is inserted at the interface near the first paired fragment.
[0051] Alternatively, it may include the following steps:
[0052] The target DNA sequence and the DNA lead editing system described in the second aspect of the present invention are added to the reaction system; the reverse transcriptase uses the first reverse transcription template, the second reverse transcription template, the third reverse transcription template and the fourth reverse transcription template to reverse transcribe and generate four single-stranded DNA sequences, which are complementary and combine to form two double-stranded regions;
[0053] When the two PAMs (protospacer motifs of the Cas protein cleavage sites) of pegRNA-A and pegRNA-C are far apart from each other relative to the two spacer sequences (the same applies to pegRNA-B and pegRNA-D), i.e., in PAM-out form, the target DNA sequence is inverted based on the DNA repair mechanism. The complementary DNA sequence of the first paired fragment is inserted near the interface of the first paired fragment, and the sequence copy between A and C is added in situ; the complementary DNA sequence of the third paired fragment is inserted near the interface of the third paired fragment, and the sequence copy between B and D is added in situ.
[0054] When the two PAMs (protospacer motifs of the Cas protein cleavage site) of pegRNA-A and pegRNA-C are close to each other relative to the two Spacer sequences (the same applies to pegRNA-B and pegRNA-D), forming a PAM-in configuration, the target DNA sequence is inverted based on the DNA repair mechanism. At the interface near the first paired fragment, a copy of the sequence between A and C is deleted and a complementary DNA sequence of the first paired fragment is inserted; at the interface near the third paired fragment, a copy of the sequence between B and D is deleted and a complementary DNA sequence of the third paired fragment is inserted.
[0055] Alternatively, it may include the following steps:
[0056] The target DNA sequence and the DNA lead editing system provided in the third aspect are added to the reaction system; the reverse transcriptase uses the fifth and sixth reverse transcription templates to reverse transcribe and generate two single-stranded DNA sequences, respectively, and the two single-stranded DNA sequences bind to their corresponding complementary genomic regions to form double-stranded regions; the target DNA sequence is inverted based on the DNA repair mechanism.
[0057] Furthermore, in the above-mentioned method for reversing the target DNA sequence, the Cas protein is a Cas9 protein containing an inactivated HNH domain, or a Cas9 protein containing an inactivated RuvC domain.
[0058] Furthermore, in the above-mentioned method for inverting the target DNA sequence, the Cas protein is SpCas9, FnCas9, St1Cas9, St3Cas9, NmCas9, SaCas9, VQR SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-NG, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, or AtCas9.
[0059] Furthermore, in the above method for inverting the target DNA sequence, the Cas protein is the Cas12 protein.
[0060] Furthermore, in the above-described method for inverting the target DNA sequence, the Cas12 protein is Cas12a, Cas12b, Cas12f, or Cas12i.
[0061] Specifically, optionally, in the above-mentioned method for inverting the target DNA sequence, the Cas12 protein is one or more of AsCpf1, LbCpf1, FnCpf1, SsCpf1, PcCpf1, BpCpf1, CmtCpf1, LiCpf1, PmCpf1, Pb3310Cpf1, Pb4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCas12b, EbCas12b, and LsCas12b.
[0062] Optionally, the target DNA sequence is located inside a cell.
[0063] Optionally, the cell may be a dividing cell or a non-dividing cell.
[0064] Optionally, the target DNA sequence is a telomere or a fragment thereof.
[0065] Optionally, the above-described method for reversing the target DNA sequence can be performed in vitro or in vivo.
[0066] Optionally, each pegRNA includes a first guide RNA or a second guide RNA, a first primer binding site or a second primer binding site, and a first reverse transcription template or a second reverse transcription template, arranged in a "5' to 3'" or "3' to 5'" orientation.
[0067] Optionally, the reverse transcriptase is M-MLV reverse transcriptase or a reverse transcriptase that can function under physiological conditions.
[0068] Optionally, the Cas protein and reverse transcriptase may be provided indirectly as nucleic acid fragments encoding the corresponding proteins, or directly as proteins.
[0069] Alternatively, each pegRNA may be provided indirectly as recombinant DNA encoding the pegRNA, or directly as an RNA molecule.
[0070] Compared with existing technologies, the advantages of this invention are: this invention provides four DNA lead editing systems based on the same technological concept. These four DNA lead editing systems all use specially designed lead editing guide RNAs (pegRNAs) to generate complementary paired fragments at both ends of the same strand of the target DNA sequence, achieving inversion editing of the target DNA through the complementary effect of the paired fragments.
[0071] The four DNA lead editing systems provided by this invention can target any DNA sequence requiring inversion editing, and are not limited to a specific species or class of species. In practical applications, as long as the technical concept of this invention is used to construct the four DNA lead editing systems for the corresponding species, good gene inversion editing results can be obtained. Attached Figure Description
[0072] Figure 1 This is a schematic diagram illustrating the process of achieving target DNA inversion using the PIEv1 method of Example 1. The PE2-pegRNA complex targets two sites on the same strand of the genomic double-stranded DNA. Two complementary 3' single-stranded DNAs were synthesized via RT using an RTT that is not homologous to the genomic sequence. When the two complementary 3' single-stranded DNAs anneal together, an inversion occurs under DNA repair (yellow portion).
[0073] Figure 2 This is a TAE gel electrophoresis image of the inversion results detected in Example 1; confirming 10 kb, 100 kb, and 1 Mb inversions on chromosome 6.
[0074] Figure 3 The results of first-generation sequencing (Sanger sequencing) of the inverted interface sequence in Example 1 determined the left and right interface sequences of the 1Mb inversion on chromosome 6. The left side was accurate, while the right side had random insertions and deletions.
[0075] Figure 4 and Figure 5 This is the verification result of the integrity of the inversion that occurred in Example 1. Figure 4 Sanger sequencing validated the wild-type genome in HEK293T cells. Figure 5 The 836 bp inversion at the VEGFA site was verified.
[0076] Figure 6 The inversion editing efficiency of PIEv1 in HEK293T and N2a cells in Example 1 is shown.
[0077] Figure 7 This is a schematic diagram illustrating the process of achieving target DNA inversion using the PIEv2 method of Example 2. The PE2-pegRNA complex targets four sites on the genomic double-stranded DNA. Two pairs of complementary 3' single-stranded DNA were synthesized via RT. When these two pairs of complementary 3' single-stranded DNA bind together, inversion occurs under DNA repair (yellow portion). After the inversion editing is completed, the four target sites still exist.
[0078] Figure 8This is a schematic diagram of the secondary editing that occurred after the PIEv2 inversion editing in Example 2; the amplification editing (AE) that occurred on the genomic DNA edited by PIEv2. After the inversion editing was completed, the four original target sites still existed, and then two pairs of complementary 3' single-stranded DNA were resynthesized on different strands of the double-stranded DNA, inducing AE.
[0079] Figure 9 This is a schematic diagram illustrating the process of achieving target DNA inversion using the PIEv3 method of Example 2. The PE2-pegRNA complex targets four sites on the genomic double-stranded DNA. Two pairs of complementary 3' single-stranded DNA were synthesized via RT. When these two pairs of complementary 3' single-stranded DNA bind together, inversion occurs under DNA repair, simultaneously deleting partial fragments at both ends (gray portions).
[0080] Figure 10 This is a TAE gel electrophoresis image of the inversion results detected using the PIEv2 method in Example 2; it identified 11 kb and 1 Mb inversions on chromosome 6.
[0081] Figure 11 The results are from the first-generation sequencing (Sanger sequencing) of the interface sequence after the inversion using the PIEv2 method in Example 2; confirming the amplification and editing of the 10kb inversion on chromosome 6.
[0082] Figure 12 This is a TAE gel electrophoresis image of the inversion results detected using the PIEv3 method in Example 2; it identifies 10 kb, 100 kb, and 1 Mb inversions on chromosome 6.
[0083] Figure 13 The results are from the first-generation sequencing (Sanger sequencing) of the interface sequence after inversion using the PIEv3 method in Example 2; the left and right interface sequences of the 1 Mb inversion on chromosome 6 were determined, with accurate results on both sides.
[0084] Figure 14 In Example 2, HEK293T ( Figure 14 a) and K562 ( Figure 14 b) The purity of the product from the interface sequences at both ends of the inversion of the 10 kb to 30 Mb region on chromosome 6 in cells, and HEK293T ( Figure 14 c) and K562 Figure 14 d) The editing efficiency of inversions in the 10 kb to 30 Mb region of chromosome 6 in cells was quantified by ddPCR.
[0085] Figure 15 and 16 This is a schematic diagram illustrating how PIEv3b reduces inaccurate editing in Example 3.
[0086] Figure 17 This is the verification result of the inversion editing efficiency of PIEv3b in different cell lines in Example 3.
[0087] Figure 18 In Example 3, the target inversion at the EGFP site is performed on a functional cell line. Figure 18 a is the design diagram of the GFP functional recovery cell line. Figure 18 b represents the editing efficiency of a 957 bp inversion in the EGFP reverse inactivation region, quantified by flow cytometry.
[0088] Figure 19 This is an illustration of the fluorescence in situ hybridization (FISH) used in Example 3 to determine intraarm inversion events on chromosome 6. The FISH diagram shows the 40 Mb inversion on chromosome 6. Column 1 from the left represents the centromere of chromosome 6 (green). In column 2, a red fluorescent probe labels approximately 198 kb of genomic sequence surrounding the ESR1 gene. Columns 1 and 2 are merged in column 3, with the nucleus shown in blue. Column 4 is a magnified view of column 3. Edit-1 and Edit-2 are different HEK293T single-cell clones that experienced the 40 Mb inversion.
[0089] Figure 20 The DNA leader editing system of Example 4 demonstrates the inversion editing effect on hemophilia A and acute myeloid leukemia (AML) sites in HEK293T cells.
[0090] Figure 21 This is a schematic diagram of a human chromosome with telomere structure created using IE technology in Example 4.
[0091] Figure 22 In Example 4, FISH was used to detect a 100 Mb inversion on Chr 2. Column 1 (from left to right) shows the centromere satellite DNA region of Chr 2 marked in green. Column 2 combines Column 1 with DAPI staining, with the cell nucleus shown in blue. Column 3 provides a magnified view of Column 2. Edit-1 and Edit-2 are different single-cell clones from IE-treated samples used for 100 Mb inversion on Chr 2 in HEK293T cells.
[0092] Figure 23To detect 30 Mb inversions on Chr 20 cells using FISH. Column 1 (from left to right) shows a ~730 kb genomic sequence near the MARPE1 gene labeled with a green fluorescent probe. Column 2 shows a ~379 kb genomic sequence near the YWHAB gene labeled with a red fluorescent probe. Columns 1 and 2 are merged in column 3, with cell nuclei stained blue. Column 4 is a magnified view of column 3, and column 5 shows only a magnified view of DAPI staining. Edit-1 and Edit-2 represent different single-cell clones from IE-treated samples, with 30 Mb inversions performed on Chr 20 cells in HEK293T cells.
[0093] Figure 24 This is a schematic diagram illustrating the design of the PIES-PAM-out method for achieving target DNA inversion in Example 5.
[0094] Figure 25 This is an efficiency graph showing the inversions of 100 kb and 1 Mb on chromosome 6 and 10 Mb on chromosome 2 in HEK293T(a) and K562(b) cells, respectively, in Example 5.
[0095] Figure 26 This is a purity statistic for the interfaces at both ends of the 100 kb and 1 Mb inverted chromosome 6 in HEK293T cells in Example 5. Detailed Implementation
[0096] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0097] It should be emphasized that the following specific embodiments are merely illustrative examples of gene inversion editing achieved by the four DNA lead editing systems provided by this invention; they are not limited to HEK293T cells, nor are the corresponding inversion editing effects only obtainable by using the following specific DNA lead editing systems.
[0098] The four DNA lead editing systems of this invention are not limited to the specific Cas protein, reverse transcriptase, or pegRNA sequences described in the following embodiments, nor are the target DNA sequences limited to any particular species or class of species. They can be used for any target DNA sequence requiring inversion editing, and are not limited to any particular species or class of species. In practical applications, for any species, as long as the four DNA lead editing systems constructed using the technical concept of this invention are employed, good gene inversion editing results can be obtained.
[0099] Example 1: Design and Feasibility Verification of PIEv1
[0100] To achieve precise and efficient inversion editing of the target DNA, this embodiment uses a complex of PE2 and two pegRNAs to perform inversion editing of the target DNA.
[0101] 1. Design of PIEv1
[0102] like Figure 1 As shown. Using a complex consisting of PE2 (nCas9-RT) and two pegRNAs (pegRNA A and pegRNA B), two sites on the same strand of the genomic double-stranded DNA are targeted. Using a reverse transcription template (RTT) that is not homologous to the genomic sequence, two complementary 3' single-stranded DNAs are introduced via reverse transcriptase (RT) (Inversion editing version 1, 3' flap).
[0103] When these two complementary 3' flaps bind together after annealing, an inversion occurs due to DNA repair. Figure 1 (The yellow part in the image). This targeted DNA inversion method was designed and named PIEv1.
[0104] 2. Feasibility verification of PIEv1
[0105] To verify the feasibility of PIEv1, this embodiment attempted to perform 10kb, 100kb, and 1Mb inversions on chromosome 6 in HEK293T cells. The specific method is as follows:
[0106] The PE2 plasmid sequence used is shown in SEQ ID NO.1.
[0107] The nucleotide sequence of the pegRNA plasmid used is shown in SEQ ID NO.2, where positions 3253-3500 are the promoter sequence, positions 3501-3520 are the spacer sequence, positions 3521-3596 are the sgRNA backbone sequence, positions 3597-3639 are the RTT and PBS sequences, and positions 3640-3684 are the linker and evopreQ1 motif sequences; where n is a degenerate base, representing any one of the bases A, T, C, and G.
[0108] For different inversion events, the spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 are different. See Table 1 below for details.
[0109] Table 1
[0110] HEK293T cells were co-transfected with pegRNA and PE2 plasmids, and the genome was harvested after 3 days of culture. Two pairs of primers were designed at the interfaces at both ends of the inverted regions to detect the occurrence of inversion events. Only the inverted sequences amplified the target band. The relevant primers are shown in Table 2 below.
[0111] Table 2
[0112] Test results as follows Figure 2 As shown in the TAE gel electrophoresis image, in the genome without any plasmid transfection (WT group), the F1 / F2 or R1 / R2 primer pairs could not amplify the bands; in the genome with inversions after editing, the F1 / F2 or R1 / R2 primer pairs could amplify the target bands, indicating that 10 kb, 100 kb and 1 Mb inversions were achieved on chromosome 6 of HEK293T cells.
[0113] like Figure 3 As shown, first-generation sequencing (Sanger sequencing) of the inverted interface revealed that the left interface contained flap sequence insertions, which ensured accuracy through flap complementation. However, the right interface exhibited significant indels, specifically random insertions and deletions. This indicates that the PIEv1 method involves some degree of imprecise editing.
[0114] To verify the integrity of the inversion, we performed an 836 bp inversion of the VEGFA gene on chromosome 6. After transfecting 293T cells, we collected the genome and constructed T / A clones from the amplified inverted bands, thus verifying the 836 bp inversion at the VEGFA site. Figure 4 First-generation sequencing (Sanger sequencing) verified the integrity of the inversion occurrence. Figure 5 The nucleotide sequence of the pegRNA plasmid used is also shown in SEQ ID NO.2. The spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differ for different inversion events. See Table 3 below for details.
[0115] Table 3
[0116] The amplification primers used, VEGFA-inv 836 bp-F, are: 5'-ctctttagccagagccgggg-3', as shown in SEQ ID NO. 23.
[0117] The amplification primers used, VEGFA-inv 836 bp-R, are: 5'-ggtggagggggtcggggct-3', as shown in SEQ ID NO. 24.
[0118] 3. Cell line validation of PIEv1
[0119] To determine the editing efficiency of PIEv1 in cell lines, we performed inversions on different regions of different chromosomes in human HEK293T cells and mouse N2a cells. The pegRNA used in human HEK293T cells is also shown in SEQ ID NO.2. For different inversion events, the spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differed. In mouse N2a cells, the spacer, RTT, and PBS sequences for different inversion events are shown in Table 4 below.
[0120] Table 4
[0121] After co-transfecting cells with pegRNA and PE2 plasmids and culturing for 3 days, the genome was harvested, and ddPCR was used to identify the inversion efficiency at the interface. The results showed that the efficiencies in HEK293T and N2a cells were 32.8-38.4% and 0.4-10.1%, respectively. Figure 6 (a and b) indicate that PIEv1 can achieve efficient inversion editing, but its efficiency is low at the chromosome level (above 1Mb).
[0122] Example 2: Design and feasibility verification of PIEv2 (PAM-in) and PIEv3 (PAM-out)
[0123] To achieve more precise inversion editing of the target DNA, this embodiment attempts to introduce a second pair of complementary 3' flaps on the other strand of the double-stranded DNA to ensure the precision of the interface at the other end. That is, in addition to the two pegRNAs (pegRNA A and pegRNA B) provided in Example 1, this embodiment also selects two additional pegRNAs (pegRNA C and pegRNA D).
[0124] like Figure 7 and 9 As shown, for two pegRNAs on the same side, PAM-in exists (e.g. Figure 9 (as shown) and PAM-out (as shown) Figure 7 (As shown) Two forms.
[0125] Figure 7In the diagram, the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-A is located to the right of the spacer sequence (ProtospacerA); the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-C is located to the left of the spacer sequence (ProtospacerB), and the two PAMs of pegRNA-A and pegRNA-C are far apart from each other relative to the two spacer sequences. Similarly, the two PAMs of pegRNA-B and pegRNA-D are also far apart from each other relative to the two spacer sequences.
[0126] Figure 9 In the diagram, the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-A is located to the right of the spacer sequence (ProtospacerA); the protospacer neighbor motif (PAM) of the Cas protein cleavage site corresponding to pegRNA-C is located to the left of the spacer sequence (ProtospacerB), and the two PAMs of pegRNA-A and pegRNA-C are close to each other relative to the two spacer sequences. Similarly, the two PAMs of pegRNA-B and pegRNA-D are also close to each other relative to the two spacer sequences.
[0127] The following designs corresponding target DNA inversion methods in two forms: PAM-in and PAM-out.
[0128] 1. Target DNA inversion method where both pegRNAs on the same side are in PAM-out form.
[0129] (1) Design of PIEv2
[0130] like Figure 7 As shown, this method involves designing two pegRNAs on the same side in the form of PAM-out, forming a complex consisting of nCas9-RT and four pegRNAs, targeting four sites on the double-stranded DNA of the genome, using a reverse transcription template (RTT) that is not homologous to the genome sequence, and introducing two pairs (four) of complementary 3' flaps via reverse transcriptase (RT).
[0131] When these two pairs of complementary 3' flaps are annealed and then joined together, the inversion is completed under the action of the repair mechanism. Figure 7 (The yellow part in the image). The precision of the interface at both ends is ensured by two pairs of complementary 3' flaps. This targeted DNA inversion method is designed and named PIEv2.
[0132] like Figure 8 As shown, amplification editing (AE) occurs on the genomic DNA edited by PIEv2. After the inversion editing is completed, the four original target sites (Protospacer-A, B, C, and D) still exist, and the positions of A and D have been interchanged. Subsequently, two pairs of complementary 3' single-stranded DNA were resynthesized on the two strands of the double-stranded DNA, and AE was performed on the inverted ends.
[0133] (2) Feasibility verification of PIEv2
[0134] Using the PIEv2 design, this embodiment performs inversions in 10kb and 1Mb regions on chromosome 6 of 293T cells. The specific method is as follows:
[0135] The nucleotide sequence of the pegRNA plasmid used is also shown in SEQ ID NO.2. The spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differ for different inversion events. See Table 5 below for details.
[0136] Table 5
[0137] ② Transfection and amplification of the target band
[0138] HEK293T cells were co-transfected with pegRNA and PE2 plasmids, and the genome was harvested after 3 days of culture. Two pairs of primers were designed at the interfaces at both ends of the inverted regions to detect the occurrence of inversion events. Only the inverted sequences amplified the target band. The specific primer sequences are shown in Table 6 below.
[0139] Table 6
[0140] Test results as follows Figure 10 As shown in the TAE gel electrophoresis image, ladder-like bands were found at both ends of the amplified and inverted sequence. First-generation sequencing was performed on the ladder-like bands, as... Figure 11 As shown, the amplification and editing at the 10kb interface on chromosome 6 were verified.
[0141] 2. Target DNA inversion method where both pegRNAs on the same side are in PAM-in form.
[0142] (1) Design of PIEv3
[0143] like Figure 9As shown, this method involves designing two pegRNAs on the same side in the form of PAM-in, a complex consisting of nCas9-RT and four pegRNAs, targeting four sites on the double-stranded DNA of the genome, using a reverse transcription template (RTT) that is not homologous to the genome sequence, and introducing two pairs (four) of complementary 3' flaps via reverse transcriptase (RT).
[0144] When these two pairs of complementary 3' flaps are annealed and then joined together, the inversion is completed under the action of the repair mechanism. Figure 9 (The yellow part in the image). Two pairs of complementary 3' flaps ensure the precision of the interfaces at both ends; this targeted DNA inversion method is named PIEv3. (As shown in the image) Figure 9 As shown in the gray area, the PIEv3 method was used to simultaneously delete the sequences between two nicks on the same side (A and C, B and D) to prevent secondary editing.
[0145] (2) Feasibility verification of PIEv3
[0146] Using the PIEv3 design, this embodiment performs inversions in 10kb, 100kb, and 1Mb regions on chromosome 6 of 293T cells. The specific method is as follows:
[0147] The nucleotide sequence of the pegRNA plasmid used is also shown in SEQ ID NO.2. The spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differ for different inversion events. See Table 7 below for details.
[0148] Table 7
[0149] HEK293T cells were co-transfected with pegRNA and PE2 plasmids, and the genome was harvested after 3 days of culture. Two pairs of primers were designed at the interfaces at both ends of the inverted regions to detect the occurrence of inversion events. Only the inverted sequences could amplify the target band. The primers used are shown in Table 8 below.
[0150] Table 8
[0151] Using the PIEv3 strategy, inversions were performed on 10kb, 100kb, 1Mb, and 30Mb regions of chromosome 6 in HEK293T and K562 cells. Figure 12 As shown, a single band was found at both ends of the amplified and inverted interface. Figure 13 As shown, the sequences at the left and right interfaces of the 1 Mb inversion on chromosome 6 were determined, and first-generation sequencing verified the accuracy of the interfaces at both ends.
[0152] like Figure 14 As shown in a and 14b, next-generation sequencing revealed that the purity of the interfaces at both ends of the 10 kb to 30 Mb region after inversion was very high in both cell lines, ranging from 95.9% to 99.0% and from 94.8% to 98.8% in the 293T and K562 cell lines, respectively.
[0153] To determine the efficiency of specific inversion editing, this embodiment uses ddPCR to precisely quantify the efficiency of the interfaces at both ends of the 10 kb to 30 Mb region where inversions occur. Figure 14 As shown in c and 14d, efficient editing was achieved in both cell lines. The inversion efficiency was 7.9-39.6% in HEK293T cells, and the efficiency of inverting 100kb and 30Mb in the K562 cell line reached 57.4% and 11.4%, respectively.
[0154] Example 3: Design and Feasibility Verification of PIEv3b
[0155] 1. Design of PIEv3b
[0156] In the PIEv3 method of Example 2, when four pegRNAs are cloned onto four separate plasmid vectors, uneven plasmid distribution may occur after cell transfection. In this case, as... Figure 15 As shown, this may result in inaccurate inversions in the PIEv1 editing results.
[0157] To reduce the number of PIEv1 editing results, such as Figure 16 As shown, in this embodiment, two pegRNAs that are distal and target different DNA strands are cloned into the same vector. No editing will occur when only one vector enters the cell. PIEv3 editing will only occur when both vectors enter the cell. We call this the PIEv3b method.
[0158] 2. Validation of inversion editing efficiency of PIEv3b in different cell lines
[0159] To verify the universality of PIEv3b, we performed inversion editing on 10 kb-1 Mb regions on chromosomes 6, 19, and 20 in human Huh-7, HAP1, HeLa, and hESCs cells; and performed inversion editing on 8 kb-50 Mb regions on chromosome 2 in mouse mESCs and haESCs cells.
[0160] For Chr 6-inv10 kb, Chr 6-inv100 kb, and Chr 6-inv1Mb, the pegRNAs used were the same as those used in Example 2 for Chr6 Inv 10kb-1-PIEv3, Chr6 Inv 100kb-PIEv3, and Chr6 Inv 1Mb-1-PIEv3, respectively. See Table 9 below for details.
[0161] Table 9
[0162]
[0163] In all cell lines, the genome was harvested 3 days after co-transfection with pegRNA and PE2, and the editing efficiency was quantified using ddPCR. In the human Huh-7 cell line, the inversion editing efficiency ranged from 6.0% to 21.3%. Figure 17 a); In the HAP1 cell line, the efficiency was 6.2-26.1% ( Figure 17 b); In HeLa cell lines, the efficiency was 18.4-46% ( Figure 17 c); In hESC cell lines, the efficiency was 8.8-62.0% ( Figure 17 d). In mouse mESCs cell lines, the inversion editing efficiency was 9.1-84.0% ( Figure 17 e); In haESCs cell lines, the editing efficiency was 2.5-47.3% ( Figure 17 f). In different cell lines, PIEv3b showed good inversion editing efficiency, demonstrating its universality.
[0164] Furthermore, compared to DSB-mediated inversions, PIEv3b produces higher precision inversion editing efficiency and fewer byproducts, while inversions constitute only a small portion of the edited products from DSB, with the majority being DSB-induced random indels. Compared to integrase-mediated inversions, PIEv3b exhibits higher editing efficiency, especially for genomic inversions larger than Mb, where integrase efficiency is very low or no editing occurs.
[0165] 3. PIEv3b Edit Completeness Verification
[0166] To further demonstrate the inversion editing capability of the PIEv3b method, we conducted a GFP gene function recovery assay using PIEv3b. The specific method is as follows:
[0167] (1) Construct a GFP expression plasmid with reverse inactivation and four pegRNA targeting sites (A, B, C and D) on both sides. The sequence of the GFP expression plasmid with reverse inactivation is shown in SEQ ID NO.116.
[0168] (2) Use lentivirus to integrate it into 293T cells to construct a stable expression cell line.
[0169] In unedited cell lines, the GFP gene cannot be expressed normally due to its inversion. After inversion editing using the PIEv3b method, the GFP gene is reversed and restored to normal expression, thus producing fluorescence, as shown in the image. Figure 18 As shown in a. (As indicated by...) Figure 18 As shown in b, flow cytometry sorting revealed that 13.1% of cells underwent editing, restoring GFP expression.
[0170] For GFP-recovery, the pegRNA used is also as shown in SEQ ID NO.2. The spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differ for different inversion events. See Table 10 below for details.
[0171] Table 10
[0172] This embodiment also used the PIEv3b method to perform intraarm inversions on a 40Mb region of chromosome 6 in HEK293T cells and verified its integrity. For the Chr6-Inv 40 Mb region, the pegRNA used was also as shown in SEQ ID NO.2. The spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differed for different inversion events. See Table 11 below for details.
[0173] Table 11
[0174] After transfecting 293T cells, single-cell clones were sorted and identified after cell growth. Positive single-cell clones exhibiting inversion editing were subjected to fluorescence in situ hybridization. The probes used in this embodiment contained both green and red fluorescent dyes; green fluorescent dye labeled the centromere of chromosome 6, and red fluorescent dye labeled the ESR1 gene at the end of the long arm of chromosome 6.
[0175] like Figure 19 As shown, in unedited cells, red fluorescence is located at the end of the long arm of chromosome 6; in cells where inversion editing has occurred, red fluorescence is located in the middle of the long arm of chromosome 6. This demonstrates that an inversion occurred in a 40Mb region.
[0176] Example 4: Application of the DNA pilot editing system of the present invention in the treatment of hemophilia A
[0177] 1. Application of disease-related loci
[0178] Some diseases and cancers are associated with large inversions in the genome. Hemophilia A (HA) is caused by recombination between inverted repeat sequences on the X chromosome, leading to the inactivation of the F8 gene. Depending on the location of the inversion, it can be classified as Inv1 (140 kb) or Inv22 (550 kb). Furthermore, some inversions can cause fusions of distant genes, forming oncogenes. For example, acute myeloid leukemia (AML) is caused by a 52 Mb inversion on chromosome 16, forming the CBFB-MYH11 oncogene. In HEK293T cells, we performed 140 kb and 550 kb inversions on the X chromosome using PIEv3b, while simultaneously deleting the inverted homologous sequences that caused the disease inversion; we also performed a 52 Mb inversion on chromosome 16 to mimic the AML genotype.
[0179] For ChrX-Inv 140 kb, ChrX-Inv 550 kb, and Chr16-Inv 52 Mb, the nucleotide sequences of the pegRNA plasmids used are also as shown in SEQ ID NO.2. The spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differ for different inversion events. See Table 12 below for details.
[0180] Table 12
[0181] Three days after co-transfecting 293T cells with pegRNA and PE2, the genome was harvested. ddPCR identified an inversion editing efficiency of 21.2-31.5% in the 140kb and 550kb inversions of hemophilia A; and an editing efficiency as high as 15% in the 52Mb inversion of acute myeloid leukemia (AML). Figure 20 These results indicate that PIEv3b has great potential in creating disease models.
[0182] 2. Creation of Chromosomal Structural Mutants
[0183] During biological evolution, human chromosomes exhibit centrocentric or subcentrocentric structures, while mice, whose genomes are highly similar to humans, all exhibit telomere structures (except for the Y chromosome). This embodiment utilizes the PIEv3b method from Example 3 to transform human cell chromosomes from an "X" shape with centrocentric structures to a "V" shape with telomere structures, as shown below. Figure 21 As shown.
[0184] Specifically, in this embodiment, a 30Mb region on chromosome 20 and a 100Mb region on chromosome 2 were selected for inversion. The nucleotide sequences of the pegRNA plasmids used for Chr20-Inv 30 Mb and Chr2-Inv 100 Mb are also shown in SEQ ID NO.2. For different inversion events, the spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 are different. See Table 13 below for details.
[0185] Table 13
[0186] After co-transfecting 293T cells with pegRNA and PE2, single-clone sorting was performed, and then the positive clones with inversion editing were subjected to fluorescence in situ hybridization.
[0187] like Figure 22 As shown, in the probes targeting chromosome 2, green fluorescence labeled the centromere of chromosome 2; in unedited cells, chromosome 2 exhibited an "X" shape with a centrocentric structure, while in cells where inversion editing occurred, chromosome 2 exhibited a "V" shape with a telocentric structure. This demonstrates the occurrence of inversion in the 100Mb region.
[0188] like Figure 23 As shown, the probe used to target chromosome 20 contains both red and green fluorescence. The red fluorescence marks the YWHAB gene on the long arm, and the green fluorescence marks the MARPE1 gene near the centromere on the long arm. The inverted end interface is located in the green fluorescently marked region, but the MARPE1 gene is not destroyed.
[0189] like Figure 23 As shown, in unedited cells, chromosome 20 exhibits two green fluorescent dots and two red fluorescent dots, and the chromosome is an "X" shape with a centrocentric structure; in cells where inversion editing has occurred, chromosome 20 has four green fluorescent dots and two red fluorescent dots, and the chromosome is a "V" shape with a telomere structure, proving that inversion occurred in the 30Mb region.
[0190] Example 5: Design and Feasibility Verification of PIES
[0191] While PIEv3b can generate efficient inversion editing on different chromosomes, it also introduces small deletions at the interfaces at both ends. To address this issue, we attempted to use a pair of pegRNAs to achieve seamless inversion editing, named PIES (PIE Seamless versions).
[0192] A third aspect of the present invention also provides a DNA leader editing system for inverting target DNA sequences, the DNA leader editing system comprising a Cas protein and a reverse transcriptase, pegRNA-E and pegRNA-F;
[0193] The pegRNA-E includes a fifth guide RNA, a fifth primer binding site, and a fifth reverse transcription template; the pegRNA-F includes a sixth guide RNA, a sixth primer binding site, and a fifth reverse transcription template; the fifth and sixth guide RNAs target the target fragments respectively.
[0194] The fifth reverse transcription template includes a fifth paired fragment, and the sixth reverse transcription template includes a sixth paired fragment; the fifth paired fragment is complementary to the gene fragment sequence adjacent to the motif near the 3' end of the original spacer sequence of the cleavage site guided by the sixth guide RNA to cut the Cas protein, and the sixth paired fragment is complementary to the gene fragment sequence adjacent to the motif near the 3' end of the original spacer sequence of the cleavage site guided by the fifth guide RNA to cut the Cas protein.
[0195] The pegRNA-E and pegRNA-F respectively guide the Cas protein to cleave the target fragment at the restriction sites on both sides of the two opposite strands of the double-stranded DNA; the original spacer sequence adjacent motif of the restriction site cleaved by the Cas protein guided by pegRNA-E and the original spacer sequence adjacent motif of the restriction site cleaved by the Cas protein guided by pegRNA-F are set far away from the corresponding spacer sequence.
[0196] like Figure 24 As shown, under the action of PE2 (nCas9-RT) and two pegRNAs (pegRNA E and pegRNA F), two 3' single-stranded DNA sequences (inversion editing version 1, 3' flap) are introduced at two sites on the two strands of the target genomic double-stranded DNA via reverse transcriptase (RT); the 3' flap generated by pegRNA E is complementary to the genomic sequence to the left of restriction site B, and the 3' flap generated by pegRNA F is complementary to the genomic sequence to the right of restriction site A.
[0197] When these two 3' flaps bind to their complementary genomic regions via annealing, inversions occur through DNA repair. pegRNA E and pegRNA F are designed for PAM-out, and this targeted DNA inversion method is named PIES-PAM-out.
[0198] 2. Feasibility verification of PIES-PAM-out
[0199] To verify the feasibility of PIES-PAM-out, this embodiment attempts to perform 100kb and 1Mb inversions on chromosome 6 and 10Mb inversions on chromosome 2 in HEK293T and K562 cells.
[0200] The pegRNA used is also shown in SEQ ID NO.2. The spacer sequence at positions 3501-3520 and the RTT and PBS sequences at positions 3597-3639 differ for different inversion events. See Table 14 below for details.
[0201] Table 14
[0202] In both cell lines, pegRNA and PE2 were co-transfected into cells, and the genome was harvested 3 days later. Editing efficiency was quantified using ddPCR. Figure 25 As shown in (a) and (b), the inversion editing efficiencies in HEK293T and K562 cells were 0.5–10.0% and 0–6.6%, respectively.
[0203] To further verify its precise and seamless inversion editing, second-generation sequencing was performed on both ends of the interface, such as... Figure 26 As shown, the product purity in HEK293T cells was 95.0-97.0%. This verifies that PIES-PAM-out can produce precise and seamless inversion editing.
[0204] The above detailed embodiments describe the implementation of the present invention; however, the present invention is not limited to the specific details described in the above embodiments. Within the scope of the claims and technical concept of the present invention, various simple modifications and changes can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.
Claims
1. A DNA leader editing system for inverting target DNA sequences, characterized in that, For inversion editing in mammalian cell lines, the DNA lead editing system includes a Cas protein and reverse transcriptase, pegRNA-A and pegRNA-B; the Cas protein is a Cas9 protein containing an inactivated HNH domain, or a Cas9 protein containing an inactivated RuvC domain; The pegRNA-A includes a first guide RNA, a first primer binding site, and a first reverse transcription template; the pegRNA-B includes a second guide RNA, a second primer binding site, and a second reverse transcription template; the first guide RNA and the second guide RNA target the target fragment respectively. The first reverse transcription template includes a first paired fragment, and the second reverse transcription template includes a second paired fragment; the sequences of the first paired fragment and the second paired fragment are complementary, and the sequence lengths of the first paired fragment and the second paired fragment are 3 to 200 nucleotides; The pegRNA-A and pegRNA-B respectively guide the Cas protein to cleave the target fragment at the restriction enzyme sites located on both sides of the same strand of the double-stranded DNA; the restriction enzyme sites located on both sides of the same strand of the double-stranded DNA are spaced 500 to 100,000,000 base pairs apart. The sequence lengths of the first and second reverse transcription templates are 15 to 2,000 nucleotides.
2. The DNA leader editing system for inverting target DNA sequences according to claim 1, characterized in that, The sequence lengths of the first and second reverse transcription templates are 15 to 500 nucleotides.
3. The DNA leader editing system for inverting target DNA sequences according to claim 1, characterized in that, The first reverse transcription template further includes a first non-complementary template sequence, and the second reverse transcription template further includes a second non-complementary template sequence; the first non-complementary template sequence and the second non-complementary template sequence are not complementary to each other.
4. The DNA leader editing system for inverting target DNA sequences according to claim 3, characterized in that, The lengths of the first non-complementary template sequence and the second non-complementary template sequence are 1 to 2,000 nucleotides.
5. The DNA leader editing system for inverting target DNA sequences according to claim 4, characterized in that, The lengths of the first non-complementary template sequence and the second non-complementary template sequence are 1 to 1,000 nucleotides.
6. The DNA leader editing system for inverting target DNA sequences according to claim 5, characterized in that, The lengths of the first non-complementary template sequence and the second non-complementary template sequence are 1 to 500 nucleotides.
7. The DNA leader editing system for inverting target DNA sequences according to claim 1, characterized in that, The restriction enzyme sites located on both sides of the same strand of the double-stranded DNA are spaced 10,000 to 30,000,000 base pairs apart.
8. The DNA leader editing system for inverting target DNA sequences according to claim 1, characterized in that, The pegRNA-A and / or pegRNA-B further include a tail sequence; the tail sequence forms a hairpin or loop with itself, or the tail sequence includes a poly(A), poly(U), or poly(C) sequence, or the tail sequence includes an RNA-binding domain.
9. A DNA leader editing system for inverting target DNA sequences, characterized in that, For inversion editing in mammalian cell lines, the DNA lead editing system includes a Cas protein and reverse transcriptase, pegRNA-A, pegRNA-B, pegRNA-C, and pegRNA-D; the Cas protein is a Cas9 protein containing an inactivated HNH domain or a Cas9 protein containing an inactivated RuvC domain. The pegRNA-A includes a first guide RNA and a first reverse transcription template; the pegRNA-B includes a second guide RNA and a second reverse transcription template; the pegRNA-C includes a third guide RNA, a third primer binding site, and a third reverse transcription template; and the pegRNA-D includes a fourth guide RNA, a fourth primer binding site, and a fourth reverse transcription template. The first reverse transcription template includes a first paired fragment, the second reverse transcription template includes a second paired fragment, the third reverse transcription template includes a third paired fragment, and the fourth reverse transcription template includes a fourth paired fragment; the sequences of the first paired fragment and the second paired fragment are complementary, and the sequences of the third paired fragment and the fourth paired fragment are complementary; the sequence lengths of the first paired fragment and the second paired fragment are 3 to 200 nucleotides. The pegRNA-A and pegRNA-B respectively guide the Cas protein to cleave the target fragment at enzyme sites located on both sides of one strand of the double-stranded DNA; the pegRNA-C and pegRNA-D respectively guide the Cas protein to cleave the target fragment at enzyme sites located on both sides of the other strand of the double-stranded DNA; the enzyme sites located on both sides of the same strand of the double-stranded DNA are spaced 500 to 100,000,000 base pairs apart. The protospacer sequence adjacent to the cleavage site of the Cas protein cleavage site guided by pegRNA-A and the protospacer sequence adjacent to the cleavage site of the Cas protein cleavage site guided by pegRNA-C are positioned close to or far from the corresponding spacer sequence. The sequence length between A and C is 2 to 10,000, and the sequence length between B and D is 2 to 10,000; A is the restriction enzyme site on one strand of the double-stranded DNA that the pegRNA-A guides the Cas protein to cut the target fragment; B is the restriction enzyme site on one strand of the double-stranded DNA that the pegRNA-B guides the Cas protein to cut the target fragment; C is the restriction enzyme site on the other strand of the double-stranded DNA that the pegRNA-C guides the Cas protein to cut the target fragment; and D is the restriction enzyme site on the other strand of the double-stranded DNA that the pegRNA-D guides the Cas protein to cut the target fragment.
10. The DNA leader editing system for inverting target DNA sequences according to claim 9, characterized in that, The sequence length between A and C is 500 to 2,000 nucleotides.
11. The DNA leader editing system for inverting target DNA sequences according to claim 9, characterized in that, The sequence length between B and D is 500 to 2,000 nucleotides.
12. The DNA leader editing system for inverting target DNA sequences according to claim 9, characterized in that, The restriction enzyme sites located on both sides of the same strand of the double-stranded DNA are spaced 10,000 to 30,000,000 base pairs apart.
13. A DNA leader editing system for inverting target DNA sequences, characterized in that, For inversion editing in mammalian cell lines, the DNA leader editing system includes a Cas protein and reverse transcriptase, pegRNA-E and pegRNA-F; the Cas protein is a Cas9 protein containing an inactivated HNH domain, or a Cas9 protein containing an inactivated RuvC domain; The pegRNA-E includes a fifth guide RNA, a fifth primer binding site, and a fifth reverse transcription template; the pegRNA-F includes a sixth guide RNA, a sixth primer binding site, and a sixth reverse transcription template; the fifth and sixth guide RNAs target the target fragments respectively. The fifth reverse transcription template includes a fifth paired fragment, and the sixth reverse transcription template includes a sixth paired fragment; the fifth paired fragment is complementary to the gene fragment sequence adjacent to the motif near the 3' end of the original spacer sequence of the cleavage site guided by the sixth guide RNA and is immediately adjacent to the gene fragment sequence near the 3' end of the original spacer sequence of the cleavage site guided by the fifth guide RNA and is immediately adjacent to the gene fragment sequence near the 3' end of the original spacer sequence; the sixth paired fragment is complementary to the gene fragment sequence adjacent to the motif near the 3' end of the original spacer sequence of the cleavage site guided by the fifth guide RNA and is immediately adjacent to the gene fragment sequence. The pegRNA-E and pegRNA-F respectively guide the Cas protein to cleave the target fragment at the restriction sites on both sides of the two opposite strands of the double-stranded DNA; the original spacer sequence adjacent motif of the restriction site cleaved by the Cas protein guided by pegRNA-E and the original spacer sequence adjacent motif of the restriction site cleaved by the Cas protein guided by pegRNA-F are set far away from the corresponding spacer sequence. The restriction enzyme sites located on opposite sides of the two opposite strands of the double-stranded DNA are spaced 500 to 50,000,000 base pairs apart.
14. A substance characterized in that, The substance is any one of the following: (1) A nucleic acid composition comprising the coding gene of the DNA lead editing system according to any one of claims 1-13; (2) A recombinant vector comprising the coding gene of the DNA lead editing system according to any one of claims 1-13; (3) A recombinant microorganism, said recombinant microorganism containing the coding gene of the DNA lead editing system according to any one of claims 1-13; (4) A kit comprising the DNA lead editing system of any one of claims 1-13; or comprising a recombinant vector comprising the coding gene of the DNA lead editing system of any one of claims 1-13.
15. The substance according to claim 14, characterized in that, The recombinant vector includes the coding gene of the DNA lead editing system according to any one of claims 9-11, and the pegRNA-A and pegRNA-D are recombined on the same recombinant vector, and the pegRNA-B and pegRNA-C are recombined on the same recombinant vector.
16. The use of the DNA lead editing system according to any one of claims 1-13 in inverted target DNA sequences not for the purpose of disease diagnosis and treatment.
17. A method for retrieving inverted target DNA sequences not for the purpose of disease diagnosis and treatment, characterized in that, Includes the following steps: The target DNA sequence and the DNA lead editing system according to any one of claims 1-8 are added to the reaction system. The reverse transcriptase uses the first reverse transcription template and the second reverse transcription template to reverse transcribe and generate two single-stranded DNA sequences, respectively. The two single-stranded DNA sequences complement each other to form a double-stranded region. Based on the DNA repair mechanism, the target DNA sequence is inverted, and the complementary DNA sequence of the first paired fragment is inserted at the interface near the first paired fragment. Alternatively, it may include the following steps: The target DNA sequence and the DNA lead editing system according to any one of claims 9-12 are added to the reaction system; the reverse transcriptase uses the first reverse transcription template, the second reverse transcription template, the third reverse transcription template and the fourth reverse transcription template to reverse transcribe and generate 4 single-stranded DNA sequences respectively, and the 4 single-stranded DNA sequences complementarily combine to form 2 double-stranded regions; When the protospacer motif of the cleavage site guided by pegRNA-A and the protospacer motif of the cleavage site guided by pegRNA-C are far apart relative to the corresponding spacer sequence, the target DNA sequence is inverted based on the DNA repair mechanism. A complementary DNA sequence of the first paired fragment is inserted near the interface of the first paired fragment, and a copy of the sequence between A and C is added in situ. Similarly, a complementary DNA sequence of the third paired fragment is inserted near the interface of the third paired fragment, and a copy of the sequence between B and D is added in situ. When the protospacer motif of the cleavage site guided by pegRNA-A and the protospacer motif of the cleavage site guided by pegRNA-C are positioned close to the corresponding spacer sequence, the target DNA sequence is inverted based on the DNA repair mechanism. At the interface near the first paired fragment, a copy of the sequence between A and C is deleted and a complementary DNA sequence of the first paired fragment is inserted; at the interface near the third paired fragment, a copy of the sequence between B and D is deleted and a complementary DNA sequence of the third paired fragment is inserted. A is the restriction enzyme site on one strand of the double-stranded DNA where the Cas protein, guided by pegRNA-A, cuts the target fragment; B is the restriction enzyme site on one strand of the double-stranded DNA where the Cas protein, guided by pegRNA-B, cuts the target fragment on the other strand of the double-stranded DNA; C is the restriction enzyme site on the other strand of the double-stranded DNA where the Cas protein, guided by pegRNA-C, cuts the target fragment on the other strand of the double-stranded DNA; D is the restriction enzyme site on the other strand of the double-stranded DNA where the Cas protein, guided by pegRNA-D, cuts the target fragment on the other strand of the double-stranded DNA. Alternatively, it may include the following steps: The target DNA sequence and the DNA lead editing system of claim 13 are added to the reaction system; the reverse transcriptase uses the fifth and sixth reverse transcription templates to reverse transcribe and generate two single-stranded DNA sequences, respectively, and the two single-stranded DNA sequences bind to their corresponding complementary genomic regions to form double-stranded regions; the target DNA sequence is inverted based on the DNA repair mechanism.
18. The method for retrieving inverted target DNA sequences according to claim 17, not for the purpose of disease diagnosis and treatment, characterized in that, The Cas protein is either a Cas9 protein containing an inactivated HNH domain or a Cas9 protein containing an inactivated RuvC domain.
19. The method for retrieving inverted target DNA sequences according to claim 18, not for the purpose of disease diagnosis and treatment, characterized in that, The Cas protein is SpCas9, FnCas9, St1Cas9, St3Cas9, NmCas9, SaCas9, VQR SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-NG, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, or AtCas9.