A method for targeted replication of DNA fragments
By using amplification editing (AE) to cut and extend the RT template on both sides of the target DNA sequence with pegRNA to form a double-stranded region for DNA polymerase replication, the problem of difficulty in inserting large fragments in prime editing is solved, and efficient and precise DNA fragment replication is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN UNIV
- Filing Date
- 2023-05-30
- Publication Date
- 2026-08-04
AI Technical Summary
Existing prime editing technologies struggle to insert large DNA fragments without relying on donor DNA, and existing methods are inefficient at the genome level and cannot achieve site-specific amplification of target DNA sequences.
The amplification-editing (AE) method uses a pair of pegRNAs to cut the target DNA sequence on both sides, and extends the target fragment through a reverse transcriptase (RT) template to form a double-stranded region to enable DNA polymerase replication and insert a new DNA fragment.
It enables efficient replication of target DNA fragments without relying on donor DNA, and can replicate DNA fragments with high precision and efficiency in a variety of cell lines and genomic loci, ranging from 20bp to 100Mb, including human chromosomal regions.
Smart Images

Figure HDA0005260159900000011 
Figure HDA0005260159900000012 
Figure HDA0005260159900000021
Abstract
Description
[0001] background
[0002] Inserting nucleic acid fragments into target nucleic acids (such as genomic sequences) is a challenging task. Targeted transgene integration is typically achieved via homologous recombination repair (HDR), however this method relies on a foreign DNA donor and is ineffective in non-dividing cells. Homolog-independent targeted integration (HITI) strategies are cell cycle independent, but they are less efficient at the genomic level (typically around 1-5%), and multiple editing events are observed to occur in combination.
[0003] A recently developed CRISPR-based gene editing tool, called Prime editing (PE), connects reverse transcriptase (RT) to the Cas9 nicking enzyme. Its RT template (RTT) is located at the 3' end of the Prime editing guide RNA (pegRNA), allowing for precise modification of the cleavage site. Prime editing enables various types of base substitutions, small fragment insertions, and deletions without relying on donor DNA, holding immense potential for applications in basic research and correction of genetic mutations related to human diseases. However, to date, Prime editing has not been used to insert large DNA fragments.
[0004] On the other hand, replicating a portion of the target nucleic acid presents another challenge. In Prime editing, the insertion is triggered by an RT template. Therefore, the length that can be replicated is limited. Furthermore, a new RT template is required for each target fragment to be replicated.
[0005] Overview
[0006] DNA sequence replication holds immense potential for treating a variety of genetic diseases and has significant commercial value in industrial settings such as protein production. Currently, there is no technology capable of site-specific amplification of fragments on a target DNA sequence, at least not based on the newly developed CRISPR technology.
[0007] Inventors have developed a new method called amplification editing (AE), which can specifically amplify target fragments within a target DNA sequence, such as a genomic sequence. One example of the AE method is the use of a pair of pegRNAs with target sites flanking the target fragment. The target fragment is extended using reverse transcriptase (RT) templates contained within the pegRNAs. Because the sequences contained in these two RT templates are at least partially complementary, they form a double-stranded region, which then serves as the starting point for DNA polymerase to synthesize a new strand, thus enabling the replication of the target fragment. In short, n rounds of amplification can produce up to 2...n A copy of the target segment is made. At the same time, the RT template-encoded sequence is inserted between the repeating target segments.
[0008] According to one embodiment, this disclosure provides a method for replicating a target fragment in a target DNA sequence in the presence of a DNA polymerase, the method comprising the target DNA sequence and components interacting therewith: (a) a Cas protein and a reverse transcriptase; and (b) a first Prime Editing guide RNA (pegRNA) containing a first CRISPR. (c) a second pegRNA containing a second crRNA and a second RT template sequence; wherein (i) the first RT template sequence contains a first paired fragment, (ii) the second RT template sequence contains a second paired fragment, (iii) the first paired fragment and the second paired fragment are complementary to each other, and (iv) the first pegRNA and the second pegRNA guide the Cas protein to cut at two sites on either side of the target fragment on the target DNA sequence, on the complementary strand, thereby allowing (1) the reverse transcriptase to extend the two opposite strands of the target fragment using the first and second RT template sequences to generate two flap DNA sequences, and (2) the two flap DNA sequences to form a double-stranded region, allowing the DNA polymerase to extend the double-stranded region to replicate the target fragment, thereby achieving replication of the target fragment and inserting an insert fragment between the two repeated target fragments, wherein one strand of the insert fragment contains the first fragment, the first complementary paired fragment and the reverse complementary sequence of the second fragment.
[0009] In some embodiments, the first pegRNA further includes a first primer binding site (PBS) and a first spacer sequence, while the second pegRNA further includes a second PBS and a second spacer sequence, such that the pegRNA can guide the Cas protein to two sites flanking the target fragment and initiate reverse transcription.
[0010] In some embodiments, the first and second RT template sequences are each 0 to 2000 nucleotides in length, preferably 15 to 500 nucleotides in length.
[0011] In some embodiments, the first and second complementary pairing fragments are each 0 to 1000 nucleotides in length, preferably 3 to 200 nucleotides or 3 to 50 nucleotides in length, and more preferably 30 to 100 nucleotides in length.
[0012] In some embodiments, the first and second RT template sequences each further comprise a non-complementary template sequence that is not complementary to each other, wherein each non-complementary template sequence is located between the corresponding complementary pairing fragment and crRNA, or between the corresponding complementary pairing fragment and PBS.
[0013] In some embodiments, each non-complementary template sequence is 1 to 2000 nucleotides long, preferably 1 to 1000 or 1 to 500 nucleotides long.
[0014] In some embodiments, the two sites on either side of the target fragment are 2 to 1,000,000,000 base pairs apart, preferably 10 to 5,000,000 base pairs apart.
[0015] In some embodiments, each RT template sequence further includes an additional sequence adjacent to a complementary pairing fragment, wherein the two additional sequences are complementary to the target DNA sequence and are at least partially complementary to each other.
[0016] This disclosure also provides a method for replicating a target fragment in a target DNA sequence in the presence of DNA polymerase, the method comprising a target DNA sequence and the following components interacting therewith: (a) a Cas protein; (b) a first single-guide RNA (sgRNA) or transcribed RNA (tracrRNA); and (c) a second sgRNA or tracrRNA, wherein both the first sgRNA or tracrRNA and the second sgRNA or tracrRNA are complementary to target site sequences flanking the target DNA sequence fragment, and the two target site sequences are at least partially complementary, wherein the first sgRNA or tracrRNA binds to one strand of the first target site in the presence of the Cas protein and cleaves the opposite strand, releasing the opposite strand as the first single-stranded DNA; and the second sgRNA or tracrRNA binds to one strand of the second target site in the presence of the Cas protein and cleaves the opposite strand, releasing the opposite strand as the second single-stranded DNA; and the first single-stranded DNA and the second single-stranded DNA bind to form a double-stranded region, allowing DNA polymerase to extend the double-stranded region to replicate the target sequence, thereby achieving sequence replication between the two target sites.
[0017] In some embodiments, the partial complement length includes complete complementarity of at least 3, 4, 5, 6, 7 or 8 consecutive nucleotides.
[0018] This disclosure also provides a method for replicating a target fragment in a target DNA sequence in the presence of DNA polymerase, the method comprising a target DNA sequence and the following components interacting therewith: (a) a Cas protein and a reverse transcriptase; (b) a prime editing guide RNA (pegRNA) containing a first CRISPR RNA (crRNA) and a reverse transcriptase (RT) template sequence; and (c) a single-stranded guide RNA (sgRNA) or transcription RNA (tracrRNA), wherein (i) the RT template sequence contains a fragment that can be used for complementary pairing, (ii) the pegRNA guides the Cas protein to cleave the opposite strand on one side of the target fragment in the target DNA sequence, thereby allowing the reverse transcriptase to extend the opposite strand of the target fragment using the RT template sequence to generate a single-stranded DNA sequence, and (iii) the sgRNA or tracrRNA guides the Cas protein to cleave the opposite strand on the other side of the target fragment in the target DNA sequence, thereby releasing the opposite strand as a second single-stranded DNA sequence; and the two single-stranded DNA sequences form a double-stranded region, enabling the DNA polymerase to extend the double-stranded region to replicate the target fragment, thereby achieving replication of the target fragment.
[0019] In some embodiments, the target DNA sequence is present inside the cell. In some embodiments, the cell is a dividing cell. In some embodiments, the cell is in a non-dividing state.
[0020] In some embodiments, the target fragment is a telomere or a fragment thereof. In some embodiments, the method is performed in vitro. In some embodiments, the method is performed in vivo.
[0021] In some embodiments, the Cas protein is a cleavage enzyme. In some embodiments, each pegRNA comprises, from 5' to 3', a first or second crRNA, a first or second complementary pair fragment, a first or second fragment, and a first or second PBS. In some embodiments, the cleavage enzyme is a Cas9 protein having an inactive HNH domain that cleaves the target strand.
[0022] In some embodiments, the nicking enzyme is a nicking enzyme derived from SpCas9, FnCas9, St1Cas9, St3Cas9, NmCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, VQR SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-NG, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, or atCas9.
[0023] In some embodiments, the Cas protein is the Cas12 protein.
[0024] In some embodiments, each pegRNA contains a first or second crRNA, a first or second PBS, a first or second fragment, and a first or second paired fragment in the 3' to 5' direction.
[0025] In some embodiments, the Cas12 protein is Cas12a, Cas12b, Cas12f, or Cas12i.
[0026] In some embodiments, the Cas12 protein is derived from a selected population, including AsCpf1, FnCpf1, SsCpf1, PcCpf1, BpCpf1, CmtCpf1, LiCpf1, PmCpf1, Pb3310Cpf1, Pb4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCas12b, EbCas12b, and LsCas12b.
[0027] In some embodiments, the first or second pegRNA further comprises (a) a tail capable of forming a hairpin or loop structure with itself, a PBS, an RT template sequence, crRNA, or a combination thereof, or (b) a poly(A), poly(U), or poly(C) sequence or an RNA-binding domain.
[0028] In some embodiments, the reverse transcriptase is M-MLV reverse transcriptase or a reverse transcriptase that functions under physiological conditions.
[0029] In some embodiments, the Cas protein and reverse transcriptase are each provided in the form of nucleotides encoding the respective proteins, or as proteins.
[0030] In some embodiments, each pegRNA is provided as recombinant DNA encoding the pegRNA or as an RNA molecule.
[0031] In some embodiments, the target fragment and the inserted fragment are further copied.
[0032] Brief description of the drawings
[0033] Figure 1 Overview of amplification editing. Schematic diagram of amplification editing (AE). Figure 1 a shows the structures of two pegRNAs positioned opposite to each other in the PAM orientation (e.g., the PAM sequences in the genome are not located within the two sites targeted by the spacer). Figure 1As shown in b, the target site is cleaved by nCas9-RT and reverse transcribed into two complementary 3' flap DNAs. These two 3' flap DNAs anneal each other and act as primers to initiate DNA replication. Subsequently, the original genomic DNA unwinds and synthesizes a new DNA strand. Synthesis stops when the new DNA strand encounters the cleavage.
[0034] Figure 2 Experimental design and results for detecting AE results. a. Schematic diagram of three PCR methods for detecting amplified editing. b. The TAE agarose gel PCR amplification products from the method in Figure a show that the 178bp and 234bp sequences at the VEGFA and HEK3 sites in HEK293T cells were amplified and inserted with a 20bp insertion. c. Sanger sequencing was performed on the editing product bands after Out-Out PCR, indicated by red triangles in Figure b, targeting the VEGFA and HEK3 sites.
[0035] Figure 3 The effects of flap DNA length and complementation length at different sites were investigated. ab, the effects of flap DNA length and complementation length of 10-100 bp on editing efficiency were quantitatively analyzed by droplet digital PCR (ddPCR) when the replication lengths of VEGFA and C-MYC sites in HEK293T cells were 178, 1014, and 7919 bp, respectively. c, DNA replication efficiency at VEGFA site 3' wings of 30, 40, 50, 80, or 100 bp and a complementation length of 30 bp. d, ddPCR was designed to quantify AE efficiency. Two primer / probe designs were used for control (reference gene / sequence) detection: one targeting the same chromosome at least 5 kb away from the replication region, and the other targeting a different chromosome. Primers for editing product detection were designed for In-In PCR, and probes were designed to target either the replication sequence or the insertion sequence.
[0036] Figure 4 The effect of GC content and multiplex editing on the AE method in 3' flaps. a, Replication mediated by VEGFA and C-MYC sites when the GC content of the RTT is 30%, 50%, or 80%. Replication efficiency was determined by ddPCR. Mean ± standard deviation, n = 3 independent biological replicates. b, AE efficiency when co-transfected with 2 or 3 pairs of pegRNAs targeting different regions, which was determined by ddPCR.
[0037] Figure 5Junction purity generated by the AE method. a, Design of head-to-tail PCR. The PCR products were determined by next-generation sequencing (NGS). b, Junction purity in the control group (WT) and the edited group during 178-bp or 234-bp amplification at the VEGFA and HEK3 loci measured by NGS. c, Junction purity during replication of 178, 1014, and 7919 bp mediated by different flap lengths and complementary lengths measured by NGS. d-e, Junction purity of replication mediated by different flaps and complements at the VEGFA (d) and C-MYC (e) loci in HEK293T cells measured by NGS.
[0038] Figure 6 Comparison of AE efficiency at each locus in different cell lines. a-e, AE efficiency of different amplification lengths in HEK293T cells (a), Huh-7 cells (b), K562 cells (c), U2OS cells (d), and N2a cells (e). Mean ± standard deviation, n = 3 independent biological replicates.
[0039] Figure 7 Editing results of AE shown by TAE agarose gel electrophoresis in different cell lines. a-c, Amplification products of In-In and In-Out PCR shown by TAE agarose gel electrophoresis.
[0040] Figure 8 Amplification editing for short sequences. a, Design of 20-bp replication. The length between the two incisions is 20 bp, and the two spacer sequences partially overlap. b, Replication efficiency of short fragment sequences. c, Out-Out PCR amplification products for 20-bp replication analyzed by NGS.
[0041] Figure 9 Tandem repeats achieved by the AE method. a, Overview of two models for multiple rounds of replication by AE. Left: The newly transcribed 3'-flap sequences are complementary to each other; they anneal to each other and then synthesize new DNA strands, resulting in 2 n repeated sequences after n rounds of replication. Right: When the 3'-flap pairs with the inserted sequence and anneals, its product has [dup*(n - 1)+x] amplified sequences (x < n), resulting in fewer than 2 n repeated sequences after n rounds of replication. b, Out-Out PCR amplification editing products of different single-cell clones shown by TAE agarose gel. The mixed results shown in the same clone are indicated by triangles. c, Sanger sequencing corresponding to the PCR amplification products in (b).
[0042] Figure 10Editing efficiency after restriction enzyme digestion was detected by ddPCR. a, Quantification of amplification efficiency of VEGFA-178bp and RUNX1-1132bp in K562 cells by ddPCR. b, Schematic diagram of restriction enzyme digestion of edited products. Before digestion, the tandem repeat sequences of the edited product genome are in one droplet (left); after digestion, each repeat is in a different droplet (right). cf, Determination of replication efficiency of VEGFA and HEK3 sites in K562 (c, d) and HEK293T (e, f) cells by ddPCR. Extracted genomic DNA was measured with or without restriction enzyme treatment. Mean ± standard deviation, n = 3 independent biological replicates. ***, P ≤ 0.001.
[0043] Figure 11 EGFP function was restored and HBA gene replication was performed using AE. a) The GFP sequence was perturbed and a 53bp coding sequence was deleted. GFP-L, the N-terminus of GFP; GFP-R, the C-terminus of GFP. b) Quantification of AE-treated cells by flow cytometry. c) Schematic diagram of generating and restoring the α-thalassemia genotype to a normal genotype. Double-strand breaks were generated in the HBA1 (α1) and HBA2 (α2) genes using Cas9 and two sgRNAs, resulting in a 3.7kb deletion and a fused α gene (named αh), which is identical to the α1 gene. The disease genotype was corrected to a normal genotype using AE. d) The replication efficiency of αh was determined by ddPCR.
[0044] Figure 12 Enhancement of miR-21 expression levels using AE. a) Amplification of the stem-loop region of the miR-21 precursor (197 bp). b) MiR-21 replication efficiency at days 3, 5, and 7 in HEK293T and K562 cells. c) Quantification of miR-21 expression levels by RT-qPCR. d) Determination of miR-21 target gene expression levels by RT-qPCR. **, P ≤ 0.01; *, P ≤ 0.05. Mean ± standard deviation, at least n = 3 independent biological replicates.
[0045] Figure 13Replication of large, mega-scale, and chromosome-scale DNA fragments in HEK293T cells. a) Determination of 30–100 kb replication on chromosomes 6 and 9 by in-in PCR. Top: Visualization of PCR products on TAE agarose gel. Bottom: Schematic diagram of Sanger sequencing of in-in PCR products. b) Determination of copy numbers of the IKZF4, RBMS2, and NEMP1 genes in the 1 Mb repeat region on chromosome 12 in control (WT) and single-cell clones. c) Quantification of 30–100 kb replication efficiency on chromosomes 6 and 9 by ddPCR. d) Quantification of 1–3 Mb replication efficiency on chromosomes 6, 9, and 12 by ddPCR. e) 10–100 Mb replication efficiency on chromosomes 6 and 9. Mean ± standard deviation, n = 3 independent biological replicates, applicable to cf. ***, P ≤ 0.001.
[0046] Figure 14 Visualization of chromosome-level replication after AE treatment in HEK293T cells. ab, Fluorescence in situ hybridization (FISH) for a 3Mb repeat on chromosome 12 (a) and a 100Mb repeat on chromosome 6 (b). The first column (from left to right) shows the centromere of chromosome 12 (a) or 6 (b). The second column shows a fluorescently labeled genomic sequence of approximately 108Kb, adjacent to the STAT6 gene (a), or a genomic sequence of approximately 198Kb, adjacent to the ESR1 gene (b). The first and second columns are merged in the third column. The fourth column is a magnified view of the third column. For (a) or (b), Edit-1 and Edit-2 are two different single-cell clones from AE-treated samples, targeting either the 3Mb (a) or 100Mb (b) repeat, respectively. The length of the fourth column in b is in micrometers.
[0047] Figure 15 To explore replication efficiency with shorter complement lengths. a, Replication efficiency for VEGFA-178bp amplification using a 0-10bp 3' flap.
[0048] Figure 16Various PAM-out amplification and editing methods. a, Schematic diagram of amplification and editing using different combinations of pegRNAs and sgRNAs. From top to bottom, the first (I) is the AE design described above; the second (II) is similar to the first, but a portion of the 3' flap is complementary to the genomic sequence near the cut; the third (III) uses one sgRNA and one pegRNA, where the 3' flap generated by the pegRNA is complementary to the genomic sequence near the sgRNA cut; the bottom design (IV) is amplification without RT enzymes using two sgRNAs and Cas9 nickase, but the two nick sequences are complementary to each other. b, Efficiency of replication in HEK293T cells using pegRNA+sgRNA or paired pegRNAs with 0, 8, 30, and 38 bp complementarity. c, Quantification of adapter purity in (b) by NGS. d, TAE agarose gel visualization of In-In PCR amplification products of paired sgRNA-mediated replication in HEK293T cells with 4 bp or 8 bp complementarity. e, adapter purity in (e) quantified by NGS. ***, P ≤ 0.001; **, P ≤ 0.01; *, P ≤ 0.05. Mean ± standard deviation, n = 3 independent biological replicates.
[0049] Detailed description
[0050] definition
[0051] It should be noted that "a" or "a certain" entity refers to one or more of that entity; for example, "an antibody" is understood to mean one or more antibodies. Therefore, the terms "a" (or "a certain"), "one or more", and "at least one" are used interchangeably here.
[0052] In this document, the term "polypeptide" is intended to encompass both single and plural "polypeptides," referring to molecules composed of monomers (amino acids) linearly linked by amide bonds (also known as peptide bonds). "Polypeptide" refers to a chain of two or more amino acids, not a product of a specific length. Therefore, peptide, dipeptide, tripeptide, oligopeptide, "protein," "amino acid chain," or any other term used to refer to two or more amino acid chains are included within the definition of "polypeptide," and the term "polypeptide" may be used in place of or interchangeably with any of the foregoing terms. The term "polypeptide" also refers to post-expression modified products of polypeptides, including but not limited to glycosylation, acetylation, phosphorylation, amination, derivatization with known protecting / blocking groups, proteolytic cleavage, or modification with non-natural amino acids. Polypeptides can be derived from natural biological sources or produced through recombinant technologies, but are not necessarily translated from a specified nucleic acid sequence. They can be generated in any manner, including chemical synthesis.
[0053] When polynucleotides are involved, the term "coding" refers to a polynucleotide that, if transcribed and / or translated in its original state or by methods well known to those skilled in the art, produces mRNA of the polypeptide and / or fragments thereof, is called a "coding" polypeptide. The antisense strand is the complementary sequence of the nucleic acid, from which the coding sequence can be deduced.
[0054] Amplification Editing
[0055] Prime editing (PE) is a genome editing technology that allows modification of an organism's genome. Prime editing directly inserts new genetic information into target DNA sites. It uses a fusion protein consisting of a catalytically impaired endonuclease (e.g., Cas9) and an engineered reverse transcriptase, along with a primeediting guide RNA (pegRNA) that recognizes the target site and provides new genetic information to replace the target DNA nucleotides. Prime editing mediates targeted insertions, deletions, and base transitions without the need for double-strand breaks (DSBs) or donor DNA templates.
[0056] PegRNAs recognize the target nucleotide sequence to be edited and encode new genetic information that replaces the target sequence. A pegRNA consists of an extended single guide RNA (sgRNA) containing a primer binding site (PBS) and a reverse transcriptase (RT) template sequence. During genome editing, the primer binding site allows the 3' end of the cut DNA strand to hybridize with the pegRNA, while the RT template serves as a template for synthesizing the edited genetic information. The sgRNA portion contains a guide sequence (spacer) that guides the prime editor to the target genome site and the sgRNA framework. When the guide sequence binds to the target genome sequence and unwinds the DNA double helix, the PBS binds to the other strand and initiates reverse transcription, using the RT template sequence as a template. The newly synthesized sequence (a "3' flap") binds to the target genome site, forming double-stranded DNA. The RT template can contain mutations or small insertions relative to the target genome sequence, but it needs to be substantially homologous to the target genome sequence because the newly synthesized DNA strand should still hybridize with one strand of the original target genome sequence.
[0057] The inventors have designed and implemented a new technique capable of efficiently and specifically amplifying target fragments, such as genomic sequences, on target DNA sequences. An example method employs a pair of pegRNAs, whose target sites are located flanking the target fragment, extending the fragment via reverse transcriptase (RT) templates contained within the pegRNAs. Since the two RT templates contain at least complementary portions, they can form a double-stranded region, serving as a starting point for DNA polymerase to synthesize new strands of each strand of the target fragment, thereby replicating the target fragment.
[0058] This new editing technique is called "Amplification Editing" (AE), such as... Figure 1 As shown. In a sample AE program, two pegRNAs are used. They are typically located at the PAM-out position. Each pegRNA consists of a CRISPR RNA (crRNA or sgRNA), a reverse transcriptase (RT) template sequence, and a primer binding site (PBS). The PBS may be complementary to the guide sequence (or “spacer sequence”) in the crRNA, but is usually a few nucleotides shorter. When the guide sequence binds to the target genomic sequence and causes the DNA double helix to unwind, the PBS can bind to the other strand and initiate reverse transcription, using the RT template sequence as a template.
[0059] Unlike pegRNAs used in traditional prime editing, in each pair of pegRNAs in an amplification-editing system, the RT template sequence does not necessarily need to be homologous to the target genome sequence. In some embodiments, the RT template preferably has reduced or even no homology to the target genome sequence.
[0060] Conversely, these two RT templates share a complementary part. For example, as Figure 1 As shown, in each pegRNA, the RT template consists of two parts: a paired RT fragment and an RT fragment. These two paired (complementary) RT fragments have complementary sequences, so the DNA sequences reverse transcribed from them can pair with each other.
[0061] It is important to note that while complementary portions are required to form flap DNA sequences that can bind to each other, the RT template does not necessarily contain non-complementary "RT fragments." In other words, in some embodiments, the two RT templates are completely complementary to each other.
[0062] When the guide sequence binds to the target genome sequence and causes the DNA double helix to unwind, PBS binds to the other strand and initiates reverse transcription, using the RT template sequence as a template. Figure 1As shown, the two pegRNAs are designed in a PAM-out design so that they cut at two positions flanking the target portion (step 110). The RT template then serves as a template to synthesize single-stranded DNA, thereby introducing two single-stranded "flaps" that extend the sequence to be replicated in opposite directions (step 120). The direction of extension is ensured by the design of the pegRNA molecule.
[0063] Each flap contains a reverse transcription of an RT fragment from the pegRNA and further includes a portion from the paired (complementary) RT fragment. Due to their complementarity, these two distal fragments can hybridize with each other to form a double-stranded region (step 130). This double-stranded region can then serve as a starting point for DNA polymerase.
[0064] Using the double-stranded region as a starting point and the single-stranded DNA of the genome as a template, a new DNA strand is synthesized, unwinding the original DNA between the two incisions (steps 140 - 150). Eventually, the target sequence is precisely replicated while inserting a small flap sequence (the sequence generated by the 3' flap) in between.
[0065] As shown in Examples 1 - 2 ( Figure 2-3 ), this newly designed amplification editing system has high precision and high efficiency. In addition, this new editing technology is extremely powerful, and the flap sequence can be as short as 10 nucleotides or as long as 100 nucleotides or more without significantly affecting the editing efficiency (Example 2, Figure 3 ). AE is active in multiple cell lines and at multiple genomic loci (Example 3, Figure 6-7 ). In addition, AE is able to replicate human genomic regions ranging from 20 bp to at least 100 megabases (Mb). Considering that the average size of a human chromosome is approximately 100 Mb, this result is quite unexpected. (Examples 2, Figure 3 、 Figure 8 、 Figure 13-14 ).
[0066] In another surprising finding, since each round of AE does not disrupt the flanking sequences containing the PAM sequence or the cleavage site, amplification can be repeated. As Figure 9 shown and demonstrated in Figure 9-10 , the AE method is indeed able to continue amplifying the target fragment. When the newly transcribed 3' flaps are complementary to each other, they anneal to each other, followed by the synthesis of a new DNA strand, resulting in a duplication of 2 n after n rounds of replication. When the 3' flap pairs with and anneals to the inserted sequence, the resulting product has [dup*(n - 1)+x] amplified sequences (x < n), resulting in a duplication number less than 2 n after n rounds of replication.
[0067] Alternative designs of AE
[0068] It is also worth noting that, in addition to paired pegRNAs, replication can be achieved through other methods. When these pairs are aligned in the PAM-out direction and share complementary sequences near the cleavage site to allow annealing to form DNA synthesis primers, they achieve pegRNA / sgRNA replication. Figure 16 a, III) or sgRNA / sgRNA ( Figure 16 a, IV) combination ( Figure 15-16 Additionally, the 3' flaps can be partially complementary to each other, and also partially complementary to genomic sequences near cleavage sites induced by another pegRNA. Figure 16 a, II).
[0069] pegRNA + sgRNA (or tracrRNA)
[0070] Prove in Example 7 ( Figure 15-16 The AE system does not require two complementary flaps. Figure 16 In design (III), the assembly on the left still includes an nCas9-RT protein and a pegRNA, which can generate a 3' flap during reverse transcription of the RT-paired fragment. On the right, only a regular sgRNA (or a regular tracrRNA, i.e., without an RT template) is used, which can work with a regular nCas9 protein or the same nCas9-RT protein, but its RT activity is not required.
[0071] In this design, the single 3' flap generated on the left is complementary to the genomic sequence near the cleavage site on the right. As shown in Example 7, their complementarity can also initiate DNA unwinding and replication, resulting in the replication of the sequence between the two cleavage sites.
[0072] sgRNA+sgRNA (or tracrRNA+tracrRNA, sgRNA+tracrRNA, tracrRNA+sgRNA)
[0073] Another alternative design is to not generate a 3' flap. Therefore, the entire system ( Figure 16 a(IV) does not contain reverse transcriptase or pegRNA. On the left and right sides, conventional nCas9 / sgRNA (or tracrRNA) assembly was used, provided that the genomic sequences near the two cleavage sites were at least partially complementary.
[0074] In the absence of a single 3' flap, at each cut site, the exposed non-target strand allows for the annealing of two complementary non-target genomic sequences, forming primers for DNA synthesis, leading to DNA replication.
[0075] Mixed pegRNA + Mixed pegRNA
[0076] Another alternative to design (I) is design (II) Figure 16 a). Design (II) still uses two nCas9-RT and two pegRNA sequences. Unlike design (I), the pegRNAs in design (II) are partially complementary to each other and partially complementary to the genomic sequences located near the contralateral cleavage sites.
[0077] Essentially, in design (II), not only can the 3' flap initiate annealing, but the adjacent genomic sequence can also play a role, thus allowing the 3' flap sequence to be shorter.
[0078] Therefore, according to one embodiment of this disclosure, this disclosure provides a method for replicating a target fragment in a target DNA sequence in the presence of a DNA polymerase. In some embodiments, the method involves interacting the target DNA sequence with the following components: (a) a Cas protein and a reverse transcriptase, (b) a first prime editing guide RNA (pegRNA) comprising a first CRISPR RNA (crRNA / sgRNA), a first reverse transcriptase (RT) template sequence, and (c) a second prime editing guide RNA (pegRNA) comprising a second crRNA / sgRNA and a second RT template sequence. In some embodiments, the first pegRNA further comprises a first primer binding site (PBS) and a first spacer, while the second pegRNA further comprises a second PBS and a second spacer.
[0079] In some embodiments, the first RT template sequence includes a first paired fragment, and the second RT template sequence includes a second paired fragment, the first and second paired fragments being complementary. Therefore, the first and second pegRNAs can guide the Cas protein to cleave the opposite strands of the target sequence on both sides of the target DNA. Subsequently, reverse transcriptase extends the two opposite strands of the target sequence using the first and second RT sequences as templates, thereby generating two single-stranded flap DNA sequences.
[0080] Furthermore, the two single-stranded flap sequences can form a double-stranded region, allowing DNA polymerase to extend this double-stranded region to replicate each target fragment. Thus, the target fragment is replicated. Simultaneously, an insert fragment is inserted between the two replicated target fragments, one of which contains the reverse complementary sequences of the first fragment, the first paired fragment, and the second fragment.
[0081] In some embodiments, each RT template sequence also includes an additional sequence (“mixed pegRNA / mixed pegRNA”) adjacent to the paired fragment, wherein the two additional sequences are complementary to the target DNA sequence and are at least partially complementary to each other.
[0082] Another method for replicating a target region in the genome in the presence of DNA polymerase includes the following components that interact with the target sequence: (a) a Cas protein, (b) a first guide RNA (sgRNA1), and (c) a second guide RNA (sgRNA2), wherein the first and second guide RNAs have complementary sequences flanking the spacer sequence in the target DNA sequence, and the two target sites are at least partially complementary; wherein, under the action of the Cas protein, the first guide RNA binds to one strand of the target site and cleaves the opposite strand, releasing the opposite strand to form a first single-stranded DNA; the second guide RNA, under the action of the Cas protein, binds to the other strand of the target site and cleaves the opposite strand, releasing the opposite strand to form a second single-stranded DNA; the two single-stranded DNAs combine to form a double-stranded region, allowing DNA polymerase to replicate the double-stranded region, thereby replicating both strands between the target sequences, ultimately achieving the replication of the target sequence.
[0083] In these methods, partial complementarity includes complete complementarity of at least 2, 3, 4, 5, 6, 7, or 8 nucleotides, which are typically sequential.
[0084] A method for replicating a targeted region in the genome in the presence of DNA polymerase is also provided, comprising the following components and interacting with the target sequence: (a) a Cas protein and reverse transcriptase, (b) a guide RNA (pegRNA) consisting of a CRISPR RNA (crRNA) and a reverse transcriptase (RT) template sequence, and (c) a single-stranded guide RNA (sgRNA); wherein: (i) the reverse transcriptase template sequence contains a pairable fragment, (ii) the pegRNA guides the Cas protein to cleave the target DNA, allowing the reverse transcriptase to extend the opposite strand of the target fragment using the RT sequence as a template, thereby generating a single-stranded DNA sequence, and (iii) the sgRNA guides the Cas protein to cleave the opposite strand of the target sequence, releasing a second single-stranded DNA sequence; wherein the two single-stranded DNA sequences combine to form a double-stranded region, allowing the DNA polymerase to replicate the double-stranded region, thereby replicating each strand between the target fragments, ultimately achieving the replication of the target sequence.
[0085] The RT template sequences encoding the flap sequence can have different lengths without affecting the efficiency of the AE system. In one embodiment, each RT template sequence is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 nucleotides long. In another embodiment, each RT template sequence is no longer than 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, or 2000 nucleotides long. In some implementations, the length of each RT template sequence is 3-2000, 10-2000, 10-500, 15-500, 15-200, 15-50, or 15-30 nucleotides.
[0086] As previously stated, the RT template sequences share at least one complementary portion (“paired fragment”). In some embodiments, each paired fragment is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 nucleotides in length. In one embodiment, each paired fragment is no longer than 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 150, or 200 nucleotides, or 500 nucleotides. In some embodiments, each paired fragment is 3-200, 5-200, 10-50, 10-25, or 15-20 nucleotides in length.
[0087] Alternatively, in addition to the paired fragments, each RT template sequence may also include a portion that does not necessarily complement the paired fragments. In some implementations, such a non-complementary template sequence is located between the corresponding paired fragment and the crRNA / sgRNA or near the PBS sequence.
[0088] In some embodiments, the length of the non-complementary template sequence is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 nucleotides. In some embodiments, the length of the non-complementary template sequence does not exceed 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, or 2000 nucleotides. In some embodiments, the length of each non-complementary template sequence is 1-2000, 1-1000, 1-500, 10-500, 15-200, 15-50, or 15-30 nucleotides.
[0089] Optionally, in addition to the paired fragment, each RT template sequence may or may not contain a portion complementary to the genomic DNA of another pegRNA near the nick site. In some embodiments, each fragment paired with the genomic DNA has a length of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, or 100 nucleotides. In one embodiment, the length of each paired fragment does not exceed 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 150, or 200 nucleotides, or 500 nucleotides.
[0090] Currently available amplification (AE) technology can amplify target fragments of various lengths. In some implementations, the length of the target fragment is at least 10, 50, 100, 200, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 30000, 40000, 50000, 100,000, 1,000,000 (1Mb) or 100,000,000 (100Mb) or bp, or the entire chromosome.
[0091] The length of the copy is the distance between the two cut sites. Therefore, in some implementations, the two cuts located on either side of the target fragment are at least 10, 50, 100, 200, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 30000, 40000, 50000, 100000, 200000, 300000, 1000000, 1000,000, 100,000,000, 100,000,000, 100,000,000, or 100,000,000, or 100,000,000, or the entire chromosome. In some embodiments, the two sites are less than 100, 200, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 30000, 40000, 50000, 200000, 300000, 1000000, 1,000,000, or 100,000,000 or 100,000,000. In some embodiments, they are spaced 2 to 300,000 base pairs apart, preferably 10 to 50,000 base pairs apart.
[0092] The pegRNA disclosed in this article contains other elements used in conventional prime editing techniques.
[0093] Prime editing is a genome editing technology that allows modification of the genome of living organisms. It directly inserts new genetic information into target DNA sites. The technology uses a fusion protein composed of a catalytically impaired endonuclease (e.g., Cas9) and an engineered reverse transcriptase (RT), bound to a prime editing guide RNA (pegRNA). The pegRNA recognizes the target DNA site and provides new genetic information to replace nucleotides in the target DNA. Prime editing enables targeted insertions, deletions, and base transitions without creating double-strand breaks (DSBs) or donor DNA templates.
[0094] PegRNAs recognize the target nucleotide sequence for editing and encode new genetic information, replacing the target sequence. A pegRNA consists of an expanded single-stranded guide RNA (sgRNA) or crRNA containing a primer-binding site (PBS) and a reverse transcriptase (RT) template sequence. During genome editing, the 3' end of the nicked DNA strand is complementary to the pegRNA's primer-binding site. The RT template serves as a template for forming the edited genetic information and is reverse transcribed into single-stranded DNA containing the edited information. The sgRNA or crRNA portion includes a guide sequence (spacer) and an sgRNA / crRNA backbone structure. The guide sequence guides the prime editor to the target genomic site.
[0095] In some implementations, the pegRNA may also include a tail that (a) is capable of forming a hairpin or loop structure with itself, PBS, an RT template sequence, crRNA, or a combination thereof, or (b) contains a poly(A), poly(U), or poly(C) sequence, or an RNA-binding domain.
[0096] Cas proteins and reverse transcriptases can be provided as fusion proteins or separately and bound to the target site (e.g., as a complex). In some embodiments, the fusion protein comprises a fusion of a Cas protein (e.g., a nickase) and a reverse transcriptase. Nickases can be derived from conventional Cas9 proteins such as SpCas9, FnCas9, St1Cas9, St3Cas9, NmCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, VQR SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-NG, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9, or atCas9. An exemplary nickase is Cas9 H840A. Cas9 enzymes contain two nuclease domains that cleave double-stranded DNA sequences, with the RuvC domain cleaving the non-target strand and the HNH domain cleaving the target strand. By introducing the H840A substitution mutation into Cas9, replacing the histidine residue at position 840 with alanine, the HNH domain is inactivated. At this point, only the RuvC domain is active, allowing the catalytically inactivated Cas9 to perform single-strand cleavage; hence, it is called a nickase.
[0097] The traditional PE2 system consists of Cas9 nickase-RT and pegRNA. However, the Cas12 protein has not been used in prime editing, primarily due to the lack of a corresponding Cas12 nickase. Traditional pegRNA is not expected to function effectively with Cas12. Cas9 nickase introduces single-stranded cleavage, while Cas12 protein cleaves double-stranded sequences. Traditional pegRNA consists of a single-stranded guide RNA (sgRNA) or only a crRNA containing a guide sequence (spacer) and a backbone (scaffold), along with a reverse transcriptase (RT) template sequence and a primer binding site (PBS), presenting a spacer-scaffold-RTT-PBS (from 5' to 3') configuration. If the target genome is cleaved into a double-stranded state by the Cas12 protein, the RTT in the pegRNA will not be effectively used as an RT template.
[0098] In Cas9-based AE systems, the RT template is flanked by PBS and crRNA sequences. When using the Cas12 protein, the RT template and crRNA are located on either side of the PBS. In some implementations, the nickase may be a nickase of SpyCas9, SauCas9, NmeCas9, StCas9, FnCas9, CjCas9, AnaCas9, or GeoCas9.
[0099] The Cas protein can be a Cas12 protein, or Cas12a, Cas12b, Cas12f, and Cas12i, without any restrictions. Examples include AsCpf1, FnCpf1, SsCpf1, PcCpf1, BpCpf1, CmtCpf1, LiCpf1, PmCpf1, Pb3310Cpf1, Pb4417Cpf1, BsCpf1, EeCpf1, BhCas12b, AkCas12b, EbCas12b, and LsCas12b.
[0100] There are no restrictions on the reverse transcriptase, which may include human immunodeficiency virus (HIV) reverse transcriptase, Moloney murine leukemia virus (M-MLV) reverse transcriptase, and avian myeloblastosis virus (AMV) reverse transcriptase, as well as any reverse transcriptase that can function under physiological conditions.
[0101] Amplification editing can be accomplished by transfecting target cells. The transfection component consists of a pair of pegRNAs or pegRNA / sgRNA with a fusion protein, or by using Cas9 nuclease and reverse transcriptase respectively. Transfection is typically achieved by introducing a vector into the cell. In some implementations, the editing tool can be introduced directly into the cell as a protein and RNA or a complex thereof. Each molecule can be introduced individually or together; there are no restrictions on the method.
[0102] Vectors can be introduced into target host cells using known methods, including but not limited to transfection, transduction, cell fusion, electroporation, and liposome transfection. Vectors can contain various regulatory elements, such as promoters. In some embodiments, this document discloses an expression vector comprising any of the polynucleotides described herein, for example, an expression vector containing a polynucleotide encoding a fusion protein and / or pegRNA.
[0103] In some embodiments, replication occurs in cells containing DNA polymerase. Such replication can occur intracellularly, in vitro, or in vivo. For example, the cells can be prokaryotic, eukaryotic, plant, animal, mammalian, or human cells.
[0104] In some implementations, the cells are not in an active dividing state, meaning they are not undergoing cell division or chromosome replication. In some implementations, the cells are further engineered to express DNA polymerase.
[0105] Application of AE technology
[0106] DNA fragment amplification has wide applications in clinical and industrial fields.
[0107] For example, the duplication / amplification of certain genes may lead to cancer. Therefore, the AE method can generate animal disease models (such as cells, mice, fruit flies, etc.) formed by gene duplication and can be used to simulate this pathogenic feature, such as duplication-amplified diseases and oncogene copy number variations.
[0108] Gene duplication has important applications in plants. The amplification of certain genes or genomic sequences (such as disease resistance genes) provides plants with a natural advantage in resisting diseases, pests, and environmental stresses. Therefore, gene duplication technology can help plants achieve these advantages. Gene duplication also has applications in microorganisms, for example, in the generation of gene clusters in fungi, bacteria, or yeast.
[0109] Some human diseases (alpha-globin deficiency) are caused by haploid defects (haploid insufficiency). Therefore, duplication of these deleted genes can increase gene expression and restore the normal phenotype in patients. Chromosomal microdeletions are typically caused by deletions of 0.1 Mb to several Mb in a single chromosome copy. Amplification of genomic regions ranging from 0.1 to 10 Mb can be highly efficient. Amplification of genomic regions can compensate for gene loss due to microdeletions by amplifying corresponding sequences on sister chromosomes.
[0110] Therapeutic proteins (such as vaccines and antibodies) are already in mass production. Production efficiency may be limited by the copy number of coding sequences in host cells (e.g., in CHO cells), especially copies integrated into the genome. AE technology can conveniently and rapidly amplify these coding sequences, thereby increasing protein yield.
[0111] Cell cycle may be limited by telomere copy number or telomere length. In one embodiment, a method is provided to amplify telomeres or portions thereof using AE technology. Such amplification may increase cell viability or replication capacity, or extend the lifespan of an organism.
[0112] This disclosure also provides compositions, kits, and packaging for such applications. In some embodiments, a composition, kit, or packaging is provided for replicating a target sequence in the presence of a DNA polymerase, comprising: (a) a Cas protein and a reverse transcriptase; (b) a guide RNA (pegRNA1) for editing, comprising a first CRISPR RNA (crRNA1) and a reverse transcriptase template sequence (RT1); and (c) another guide RNA (pegRNA) for editing, comprising a second crRNA2 and a second RT2 template sequence, wherein (i) the first RT template sequence comprises a first complementary DNA; (ii) the second RT template sequence comprises a second complementary DNA; (iii) the first and second complementary DNA fragments are complementary to each other; and (iv) the first and second pegRNAs are capable of guiding the Cas protein to cleave at two sites flanking the target fragment, with cleavage occurring on the complementary strands.
[0113] This patent also provides another composition, kit, or package for replicating a target fragment of a target DNA sequence in the presence of DNA polymerase, comprising: (a) a Cas protein, (b) a first single-stranded guide RNA (sgRNA1), and (c) a second single-stranded guide RNA (sgRNA2), wherein the first sgRNA and the second sgRNA each have sequences complementary to the sequences flanking the spacer sequence in the target DNA sequence, and the two target sites are at least partially complementary.
[0114] Furthermore, this disclosure provides a composition, kit, or package for replicating a target fragment of a target DNA sequence in the presence of DNA polymerase, comprising: (a) a Cas protein and a reverse transcriptase, (b) a master editing guide RNA (pegRNA) including a first CRISPR RNA (crRNA) and a reverse transcriptase (RT) template sequence, and (c) a single guide RNA (sgRNA), wherein: (i) the RT template sequence includes a paired fragment, (ii) the pegRNA can guide the Cas protein to cleave at a first site near the target fragment of the target DNA sequence, thereby allowing the reverse transcriptase to extend the complementary strand of the target fragment using the RT template sequence as a template to generate a single-stranded DNA sequence, and (iii) the sgRNA can guide the Cas protein to cleave at a second site near the target fragment of the target DNA sequence, thereby releasing the complementary strand to form a second single-stranded DNA sequence.
[0115] The preceding sections will introduce each component in more detail; they will be combined here.
[0116] example
[0117] Example 1. Design and Experimentation of Augmentation Editor
[0118] A new system based on guided editing (PE) technology has been designed to replicate target fragments of target sequences. This new technology is called "Amplification Editing" (AE). Figure 1 ab is an example of the After Effects editing process.
[0119] In the example shown, two pegRNA molecules were used. (Reference) Figure 1 a. Each pegRNA, in addition to containing a CRISPRRNA (crRNA) / single-stranded guide RNA (sgRNA), also includes a reverse transcriptase (RT) template sequence and a primer binding site (PBS). The PBS can be complementary to the guide sequence (or "spacer sequence") in the crRNA / sgRNA, but is usually a few nucleotides shorter. When the guide sequence binds to the targeted genomic sequence and unwinds the DNA double helix, the PBS can bind to the complementary strand, initiating the reverse transcription process using the RT template sequence as a template.
[0120] Unlike pegRNAs in traditional guided editing technologies, in each pegRNA of an amplification editing system, the RT template sequence does not need to be homologous to the target genomic sequence. In some embodiments, the RT template sequence is preferably a sequence with low or no homology to the target sequence.
[0121] Instead, these two RT templates share a complementary part. For example, as Figure 1As shown in diagram a, in each pegRNA, the RT template comprises two parts: one is the RT-pairing fragment (or "complementary part") close to the crRNA / sgRNA, and the other is the RT fragment further away from the crRNA / sgRNA, or... Figure 16 As shown in the diagram, two paired (complementary) RT fragments have complementary sequences, allowing the DNA sequences reverse transcribed from them to pair with each other.
[0122] When the guide sequence binds to the target genome sequence and unwinds the DNA double helix, PBS binds to the complementary strand and initiates reverse transcription, using the RT template sequence as a template. Figure 1 As shown in step b, the two pegRNAs are designed to cleave at “PAM-out” sites on both sides of the target fragment, with cleavage occurring at two locations on the target fragment (step 110). The RT template is then used as a template for synthesizing single-stranded DNA, thereby introducing two mutually extending single-stranded DNA regions (step 120). The term “PAM-out” here refers to the PAM sequence being located outside the region sandwiched between the two target sites (sites identified by the spacer sequence).
[0123] The two single-stranded DNAs consist of a reverse transcription product from the RT fragment of pegRNA and a reverse transcription product from the distally paired (complementary) RT fragment. Due to their complementarity, these two distal fragments can hybridize to form a double-stranded region (step 130). This double-stranded region can then serve as the starting point for DNA polymerase.
[0124] Starting with a double-stranded region and using single-stranded genomic DNA as a template, a new DNA strand is synthesized, and the original genomic DNA is unwound between two gaps (steps 140-150). Finally, the target sequence (sequence A) is replaced with a new fragment comprising a first copy of sequence A, an insertion sequence based on two pegRNA RT templates, and a second copy of sequence A. Essentially, the amplification and editing process (A) replicates sequence A, and (B) inserts a new sequence based on an RT template between them.
[0125] This example further tested the amplification editing technique in the laboratory. We designed three types of PCR primers to examine the editing results of AE ( Figure 2 a) In-In PCR uses a pair of reverse primers located inside the amplification region and cannot amplify unedited sequences. Out-Out PCR uses a pair of primers targeting the outside of the amplification region. In-Out PCR uses one primer targeting the outside of the amplification region and one primer targeting the inserted sequence (generated from a 3' flap).
[0126] Then, we used two pairs of pegRNAs designed to replicate the 178 bp sequence at the VEGFA site and the 234 bp sequence at the HEK3 site in HEK293T cells. PCR bands of the expected size were detected in AE-edited cells by in-in PCR and in-out PCR, but not in control cells, indicating that the desired replication was achieved. Figure 2 b). Out-Out PCR detected two distinct bands in AE-edited cells, while only one band was detected in control cells, indicating that the target region was amplified. Figure 2 b). Analysis of the large bands in the Out-Out PCR using Sanger sequencing revealed precisely replicated sequences at the target site, with a 20 bp flap sequence inserted between them. Figure 2 c).
[0127] Example 2. Characterization and optimization of AE
[0128] To quantify the efficiency of AE, droplet digital PCR (ddPCR) was used. The primer design was the same as in-inPCR, while the probe was designed to target sequences in the replication region. Figure 3 d).
[0129] The paired 3' single-stranded DNA regions ranged in length from 10 bp to 100 bp and were complementary to each other. We examined different replication sizes at the VEGFA and C-MYC sites. For replications ranging from ~200 bp to ~8 Kb, 3' single-stranded DNA lengths of 30–50 bp showed high efficiency. In HEK293T cells at the VEGFA site, the replication efficiency for 178 bp exceeded 60%, for 1 Kb it exceeded 40%, and for 8 Kb it was approximately 20%. Figure 3 a). The trend at C-MYC sites is similar ( Figure 3 b). Paired 3' single-stranded DNAs can be partially complementary at their 3' ends (i.e., "overlap"). We used a 30 bp overlap and extended the length of the 3' single-stranded DNA. Our data show that 30 bp of 3' single-stranded DNA exhibits higher efficiency in ~200 bp replication, while 50 bp of 3' single-stranded DNA shows the best efficiency in 8 kb replication. Figure 3 c). We also investigated the effects of different GC contents in 3' single-stranded DNA, and found that 3' single-stranded DNA with GC contents ranging from 30% to 80% effectively promoted replication. Figure 4 a). To explore the feasibility of simultaneous replication at multiple sites, we co-transfected plasmids encoding two or three pairs of pegRNAs into HEK293T cells. The replication efficiency at each site was comparable to that of single-site editing via multiple editing. Figure 4 b).
[0130] To determine the purity of the edited results, we performed deep sequencing on the left and right interfaces of the replicated region, as well as the middle interface containing the flap insertion. Figure 5 a). The purity at all three interfaces of VEGFA-178bp replication and HEK3-234bp replication is close to 100%. Figure 5 b). The purity at the ~200bp to ~8Kb replication interfaces of VEGFA and C-MYC sites was also close to 100%, indicating that AE replication has high precision. Figure 5 ce).
[0131] Example 3. AE is active at different endogenous sites in various cell lines.
[0132] AE has been edited at the three endogenous sites mentioned above. We will further investigate the replication ability of AE at other endogenous sites, including AAVS1, RUNX1, and HEK4, and quantify the replication efficiency by ddPCR. We found that the replication rate of ~200 bp ranged from 56.3% to 68.5%, the replication rate of ~1-2 kb ranged from 28.5% to 47.8%, and the replication rate of ~7-9 kb ranged from 20.4% to 33.3%. These results are all based on the editing of AE in HEK293T cells. Figure 6 a).
[0133] Next, we examined whether AE was active in other cell lines, such as human Huh-7 cells, human K562 cells, human U2OS cells, and mouse N2a cells. Using ddPCR, we found that the replication efficiency of AE ranged from 1.7% to 34.6% in Huh-7 cells, 18.9% to 66.1% in K562 cells, 4.9% to 25.2% in U2OS cells, and 20.0% to 85.5% in N2a cells. Figure 6 We validated replication events in these cells using in-in and in-out PCR. Figure 7 ac).
[0134] Example 4. AE generates a series of repeating events.
[0135] Next, we investigated whether the AE could replicate DNA fragments smaller than 150 bp. To this end, we designed the AE to replicate DNA fragments in the range of 20-130 bp, with editing efficiencies ranging from 28.6% to 52.3%. Figure 8(ab). It is important to note that when the replicated fragment size is 20 bp, the spacer sequences of the two pegRNAs are complementary, and the primer binding site (PBS) overlaps with the spacer sequence of the other pegRNA. Figure 8 a). Therefore, each 3' single-stranded DNA generation process may occur sequentially. We amplified the edited product of the small fragment using out-out PCR and performed deep sequencing analysis. The results showed that the replicated sequence was consistent with the expected sequence (a). Figure 8 c). Unsurprisingly, we also found replication fragments with small deletions, possibly because the two cuts were too close together, making it difficult for the 20bp replication to proceed smoothly. More interestingly, we found three sets of tandem repeats consisting of a 20bp sequence and an insertion sequence interval, indicating that replication may occur multiple times. Figure 8 ac, Figure 9 a).
[0136] Subsequently, we amplified the 234bp replication region of the HEK3 site in single-cell clones using out-out PCR. Gel electrophoresis analysis showed that the number of tandem repeats ranged from two to nine, and these repeat sequences were validated by Sanger sequencing. Figure 9 (bc). Although theoretically more tandem repeats might occur, their length may exceed the detection range of genomic PCR. After each round of replication, the recognition site of each pegRNA remains intact, thus supporting the occurrence of the next round of replication. When the mechanism of the next round of replication is consistent with that of the first round, 2 n Repeating sequences of multiples (e.g., 2A, 4A, 8A) Figure 9 (a. Left figure). Furthermore, 3' single-stranded DNA may also pair with the insertion fragment generated in the previous replication cycle, thus generating varying numbers of repeats. Figure 9 (Right image of a).
[0137] Consistent with the above observations, ddPCR analysis revealed repeatability exceeding 100% (relative to the reference gene) at the VEGFA and RUNX1 sites in K562 cells. Figure 10 a). To further evaluate the efficiency of AE sequential replication, we fragmented the generated tandem repeats using restriction endonucleases. In the fragmented samples, the sequentially replicated fragments were distributed into different ddPCR droplets, thereby improving replication efficiency. Figure 10 b). In contrast, samples with only one replication event showed no significant difference in results regardless of whether fragmentation was performed. In K562 cells, the replication efficiency of VEGFA-178bp reached 808.7%, while the replication efficiency of VEGFA-1Kb increased from 66.1% to 199.0%. After fragmentation, the replication efficiency was significantly higher than that of the unfragmented sample. Figure 10 c). A similar trend was also observed at the HEK3 site in K562 cells. Figure 10 d). In HEK293T cells, replication of fragments approximately 200 bp and 1 kb in size also showed a consistent trend ( Figure 3 fg). When the size of the replicated fragment reached 8Kb, no significant difference was observed before and after fragmentation in K562 cells, while in HEK293T cells, the replication efficiency was slightly improved after fragmentation (fg). Figure 10 These data indicate that the frequency of successive replication decreases as the size of the replicated fragment increases.
[0138] Example 5. Functional determination and potential applications of AE
[0139] To demonstrate that AE can restore gene expression through replication, we established a stable cell line with interference from a small deletion (53 bp) in the GFP sequence. A pair of pegRNAs with PAM-out design were used to replicate the GFP region and insert the small deletion fragment to restore GFP expression. Figure 11 a). AE successfully produced ~30% GFP-positive cells through gene replication ( Figure 11 b).
[0140] Alpha-thalassemia is a common blood disorder, usually caused by mutations in the HBA1 and HBA2 genes (two nearly identical genes). The most common form is a 3.7 kb deletion (-α3.7) in both HBA1 and HBA2 genes, resulting in a fusion HBA gene that is identical to the HBA1 gene. We constructed the -α3.7 genotype in HEK293T cells using CRISPR technology and replicated the fusion HBA gene using AE technology, thereby repairing the -α3.7 deletion. Figure 11 c). The replication efficiency of the HBA gene reached 16.4% ( Figure 11 d).
[0141] To verify whether AE treatment can induce functional changes in endogenous genes, we applied AE to stem-loop region amplification of microRNA-21 (miR-21). Figure 12 a). In HEK293T and K562 cells, the replication efficiency of the miR-21 stem-loop region was 94.6% and 189.0%, respectively, indicating that this region underwent multiple rounds of expansion. Figure 12 b). Compared with control cells, after 7 days of AE treatment, the expression levels of miR-21 in HEK293T and K562 cells were significantly increased by 5.9-fold and 6.2-fold, respectively. Figure 12c). Meanwhile, the expression levels of miR-21 target genes decreased, indicating that this replication process is functionally efficient. Figure 12 d).
[0142] Example 6. AE can copy from 30Kb to 100Mb.
[0143] We investigated whether AE (anti-replication) could replicate large genomic regions, ranging from 30 kb to 100 Mb. By designing multiple pairs of pegRNAs with intervals of 30, 60, and 100 kb and applying them to chromosome 6 (Chr 6) or chromosome 9 (Chr 9), we detected the presence of replication events using in-in PCR. To obtain the longest possible amplified sequences, we performed in-in PCR experiments using multiple primer pairs. Control samples failed to show any bands in in-in PCR, while AE-treated samples showed the expected PCR products ranging in size from 3.6 to 5.0 kb. Figure 13 a). These PCR products were subjected to Sanger sequencing using multiple primers, and the results showed that the replicated sequences were correct. Figure 13 a). The replication efficiency, as determined by ddPCR, ranged from 7.5% to 31.7% for 30-60 kb replication and 28.8% for 100 kb replication. Figure 13 c).
[0144] Inspired by the above results, we further explored the possibility of AE replicating the genome at the Mb level. We first tested the replication efficiencies of 1Mb and 3Mb on chromosomes 6, 9, and 12. The replication efficiencies for the 1Mb region ranged from 9.0% to 27.7%, and for the 3Mb region, they ranged from 2.7% to 7.6%. Figure 13 d). To confirm the presence of Mb-level duplication, we examined the copy number of three distinct genes within the duplicative region of chromosome 12. In control cells, the copy number of these genes was approximately 3, indicating triploidy in this region. In AE-treated monoclonal cells, the copy number of these genes was approximately 4, indicating duplication in this region of chromosome 12. Figure 13 b).
[0145] We then explored the feasibility of replicating 10 to 100 Mb genomic regions. The efficiency of replicating chromosome-scale regions ranged from 0.55% to 2.5% on chromosomes 6 and 9. Figure 13 f). It is noteworthy that chromosome 6 is 172 Mb in size, while AE successfully replicated a 100 Mb region, achieving an efficiency of 1.1% (f). Figure 13 f).
[0146] To further confirm the occurrence of large-scale genome duplication, we used fluorescence in situ hybridization (FISH) with DNA probes targeting the duplication region of the genome. The STAT6 gene is located in a 3Mb duplication region on chromosome 12. We used previously validated DNA probes targeting the STAT6 gene to visualize this duplication region, and used DNA probes targeting the centromere region as controls. Figure 14 a). In wild-type cells, two red dots surround a green dot, while AE-treated monoclonal cells show four red dots surrounding a green dot, indicating that the STAT6 region has been replicated. Figure 14 a). On chromosome 6, a 100 Mb replication region occupies most of its long arm (110 Mb). We used a previously validated DNA probe targeting the ESR1 gene to locate this replication region (a). Figure 14 b). AE-edited cells exhibited a significantly elongated long arm of chromosome 6 (from 110 Mb in the control group to 210 Mb in the edited group), and the replicated region within this area was confirmed. Figure 14 b). These data collectively confirm the existence of AE chromosome arm replication in the 1–100 Mb range within cells.
[0147] Example 7. Various PAM-out methods used for DNA replication
[0148] In this study, we explored whether DNA replication could be achieved using methods other than paired pegRNAs. When the Cas9 / sgRNA complex targets DNA, the spacer sequence of the sgRNA binds to its complementary sequence, resulting in a free, untargeted DNA strand, forming a small 3' single-stranded DNA. We first examined the replication efficiency of 3' single-stranded DNA with complementary sequences of 10 bp or less. The results showed that the replication efficiencies of 3 bp and 8 bp complementary 3' single-stranded DNA were 8.5% and 12.6%, respectively. Figure 15 a). Interestingly, in samples without any complementary 3' single-stranded DNA generated by the RTT of pegRNA, the replication efficiency was 1.9%. This low replication frequency may be due to the slight homology at the sgRNA binding site. Based on this observation, we suggest that pegRNA / sgRNA or sgRNA / sgRNA combinations could be used to design shared complementary sequences in the PAM-out direction to enable the 3' single-stranded DNA to anneal ( Figure 16 a). In addition, it is also feasible to design a 3' single-stranded monoclonal genomic sequence that is partially complementary to another pegRNA cleavage site and partially complementary to each other.
[0149] We further investigated the effect of pegRNA / sgRNA combinations with complementary sequences on DNA replication. Specifically, we designed pegRNAs with an 8 bp or 0 bp complementary sequence between their RTT and the sgRNA target site. Experimental results showed that paired pegRNAs achieved replication efficiencies of 32.8%–71.0% at the RUNX1, VEGFA, and AAVS1 gene sites, while the replication efficiency of pegRNA / sgRNA combinations with 8 bp complementary sequences ranged from 3.9% to 24.3%. Figure 16 b). Although the pegRNA / sgRNA combination with a 0bp complementary sequence was less efficient than the combination with an 8bp complementary sequence, they still exhibited significant editing effects, possibly due to the 3-4bp microhomology between the RTT and the sgRNA target site. Figure 16 b). We analyzed the replication interface of the pegRNA / sgRNA combination using deep sequencing. The results showed that the purity of the combination samples with an 8bp complementary sequence ranged from 53.9% to 87.9%, while the purity of the samples with a 0bp complementary sequence ranged from 0% to 9.5%. Figure 16 c). These results demonstrate that precise complementary sequence design is key to maintaining efficient and accurate editing in AE.
[0150] Next, we experimented with a pair of sgRNAs and designed 4-8 bp complementary sequences in the PAM-out direction to target genomic sequences near the cut. Figure 16 a). Through in-in PCR analysis, we found that bands were clearly observed in samples treated with sgRNAs and Cas9 nuclease, while no bands were observed in the control group. Figure 16 d). Second-generation data at the edit group interface showed that the purity of the 8bp complementary sequence was as high as 96.6%, indicating that high-precision DNA replication can still be achieved even without the participation of nCas9-RT and pegRNA. Figure 16 e).
[0151] ***
[0152] The scope of this patent is not limited to the specific embodiments described. These examples are merely single instances of this patent, and any functionally equivalent components or methods fall within the scope of this patent. Those skilled in the art can make various modifications and variations to the methods and components disclosed herein without departing from the spirit and scope of this disclosure. Therefore, any modifications and variations to the content described herein, as long as they conform to the scope of the appended claims and their equivalents, are within the protection scope of this patent.
[0153] All publications and patent applications mentioned in this specification are incorporated herein by reference to the same extent that each individual publication or patent application is specifically and individually indicated to be incorporated by reference.
Claims
1. A method for replicating a target fragment of a target DNA sequence in vitro in the presence of a DNA polymerase, comprising contacting the target DNA sequence with the following components: (a) Cas protein and reverse transcriptase; (b) The first PE guide RNA (pegRNA), which includes the first CRISPRRNA (crRNA) and the first reverse transcriptase (RT) template sequence; (c) A second pegRNA, which includes a second crRNA and a second RT template sequence; in: (i) The first RT template sequence includes the first paired segment; (ii) The second RT template sequence includes a second paired segment; (iii) The first paired segment and the second paired segment are complementary; (iv) The first and second pegRNAs guide the Cas protein to cleave the two strands of the target DNA at two locations flanking the target fragment within the target DNA sequence. Therefore, it is permitted that: (1) The reverse transcriptase uses the first and second RT template sequences as templates to extend the two complementary strands of the target fragment, generating two single-stranded DNA sequences. (2) Two single-stranded flap sequences form a double-stranded region, which allows DNA polymerase to expand the double-stranded region to replicate each strand of the target fragment, thereby replicating the target fragment and inserting an insert fragment between the two replicated target fragments, wherein one strand of the insert fragment includes the reverse complementary sequence of the first fragment, the first paired fragment, and the second fragment. The method is capable of repeating replication, generating tandem repeat sequences in the target DNA sequence; The Cas protein is the nCas9 protein; The target DNA sequence is located in vitro; the target DNA sequence is located inside mammalian cells; The length of the complementary fragment is 3bp, 8bp, or 10-100bp; The size of the repeated copy fragment is less than 100Mb.
2. The method of claim 1, wherein the first pegRNA further comprises a first primer binding site PBS and a first spacer sequence, and the second pegRNA further comprises a second PBS and a second spacer sequence, such that the pegRNA can guide the Cas protein to two positions on either side of the target fragment and initiate reverse transcription.
3. The method of claim 2, wherein the length of the first and second RT template sequences is 0 to 2000 nucleotides.
4. The method of claim 2, wherein the length of the first and second RT template sequences is 15 to 500 nucleotides.
5. The method of claim 2, wherein the length of the first and second paired fragments is 0 to 1000 nucleotides.
6. The method of claim 2, wherein the length of the first and second paired fragments is 3 to 200 nucleotides.
7. The method of claim 2, wherein the length of the first and second paired fragments is 3 to 50 nucleotides.
8. The method of claim 2, wherein the length of the first and second paired fragments is 30-100 nucleotides.
9. The method of any one of claims 2-8, wherein each of the first and second RT template sequences further comprises a non-complementary template sequence that is not complementary to the other sequence, wherein each non-complementary template sequence is located between the corresponding paired fragment and crRNA, or between the corresponding paired fragment and PBS.
10. The method of claim 9, wherein the length of each non-complementary template sequence is from 1 to 2000 nucleotides.
11. The method of claim 9, wherein the length of each non-complementary template sequence is from 1 to 1000 nucleotides.
12. The method of claim 9, wherein each non-complementary template sequence is 1 to 500 nucleotides in length.
13. The method of claim 2, wherein the two sites on either side of the target fragment are 2 to 1 billion base pairs apart.
14. The method of claim 2, wherein the two sites on either side of the target fragment are 10 to 5 million base pairs apart.
15. The method of claim 2, wherein each RT template sequence further comprises an additional sequence adjacent to the paired fragment, wherein the two additional sequences are complementary to the target DNA sequence and are at least partially complementary to each other.
16. A method for replicating a target fragment of a target DNA sequence in vitro in the presence of a DNA polymerase, comprising contacting the target DNA sequence with the following components: (a) Cas protein; (b) First sgRNA or tracrRNA; (c) A second sgRNA or tracrRNA, wherein the first sgRNA or tracrRNA and the second sgRNA or tracrRNA each have sequence complementarity with target sites flanking the target fragment in the target DNA sequence, and the two target sites have at least partial complementarity. wherein: The first sgRNA or tracrRNA, in the presence of the Cas protein, binds to one strand and cleaves the complementary strand of the target site, releasing the complementary strand as the first single-stranded DNA; the second sgRNA or tracrRNA, in the presence of the Cas protein, binds to one strand and cleaves the complementary strand of the target site, releasing the opposing strand as the second single-stranded DNA; the first single-stranded DNA and the second single-stranded DNA bind to form a double-stranded region, allowing DNA polymerase to expand this double-stranded region to replicate the target sequence, thereby replicating the sequence between the two target sites; The Cas protein is the nCas9 protein; The target DNA sequence is located in vitro; the target DNA sequence is located inside mammalian cells; The length of the complementary fragment is 3bp, 8bp, or 10-100bp; The size of the repeated copy fragment is less than 100Mb.
17. The method according to claim 16, wherein partial complementarity comprises complete complementarity of at least three consecutive nucleotides.
18. The method according to claim 16, wherein partial complementarity comprises complete complementarity of at least four consecutive nucleotides.
19. The method according to claim 16, wherein partial complementarity comprises complete complementarity of at least 5 consecutive nucleotides.
20. The method according to claim 16, wherein partial complementarity comprises complete complementarity of at least 6 consecutive nucleotides.
21. The method according to claim 16, wherein partial complementarity comprises complete complementarity of at least 7 consecutive nucleotides.
22. The method according to claim 16, wherein partial complementarity comprises complete complementarity of at least eight consecutive nucleotides.
23. A method for replicating a target fragment of a target DNA sequence in vitro in the presence of a DNA polymerase, comprising contacting the target DNA sequence with the following components: (a) Cas protein and reverse transcriptase; (b) A pegRNA comprising a first CRISPR RNA (crRNA) and a reverse transcriptase (RT) template sequence; (c) One sgRNA or tracrRNA; in: (i) The RT template sequence includes a paired segment; (ii) pegRNA guides the Cas protein to cut the target DNA sequence at the first position near the target fragment, thereby allowing the reverse transcriptase to extend the opposite strand of the target fragment using the RT template sequence as a template to generate a single-stranded DNA sequence; (iii) sgRNA or tracrRNA guides the Cas protein to cut the target DNA sequence at the second position closest to the target fragment, thereby releasing the opposing strand as a second single-stranded DNA sequence; thus, the two single-stranded DNA sequences form a double-stranded region, which allows DNA polymerase to expand the double-stranded region to replicate the target fragment, thereby replicating the target fragment; The Cas protein is the nCas9 protein; The target DNA sequence is located in vitro; the target DNA sequence is located inside mammalian cells; The length of the complementary fragment is 3bp, 8bp, or 10-100bp; The size of the repeated copy fragment is less than 100Mb.
24. The method of claim 23, wherein the cell is a dividing cell.
25. The method of claim 23, wherein the cell is not a dividing cell.
26. The method of claim 23, wherein the Cas protein is the SpCas9 cleavage enzyme.
27. The method of claim 23, wherein each pegRNA comprises a first or second crRNA in the 5' to 3' orientation, a first or second paired fragment, a first or second fragment, and a first or second PBS.
28. The method of claim 2 or 23, wherein the cleavage enzyme is a Cas9 protein containing an inactive HNH domain that cleaves the target strand.
29. The method of claim 2 or 23, wherein the nicking enzyme is a nicking enzyme of SpCas9, FnCas9, St1Cas9, St3Cas9, NmCas9, SaCas9, AsCpf1, LbCpf1, FnCpf1, VQR SpCas9, EQR SpCas9, VRER SpCas9, SpCas9-NG, xSpCas9, RHA FnCas9, KKH SaCas9, NmeCas9, StCas9, CjCas9 or atCas9.
30. The method of claim 2 or 23, wherein the first pegRNA or the second pegRNA further comprises a tail that (a) is capable of forming a hairpin structure or loop with itself, PBS, an RT template sequence, crRNA, or a combination thereof, or (b) comprises a poly(A), poly(U), or poly(C) sequence, or an RNA-binding domain.
31. The method according to claim 2, claim 16 or claim 23, wherein the reverse transcriptase is M-MLV reverse transcriptase or a reverse transcriptase capable of functioning under physiological conditions.
32. The method of claim 31, wherein the Cas protein and the reverse transcriptase are provided as nucleotide sequences encoding their respective proteins, or as proteins.
33. The method of claim 32, wherein each pegRNA is provided as a recombinant DNA or RNA molecule encoding the pegRNA.
34. The method of claim 33, wherein the copied target fragment and the inserted fragment are further copied.