A method for predicting potential off-target events in genome editing processes using prime editing systems.
The method using prime editor protein and tpegRNA in the prime editing system addresses the challenge of predicting off-target effects, ensuring safer and more precise genome editing by analyzing tagmentation information.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-02-08
- Publication Date
- 2026-03-25
Smart Images

Figure 0007835414000073 
Figure 0007835414000074 
Figure 0007835414000075
Abstract
Description
[Technical Field]
[0001] This application relates to a method for predicting off-target effects in a prime editing system, which is a type of gene editing system. [Background technology]
[0002] Genome editing using the CRISPR / Cas system is an actively researched field. Various studies have been conducted on modified guide RNAs and other related materials, and Cas proteins for gene manipulation have been developed. However, there are still problems with gene editing methods using the CRISPR / Cas system. Many of the problems arising from CRISPR / Cas gene manipulation methods have motivated the development of more sophisticated genome editing technologies. Such motivation led to the development of base editing, a more sophisticated genome editing technique. However, the scope of application for base editing is still limited.
[0003] David R. Liu and colleagues developed prime editing technology, a "search-and-replace" genome editing technique effective in inducing insertions, deletions, all 12 base pair conversions, and combinations thereof in the genome.
[0004] A new genome editing platform called "prime editing" has been developed by David R. Liu et al., but methods for predicting potential off-target effects that may occur with prime editing have not yet been developed. The development of prime editing, a new genome editing platform, necessitates the development of new methods for predicting off-target effects suitable for prime editing systems. [Overview of the project] [Problems that the invention aims to solve]
[0005] Off-targets that occur in the gene editing process cause strong side effects. Accordingly, various methods for predicting off-targets have been developed. However, the methods known to date have been developed to target conventional CRISPR / Cas systems, and thus it is difficult to apply them to new gene editing systems such as prime editing systems. Therefore, the present application discloses a method or system for predicting off-targets in a prime editing system, which has been developed to target a prime editing system.
Means for Solving the Problem
[0006] Some embodiments of the present application are methods for predicting off-targets that occur in the process of genome editing by a prime editing system, comprising: (a) obtaining the engineered cells, where the engineered cells contain engineered genomic DNA, the engineered genomic DNA contains a tag sequence, and the engineered genomic DNA is involved in a prime editor protein and tpegRNA, as follows: (i) contacting the genomic DNA with a prime editor protein and tpegRNA (tagmentation pegRNA), where the prime editor includes a Cas protein and a reverse transcriptase, and the tpegRNA includes a spacer and an extension region containing a tag template; (ii) inserting the tag sequence into the genomic DNA by a reverse transcription process performed by the reverse transcriptase using the tag template of the tpegRNA as a template for reverse transcription; and being generated through a process comprising: (b) analyzing the engineered genomic DNA to obtain information on tagmentation, where the information on tagmentation includes information on the region of the genomic DNA into which the tag sequence has been inserted. A method is provided comprising the above.
[0007] In a particular embodiment, a method for predicting off-target events is: The step of obtaining off-target information based on tagging information, where off-target information includes information on whether or not off-target candidates exist, and, if off-target candidates exist, information on the region of the off-target candidates. It can also be equipped with.
[0008] In a particular embodiment, a method for predicting off-target events is: This stage involves verifying on-target information and comparing it with tagging information. It can also be equipped with.
[0009] In a particular embodiment, a method for predicting off-target events is: This stage involves reviewing on-target information and comparing it with tagging information to determine whether or not off-target candidates exist. It can also be equipped with.
[0010] In certain embodiments, the tag sequence may be inserted into a region of genomic DNA designated by a tpegRNA spacer.
[0011] In certain embodiments, the region into which the tag sequence is inserted may be related to (correspond to) an off-target candidate region or an on-target region.
[0012] In certain embodiments, information regarding the region in which the tag sequence is inserted may include information regarding the chromosome in which the tag sequence is located, and information regarding the region within the chromosome in which the tag sequence exists.
[0013] In certain embodiments, information regarding candidate off-target regions may include information regarding the chromosome in which the off-target candidate is located, and the region within the chromosome in which the off-target candidate resides.
[0014] In certain embodiments, information regarding tagmentation may further include the insertion ratio of the tag sequence by the region in which the tag sequence is inserted.
[0015] In certain embodiments, off-target information may further include off-target prediction scores for off-target candidates.
[0016] In certain embodiments, off-target information may further include the number of predicted off-target candidates.
[0017] In certain embodiments, the manipulated cells can be obtained by a method comprising the steps of contacting the cells with a prime editor protein or a nucleic acid encoding the prime editor protein, and tpegRNA or a nucleic acid encoding the tpegRNA.
[0018] In certain embodiments, the engineered cells may be obtained by a method comprising the steps of introducing a prime editor protein or a nucleic acid encoding the prime editor protein, and tpegRNA or a nucleic acid encoding the tpegRNA into the cells.
[0019] In certain embodiments, the method may further comprise a step of obtaining DNA from manipulated cells, the step of which is performed prior to (b).
[0020] In certain embodiments, the tpegRNA may include a spacer; a gRNA core; and an elongation region comprising a primer binding site, a tag template, and a reverse transcription template.
[0021] In certain embodiments, the reverse transcription template of tpegRNA may include an editing template and a homology region.
[0022] In certain embodiments, the manipulated genomic DNA may include editing.
[0023] In certain embodiments, the spacer, gRNA core, and elongation region may be positioned in the order of spacer, gRNA core, and elongation region in the 5'→3' direction.
[0024] In certain embodiments, the tag template may be located between the primer binding site and the reverse transcription template in the extension region.
[0025] In certain embodiments, the tpegRNA may further include a 3' engineering region containing an RNA protection motif.
[0026] In a particular embodiment, a method for predicting off-target events is: Information about a specified cell, information about a specified pegRNA, and information about a specified prime editor protein. A stage of verifying a predetermined prime editing system, which includes one or more of the following: It can also be equipped with.
[0027] In certain embodiments, the given cells may be different from the cells used in the method for predicting off-target behavior.
[0028] In a particular embodiment, the spacer sequence of tpegRNA may be the same as the spacer sequence of a given pegRNA, and the primer binding site sequence of tpegRNA may be the same as the primer binding site sequence of a given pegRNA.
[0029] In a particular embodiment, the spacer sequence of tpegRNA may be the same as the spacer sequence of a predetermined pegRNA, the primer binding site sequence of tpegRNA may be the same as the primer binding site sequence of a predetermined pegRNA, and the reverse transcription template sequence of tpegRNA may be the same as the reverse transcription template sequence of a predetermined pegRNA.
[0030] In a particular embodiment, the prime editor protein used in the method for predicting off-target reactions may be the same as or different from a given prime editor protein.
[0031] In certain embodiments, the length of the tag template may be between 5 and 60 nt.
[0032] In certain embodiments, the length of the tag template may be between 10 and 50 nt.
[0033] In certain embodiments, the prime editor protein may be a PE nuclease containing a Cas protein having double-strand cleavage activity.
[0034] In certain embodiments, the prime editor protein may be a PEmax nuclease.
[0035] In certain embodiments, the Cas protein contained in the prime editor protein may be niccasse.
[0036] In certain embodiments, the prime editor protein may be the PE2 prime editor protein.
[0037] In certain embodiments, the manipulation of the DNA genome may further involve dnMLH1, gRNA, and one or more of additional Cas proteins and additional prime editor proteins.
[0038] In certain embodiments, (b) may include the step of specifically analyzing the manipulated genomic DNA tag.
[0039] In certain embodiments, (b) may include the step of sequencing the manipulated genomic DNA.
[0040] In a particular embodiment, (b) is The steps include: generating a tag-specific library from manipulated genomic DNA; generating an amplified tag-specific library by amplifying the tag-specific library; and sequencing the amplified tag-specific library. Includes.
[0041] Some embodiments of the present invention are methods for predicting off-target events that occur in the genome editing process by a prime editing system, (a) A step of preparing a population of cells, which includes one or more manipulated cells. Here, the manipulated cells contain manipulated genomic DNA, the manipulated genomic DNA contains a tag sequence, and the manipulated genomic DNA involves prime editor protein and tpegRNA, as follows: (i) A process of contacting genomic DNA with a prime editor protein and tpegRNA (tagmentation pegRNA), where the prime editor protein comprises a Cas protein and a reverse transcriptase, and where the tpegRNA comprises an elongation region containing a spacer and a tag template. (ii) The process of inserting a tag sequence into genomic DNA, where the insertion of the tag sequence is achieved through a reverse transcription process performed by reverse transcriptase using a tpegRNA tag template as a reverse transcription template; Includes, generated through a process, (b) A step of obtaining tagment information by analyzing the results obtained through a process that includes sequencing the manipulated genomic DNA of one or more manipulated cells, Here, tagmentation information includes information about one or more sites where each tag sequence is inserted; and (c) The step of obtaining off-target information based on tagment information, Here, off-target information includes information on whether or not off-target candidates exist, and information on the regions of one or more off-target candidates. To provide a method that includes [this].
[0042] Some embodiments of this application are Spacer; gRNA core; and extension region containing tag template This provides tagged pegRNA (tpegRNA).
[0043] In certain embodiments, the elongation region containing the spacer, gRNA core, and tag template may be positioned in the order of the spacer, gRNA core, and tag template in the 5'→3' direction.
[0044] In certain embodiments, the extension region may include a tag template, a primer binding site, and a reverse transcription template.
[0045] In certain embodiments, the tag template may be located between the primer binding site and the reverse transcription template in the extension region.
[0046] In certain embodiments, the reverse transcription template may be located between the tag template and the primer binding site.
[0047] In certain embodiments, the primer binding site, tag template, and reverse transcription template may be located in the extension region in the 5'→3' direction in the order of reverse transcription template, tag template, and primer binding site.
[0048] In certain embodiments, the reverse transcription template may include an editing template and homology regions.
[0049] In certain embodiments, the tag template may have a length of 5 to 60 nt.
[0050] In certain embodiments, the tag template may have a length of 10 to 50 nt.
[0051] In certain embodiments, the tpegRNA may further include a 3' engineering region containing an RNA protection motif.
[0052] In certain embodiments, the RNA protection motif may have a length of 10 nt to 60 nt.
[0053] In certain embodiments, the tpegRNA may have a length of 100 nt to 350 nt.
[0054] Some embodiments of the present invention are compositions for predicting off-target events that may occur in the genome editing process by a prime editing system, tpegRNA; and Prime Editor containing Cas protein and reverse transcriptase The present invention provides a composition containing the following:
[0055] Favorable effects Methods for predicting off-target events in prime editing systems, according to some embodiments of the present invention, utilize the molecular mechanisms of prime editing systems and thus offer numerous advantages in predicting off-target events in prime editing systems compared to other known off-target prediction methods. [Brief explanation of the drawing]
[0056] [Figure 1] Examples of the structures of conventional guide RNA (gRNA), prime editing guide RNA (pegRNA), and tagmentation pegRNA (tpegRNA) are shown. [Figure 2] With regard to exemplary embodiments of tpegRNA, the tpegRNA shown in Figure 2 comprises an elongation region including a DNA synthesis template, a tag template, and a primer binding site (PBS); [Figure 3]Regarding exemplary embodiments of tpegRNA, the tpegRNA shown in Figure 3 includes a primer binding site, a tag template, an editing template, and an elongation region containing a homology region. [Figure 4] Regarding the tag insertion mechanism using tpegRNA in the off-target prediction system of this application, specifically, examples of DNA molecules in which nicks have been created at on-target or off-target candidate sites, and prime editor protein / tpegRNA complexes that induce nicks are shown. [Figure 5] Regarding the tag insertion mechanism using tpegRNA in the off-target prediction system of this application, specifically, the primer binding site of tpegRNA functions as a primer and annealing site for genomic DNA, and subsequently, it indicates a region where reverse transcription is performed by reverse transcriptase (RT) using a tag template or the like as a template. [Figure 6] The present invention relates to a tag insertion mechanism using tpegRNA in an off-target prediction system. This mechanism involves adding a tag sequence to an endogenous DNA strand (3' DNA flap) by reverse transcription, followed by a process including removal of the 5' DNA flap and DNA repair, to position the tag sequence and its complementary sequence at an on-target site or off-target candidate site of genomic DNA. [Figure 7] This document illustrates an exemplary process of tagmentation of Prime Editor sequencing (TAPE-seq), which is the off-target prediction system of this application. [Figure 8] The results of the tag sequence insertion rate during the incubation period are shown. [Figure 9] This shows a map of the green fluorescent protein (GFP)-piggyBac vector. [Figure 10] The enrichment results for GFP-positive cells (HEK293T) are shown. [Figure 11]The enrichment results for GFP-positive cells (HEK293T) are shown. [Figure 12] The enrichment results for GFP-positive cells (HeLa) are shown. [Figure 13] The enrichment results for GFP-positive cells (HeLa) are shown. [Figure 14] The enrichment results for GFP-positive cells (K562) are shown. [Figure 15] The enrichment results for GFP-positive cells (K562) are shown. [Figure 16] This shows the number of candidate off-target regions identified by TAPE-seq, categorized by incubation period after transfecting HEK294 T cells with HEK4(+2G→T)pegRNA. [Figure 17] The experimental results for determining the optimal amount of piggyBac vector for cotransfection with transposase plasmids are shown, and Figure 17 is a graph showing the number of copies of piggyBac constructs found in cells by quantitative polymerase chain reaction (PCR) in relation to the amount of piggyBac plasmid (PB plasmid). [Figure 18] The experimental results for determining the optimal amount of piggyBac vector for cotransfection with transposase plasmids are shown, and Figure 18 is a graph showing the tagging rate at the on-target site based on the amount (ng) of piggyBac plasmid used for HEK293T transfection. [Figure 19] The experimental results for determining the optimal amount of piggyBac vector for cotransfection with transposase plasmids are shown, and Figure 19 is a graph showing the tagmentation rate at off-target site 1 based on the amount (ng) of piggyBac plasmid used for HEK293T transfection. [Figure 20] The results of the tagging rate analysis by probe sequence length are shown, and here the tagging rate at the on-target site was analyzed. [Figure 21] The results of the tagging rate analysis by probe sequence length are shown, and here the tagging rate in off-target sites was analyzed. [Figure 22] This paper presents the results of an analysis of the prime editing rate and tagging rate at on-target sites for nine different pegRNAs. [Figure 23] The results of the tagging rate analysis at six target sites of HEK4 (+2G→T)pegRNA and HBB (+4A→T)pegRNA are shown. [Figure 24] The editing rates for Case 1 and Case 2, determined by targeted deep sequencing using a PE analyzer, are shown, where nine different pegRNAs were analyzed. [Figure 25] The results of tagging tests with and without prime editing are shown for 10 different on-target and off-target sites. [Figure 26] The following shows the comparative results for validated regions containing off-target sites of HEK4 pegRNA predicted by TAPE-seq. Figure 26 shows the comparative results for validated regions containing off-target sites of HEK4(+2G→T)pegRNA predicted by TAPE-seq. [Figure 27] The following shows the comparative results for validated regions containing off-target sites of HEK4 pegRNA predicted by TAPE-seq. Figure 27 shows the comparative results for combinations of validated sites of HEK4 (+3TAA ins), off-target sites of HEK4 (+2G→T) predicted by TAPE-seq using Mi-seq, and off-target sites of HEK4 (+2G→T) predicted by TAPE-seq using Hi-seq, as well as off-target sites of HEK4 (+3TAA ins) predicted by TAPE-seq (Mi-seq). [Figure 28]The following shows the comparative results for validated regions containing off-target sites of HEK4 pegRNA predicted by TAPE-seq. Figure 28 shows the comparative results for the validated sites of HEK4 (+2G→T), the off-target sites of HEK4 (+2G→T) predicted by TAPE-seq using Mi-seq, the off-target sites of HEK4 (+2G→T) predicted by TAPE-seq using Hi-seq, and the off-target sites of HEK4 (+3TAA ins) predicted by TAPE-seq (Mi-seq). [Figure 29] Regarding a comparison between the results predicted by TAPE-seq and those predicted by other off-target prediction methods, Figure 29 shows the results for HEK4(+2G→T)pegRNA. [Figure 30] Regarding the comparison between the results predicted by TAPE-seq and the results predicted by other off-target prediction methods, Figure 30 shows the results for HEK4(+3TAA ins)pegRNA. [Figure 31] Regarding a comparison between the results predicted by TAPE-seq and those predicted by other off-target prediction methods, Figure 31 shows the results for EMX1(+5G→T)pegRNA. [Figure 32] Regarding the comparison between the results predicted by TAPE-seq and the results predicted by other off-target prediction methods, Figure 32 shows the results for FANCF(+6G→C)pegRNA. [Figure 33] Regarding the comparison between the results predicted by TAPE-seq and the results predicted by other off-target prediction methods, Figure 33 shows the results for HEK3(+1CTT ins)pegRNA. [Figure 34] Regarding the comparison between the results predicted by TAPE-seq and the results predicted by other off-target prediction methods, Figure 34 shows the results for RNF2(+6G→A)pegRNA. [Figure 35]Regarding a comparison between the results predicted by TAPE-seq and those predicted by other off-target prediction methods, Figure 35 shows the results for DNMT1(+6G→A)pegRNA. [Figure 36] Regarding the comparison between the results predicted by TAPE-seq and the results predicted by other off-target prediction methods, Figure 36 shows the results for HBB(+4A→T)pegRNA. [Figure 37] Regarding a comparison between the results predicted by TAPE-seq and those predicted by other off-target prediction methods, Figure 37 shows the results for RUNX1(+6G→C)pegRNA. [Figure 38] Regarding the comparison between the results predicted by TAPE-seq and the results predicted by other off-target prediction methods, Figure 38 shows the results for VEGFA(+5G→T)pegRNA. [Figure 39] Figures 29-38 show the results of the analysis of validated off-target data missed by each prediction method, related to the results shown in Figures 29-38. [Figure 40] The results of the tagging rate analysis for PE2 TAPE-seq and PE4 TAPE-seq are shown. [Figure 41] Figure 41 shows a comparison of off-target results predicted by PE2 TAPE-seq, off-target results predicted by PE4 TAPE-seq, and true off-target results validated by targeted deep sequencing, with Figure 41 showing results related to HEK293T. [Figure 42] Figure 42 shows a comparison of off-target results predicted by PE2 TAPE-seq, off-target results predicted by PE4 TAPE-seq, and true off-target results validated by targeted deep sequencing. The results related to HeLa are shown in Figure 42. [Figure 43]Figure 43 shows a comparison of off-target results predicted by PE2 TAPE-seq, off-target results predicted by PE4 TAPE-seq, and true off-target results validated by targeted deep sequencing. The results related to K562 are shown. [Figure 44] Figures 41-43 show the analysis results summarizing the number of overlooked target sites. Here, Figure 44A shows the analysis results for each prediction method, and Figure 44B shows the analysis results for each cell. [Figure 45] The TAPE-seq off-target prediction results and cell-specific validation results are shown, with Figure 45 comparing the validation results for HEK293T with the TAPE-seq prediction results for each cell type. [Figure 46] The TAPE-seq off-target prediction results and cell-specific validation results are compared, with Figure 46 showing a comparison of validation results in HeLa cells and TAPE-seq prediction results in each cell type. [Figure 47] The TAPE-seq off-target prediction results and cell-specific validation results are compared, with Figure 47 showing a comparison of validation results in K562 cells and TAPE-seq prediction results in each cell type. [Figure 48] This shows the analysis results of the number of validated off-target cells missed based on TAPE-seq prediction results in each cell. [Figure 49] The results of the tagging rate analysis of TAPE-seq using PE2, PE2-nuclease, and PEmax-nuclease used with epegRNA are shown. [Figure 50]Figure 50 shows the results of comparing the off-target regions predicted by each TAPE-seq method (PE2 TAPE-seq, PE2-nuclease TAPE-seq, and TAPE-seq using PEmax-nuclease together with epegRNA) with the validated off-target regions. [Figure 51] The results of comparing the off-target regions predicted by each TAPE-seq method (PE2 TAPE-seq, PE2-nuclease TAPE-seq, and TAPE-seq using PEmax-nuclease along with epegRNA) with validated off-target regions are shown. Figure 51 shows the results for HBB (+4A→T)pegRNA and DNMT1 (+6G→C)pegRNA. [Figure 52] Figure 52 shows the results of comparing the off-target regions predicted by each TAPE-seq method (PE2 TAPE-seq, PE2-nuclease TAPE-seq, and TAPE-seq using PEmax-nuclease along with epegRNA) with the validated off-target regions. [Figure 53] Figure 53 shows the results of comparing the off-target regions predicted by each TAPE-seq method (PE2 TAPE-seq, PE2-nuclease TAPE-seq, and TAPE-seq using PEmax-nuclease along with epegRNA) with the validated off-target regions. [Figure 54] Figure 54 shows the results of comparing the off-target regions predicted by each TAPE-seq method (PE2 TAPE-seq, PE2-nuclease TAPE-seq, and TAPE-seq using PEmax-nuclease along with epegRNA) with the validated off-target regions. [Figure 55] Figure 55 shows the off-target prediction results for nDigenome-seq, GUIDE-seq, and TAPE-seq (TAPE-seq using PEmax-nuclease along with epegRNA) compared with validated off-target results. Figure 55 shows the results for HEK4(+2G→T)pegRNA and HEK4(+3TAA ins)pegRNA. [Figure 56] Figure 56 shows the off-target prediction results for nDigenome-seq, GUIDE-seq, and TAPE-seq (TAPE-seq using PEmax-nuclease along with epegRNA) compared with validated off-target results. Figure 56 shows the results for HBB (+4A→T)pegRNA and DNMT1 (+6G→C)pegRNA. [Figure 57] Figure 57 shows the results of comparing off-target predictions and validated off-target results for nDigenome-seq, GUIDE-seq, and TAPE-seq (TAPE-seq using PEmax nuclease along with epegRNA). [Figure 58] The results of off-target predictions for nDigenome-seq, GUIDE-seq, and TAPE-seq (TAPE-seq using PEmax-nuclease along with epegRNA) are shown, along with a comparison of validated off-target results. Figure 58 shows the results for FANCF (+6G→C)pegRNA and HEK3 (+1CTT ins)pegRNA. [Figure 59] The results of off-target predictions for nDigenome-seq, GUIDE-seq, and TAPE-seq (TAPE-seq using PEmax-nuclease along with epegRNA) are shown in comparison with validated off-target results. Figure 59 shows the results for RNF2 (+6G→A)pegRNA and RUNX1 (+6G→C)pegRNA. [Figure 60]The results of the analysis of the missed detection rates for GUIDE-seq, nDigenome-seq, TAPE-seq(PE2), TAPE-seq(PE2-nuclease), and TAPE-seq (using PEmax-nuclease together with epegRNA) are shown. [Figure 61] The results of a comparison of GUIDE-seq, nDigenome-seq, TAPE-seq(PE2), TAPE-seq(PE2-nuclease), and TAPE-seq (using PEmax-nuclease along with epegRNA) using receiver operating characteristic (ROC) curves are shown below. Figure 61 shows the results for HEK4(+2G→T)pegRNA and HEK4(+3TAA ins)pegRNA. [Figure 62] The results of a comparison of GUIDE-seq, nDigenome-seq, TAPE-seq(PE2), TAPE-seq(PE2-nuclease), and TAPE-seq (using PEmax-nuclease along with epegRNA) using receiver operating characteristic (ROC) curves are shown below. Figure 62 shows the results for HBB(+4A→T)pegRNA and DNMT1(+6G→C)pegRNA. [Figure 63] The results of a comparison of GUIDE-seq, nDigenome-seq, TAPE-seq(PE2), TAPE-seq(PE2-nuclease), and TAPE-seq (using PEmax-nuclease along with epegRNA) using receiver operating characteristic (ROC) curves are shown below. Figure 63 shows the results for HEK3(+1CTT ins)pegRNA. [Figure 64]The results of a comparison of GUIDE-seq, nDigenome-seq, TAPE-seq(PE2), TAPE-seq(PE2-nuclease), and TAPE-seq (using PEmax-nuclease along with epegRNA) using receiver operating characteristic (ROC) curves are shown below. Figure 64 shows the results for EMX1(+5G→T)pegRNA and FANCF(+6G→C)pegRNA. [Figure 65] The results of a comparison of GUIDE-seq, nDigenome-seq, TAPE-seq(PE2), TAPE-seq(PE2-nuclease), and TAPE-seq (using PEmax-nuclease along with epegRNA) using receiver operating characteristic (ROC) curves are shown below. Figure 65 shows the results for RNF2(+6G→A)pegRNA and RUNX1(+6G→C)pegRNA. [Figure 66] The results of a comparison of GUIDE-seq, nDigenome-seq, TAPE-seq(PE2), TAPE-seq(PE2-nuclease), and TAPE-seq (using PEmax-nuclease along with epegRNA) using receiver operating characteristic (ROC) curves are shown below. Figure 66 shows the results for VEGFA(+5G→T)pegRNA. [Figure 67] The analysis results for the area under the ROC curve, calculated based on the analysis results in Figures 61-66, are shown. [Figure 68] The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown, with Figure 68 showing the results related to editing patterns induced by HEK4(+3TAA ins)pegRNA. [Figure 69] The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown, with Figure 69 showing the results related to editing patterns induced by HEK4(+2G→T)pegRNA. [Figure 70] The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown, with Figure 70 showing the results related to editing patterns induced by HEK4(+2G→T)pegRNA. [Figure 71] The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown, with Figure 71 showing the results related to editing patterns induced by HEK4(+2G→T)pegRNA. [Figure 72] The results of the analysis of editing patterns at off-target sites, analyzed by targeted deep sequencing, are shown here. Figure 72 shows the results of the editing patterns at validated off-target sites associated with HEK4(+2G→T)pegRNA. [Figure 73] The results of the analysis of editing patterns at off-target sites, analyzed by targeted deep sequencing, are shown here. Figure 73 shows the results of the editing patterns at validated off-target sites associated with HEK4(+2G→T)pegRNA. [Figure 74] The results of the analysis of editing patterns at off-target sites, analyzed by targeted deep sequencing, are shown here. Figure 74 shows the results of the editing patterns at validated off-target sites associated with HEK4(+2G→T)pegRNA. [Figure 75] Figure 75 shows the results of the analysis of editing patterns at off-target sites analyzed by targeted deep sequencing, specifically at validated off-target sites associated with HEK4(+2G→T)pegRNA. [Figure 76] The results of the analysis of editing patterns at off-target sites, analyzed by targeted deep sequencing, are shown here. Figure 76 shows the results of the editing patterns at validated off-target sites associated with HBB(+4A→T)pegRNA. [Figure 77]The results of the analysis of editing patterns at off-target sites, analyzed by targeted deep sequencing, are shown here. Figure 77 shows the results of the editing patterns at validated off-target sites associated with HEK4(+3TAA ins)pegRNA. [Figure 78] The results of the analysis of editing patterns at off-target sites, analyzed by targeted deep sequencing, are shown here. Figure 78 shows the results of the editing patterns at validated off-target sites associated with HEK4(+3TAA ins)pegRNA. [Figure 79] The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown here. Figure 79 shows the results of the editing patterns at validated off-target sites associated with HEK4(+3TAA ins)pegRNA. [Figure 80] The results of the analysis of editing patterns at off-target sites, analyzed by targeted deep sequencing, are shown here. Figure 80 shows the results of the editing patterns at validated off-target sites associated with HEK4(+3TAA ins)pegRNA. [Figure 81] The results of the analysis of off-target site editing patterns analyzed by targeted deep sequencing are shown, with Figure 81 showing the results in HeLa cells, specifically the results for HEK4(+3TAA ins)pegRNA and HEK4(+2G→T)pegRNA. [Figure 82] The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown here. Figure 82 shows the results in HeLa cells, specifically the results for HEK4(+3TAA ins)pegRNA and HEK4(+2G→T)pegRNA. [Figure 83]The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown here. Figure 83 shows the results in K562 cells, specifically the results for HEK4(+3TAA ins)pegRNA and HEK4(+2G→T)pegRNA. [Figure 84] The results of the analysis of off-target site editing patterns, analyzed by targeted deep sequencing, are shown here. Figure 84 shows the results in K562 cells, specifically the results for HEK4(+3TAA ins)pegRNA and HEK4(+2G→T)pegRNA. [Figure 85] The results of the analysis of off-target site editing patterns analyzed by targeted deep sequencing are shown. Here, Figure 85 shows the results of the off-target site editing patterns validated by TAPE-seq performed using PEmax-nuclease, specifically the results for HEK4 (+2G→T)pegRNA, DNMT1 (+6G→C)pegRNA, HBB (+4A→T)pegRNA, and VEGFA (+5→T)pegRNA. [Figure 86] The results of the analysis of off-target site editing patterns analyzed by targeted deep sequencing are shown. Here, Figure 86 shows the results of the off-target site editing patterns validated by TAPE-seq performed using PEmax-nuclease, specifically the results for HEK4 (+2G→T)pegRNA, DNMT1 (+6G→C)pegRNA, HBB (+4A→T)pegRNA, and VEGFA (+5→T)pegRNA. [Figure 87] The results of the analysis of off-target site editing patterns analyzed by targeted deep sequencing are shown. Here, Figure 87 shows the results of the off-target site editing patterns validated by TAPE-seq using PEmax-nuclease, specifically the results for HEK4 (+2G→T)pegRNA, DNMT1 (+6G→C)pegRNA, HBB (+4A→T)pegRNA, and VEGFA (+5→T)pegRNA. [Figure 88] Figure 88 shows the results of ROC curve analysis constructed using the number of mismatches in each region of tpegRNA (target region, PBS, and RT template). The results for HEK4(+2G→T)pegRNA, HEK4(+3TAA ins)pegRNA, and HBB(+4A→T)pegRNA are shown. [Figure 89] Figure 89 shows the results of ROC curve analysis constructed using the number of mismatches in each region of tpegRNA (target region, PBS, and RT template). The results for HEK3 (+1CTT ins)pegRNA, FANCF (+6G→C)pegRNA, and EMX1 (+5G→T)pegRNA are shown. [Figure 90] Figure 90 shows the results of ROC curve analysis constructed using the number of mismatches in each region of tpegRNA (target region, PBS, and RT template). The results for DNMT1 (+6G→C)pegRNA, RUNX1 (+6G→C)pegRNA, and VEGFA (+5G→T)pegRNA are shown. [Figure 91] The analysis results for the area under the ROC curve, calculated based on the analysis results in Figures 88-90, are shown. [Figure 92] This shows the results of the analysis of the mismatch rate between false positive sites and validated sites predicted by TAPE-seq. [Figure 93] The vector map for the piggyBac PE2 all-in-one plasmid (pAllin1-PE2) is shown. [Modes for carrying out the invention]
[0057] Definition of Terms Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art in the field to which this disclosure relates. The following references provide common definitions of many terms used herein: [Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991)]. Unless otherwise specified, the following terms have the meanings assigned to them.
[0058] "Linking" or "Chaining" As used herein, the terms “linked” or “chained” mean that two or more elements present within a single conceptualizable structure are linked, either directly or indirectly (e.g., via other elements such as linkers), and are not intended to imply that no other additional elements can exist between the two or more elements. For example, a statement such as “element B is linked to element A” is intended to include, but should not be limited to, both cases where one or more other elements exist between elements A and B (i.e., element A is linked to element B via one or more other elements) and cases where one or more other elements do not exist between elements A and B (i.e., elements A and B are directly linked).
[0059] Sequence identity As used herein, the term “sequence identity” refers to the degree of similarity between two or more sequences. For example, the term “sequence identity” is used in conjunction with terms referring to a reference sequence and terms expressing ratios (such as percentages). For example, the term “sequence identity” may be used to describe a sequence that is similar to or substantially identical to a reference nucleotide sequence. When a sequence is described as having “at least 90% sequence identity with sequence A,” the reference sequence in this specification is sequence A. For example, the percentage of sequence identity can be calculated by aligning the reference sequence and the sequence for which the percentage of sequence identity is to be measured. Additionally, the percentage of sequence identity may be calculated by including all mismatches, deletions, and insertions for one or more nucleotides. The method for calculating and / or determining the percentage of sequence identity is not particularly limited and may be calculated and / or determined by any reasonable method or algorithm available to those skilled in the art.
[0060] Amino acid sequence notation Unless otherwise specified, amino acid sequences in this specification are described using either single-letter or three-letter abbreviations for amino acids, in the N-terminal to C-terminal direction. For example, RNVP means a peptide in which arginine, asparagine, valine, and proline are linked in that order from the N-terminal to the C-terminal direction. Another example is Thr-Leu-Lys, which means a peptide in which threonine, leucine, and lysine are linked in that order from the N-terminal to the C-terminal direction. Amino acids that cannot be represented by single-letter abbreviations may be described in more detail using other letters.
[0061] Each amino acid is denoted as follows: alanine (Ala, A); arginine (Arg, R); asparagine (Asn, N); aspartic acid (Asp, D); cysteine (Cys, C); glutamic acid (Glu, E); glutamine (Gln, Q); glycine (Gly, G); histidine (His, H); isoleucine (Ile, I); leucine (Leu, L); lysine (Lys, K); methionine (Met, M); phenylalanine (Phe, F); proline (Pro, P); serine (Ser, S); threonine (Thr, T); tryptophan (Trp, W); tyrosine (Tyr, Y); and valine (Val, V).
[0062] Nucleic acid sequence notation Where used herein, the symbols A, T, C, G, and U are interpreted as having meanings understood by those skilled in the art, which may be interpreted as bases, nucleosides, or nucleotides in DNA or RNA, depending on the context and the art. For example, when referring to a base, each symbol may be interpreted as adenine (A), thymine (T), cytosine (C), guanine (G), or uracil (U); when referring to a nucleoside, it may be interpreted as adenosine (A), thymidine (T), cytidine (C), guanosine (G), or uridine (U); when referring to a nucleotide in a sequence, it may be interpreted as meaning a nucleotide containing each of the above nucleosides.
[0063] Orientation of the disclosed sequence Unless otherwise specified or stated, nucleotide sequences (such as DNA sequences, RNA sequences, or DNA / RNA hybrid sequences) disclosed herein should be understood to begin in the 5'→3' direction. Amino acid sequences disclosed herein should be understood to begin in the N→C-terminal direction unless otherwise specified or stated.
[0064] target sequence As used herein, the term “target sequence” means a specific sequence recognized by a guide RNA or gene editing tool (e.g., Cas / conventional gRNA complex, prime editor enzyme / pegRNA complex, etc.) for cleaving a target gene or target nucleic acid. The target sequence may be appropriately selected depending on the purpose. For example, “target sequence” may refer to a sequence contained in the sequence of a target gene or target nucleic acid that is complementary to a spacer sequence contained in a guide RNA (e.g., pegRNA) (in this case, the target sequence may bind complementaryly to the spacer sequence of the guide RNA). In another example, “target sequence,” which is a sequence contained in the sequence of a target gene or target nucleic acid, may refer to a sequence complementary to a spacer sequence contained in a guide RNA (in this case, the target sequence may have substantially the same sequence as the spacer sequence of the guide RNA). As described above, the term “target sequence” may be used to refer to a sequence complementary to a spacer sequence contained in a guide RNA and / or a sequence substantially identical to the spacer sequence of the guide RNA, but should not be interpreted as being limited thereto. In some embodiments, the target sequence may be disclosed as a sequence containing a PAM sequence. In some embodiments, the target sequence may be disclosed as a sequence that does not contain a PAM sequence. The target sequence will be interpreted as appropriate depending on the context in which its contents are described. Typically, the spacer sequence is determined considering the sequence of the target gene or target nucleic acid and the PAM sequence recognized by the editing protein of the CRISPR / Cas system. The target sequence may refer only to the sequence of a particular strand that binds complementaryly to the guide RNA of the CRISPR / Cas complex, only to the sequence of a particular strand that does not bind complementaryly to the guide RNA, or to the entire target double helix containing a portion of a particular strand, which will be interpreted as appropriate depending on the context. The definition of the term "target sequence" disclosed herein is disclosed to describe the strand on which the target sequence may reside, and therefore the use of the term "target sequence" is not intended to distinguish between on-target and off-target sequences. The term "target sequence" may be used in reference to on-target sequences.Additionally, the term “target sequence” may be used in relation to off-target sequences. In other words, in some embodiments, an intended target sequence may be called an on-target sequence, and an unintended target sequence may be called an off-target sequence. For example, in some embodiments, an on-target sequence may be called a target sequence (in this case, the target sequence and the spacer sequence of the guide RNA may, for example, be substantially the same). In another example, in some embodiments, an off-target sequence may be called a target sequence (in this case, there may be, for example, no mismatch at all or one or more mismatches between the target sequence and the spacer sequence of the guide RNA). The term “target sequence” may be interpreted appropriately in relation to on-target and off-target sequences, depending on the context of the relevant paragraph.
[0065] Spacer linkage chain As used herein, the term "spacer-binding strand" refers to a strand in a gene editing system (e.g., CRISPR / Cas gene editing system, prime editing system, etc.) that uses a guide nucleic acid (e.g., guide RNA) to form a complementary bond with part or all of the spacer region sequence of the guide nucleic acid. Typically, DNA molecules such as genomes have a double-stranded structure. In a double-stranded structure, a strand that has a complementary sequence with part or all of the spacer region sequence of the guide nucleic acid, thereby forming a complementary bond, may be called a spacer-binding strand.
[0066] Spacer unbound chain As used herein, the term "spacer-unbound strand" refers to any strand other than the "spacer-bound strand" in a gene editing system (e.g., CRISPR / Cas gene editing system, prime editing system, etc.) that includes a guide nucleic acid (e.g., guide RNA), which contains a sequence that forms a complementary bond to some or all of the spacer region sequence of the guide nucleic acid. Typically, DNA molecules such as genomes have a double-stranded structure, and the term "spacer-unbound strand" can be used to refer to any strand other than the spacer-bound strand in the double helix. For example, in the editing of DNA molecules by a prime editing system, a strand containing a sequence that forms a complementary bond to some or all of the spacer region sequence of pegRNA may be called a "spacer-bound strand," while a strand containing a sequence that forms a complementary bond to the primer-binding site (PBS) of pegRNA may be called a "spacer-unbound strand." For example, in prime editing version 2, Cas9 (H840A) induces a nick on the spacer-unbound strand, forming a 3' DNA flap on the spacer-unbound strand.
[0067] The first and second strands of the DNA molecule Typically, DNA molecules, such as genomes, have a double helix structure composed of two strands. Such double-stranded DNA molecules are sometimes called double-stranded DNA. In descriptions of CRISPR / Cas-based gene editing systems, it may be necessary to distinguish between the two strands of a DNA molecule. One strand of a DNA molecule may be called the first strand. In this case, the strand other than the first strand of double-stranded DNA may be called the second strand. In each embodiment, the first and second strands can be set randomly. For example, in some embodiments, one strand of a DNA molecule may be called the first strand, and the other strand of the DNA molecule may be called the second strand. In some embodiments, for example, the spacer-bound strand may be called the first strand. In some embodiments, the spacer-unbound strand may, in another example, be called the first strand. As described above, one strand of a DNA molecule may be called the first strand, and the other strand may be called the second strand, as needed.
[0068] Upstream and downstream As used herein, the terms “upstream” and “downstream” are relative terms defining the linear positions of at least two elements located in a nucleic acid molecule (whether single-stranded or double-stranded) oriented in the 5’→3’ direction. For example, when describing a nucleic acid molecule in which a first element is upstream of a second element, the first element may be located somewhere in the 5’ direction relative to the second element. For example, if a single-nucleotide polymorphism (SNP) is located on the 5' side of a nick site, the SNP may be described as being upstream of the Cas9-induced nick site. In another example, when describing a nucleic acid molecule in which a first element is downstream of a second element, the first element may be located somewhere in the 3’ direction relative to the second element. For example, if an SNP is located on the 3’ side of a nick site, the SNP may be described as being downstream of the Cas9-induced nick site. Nucleic acid molecules may be DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA.
[0069] Nuclear localization signal or sequence (NLS) The term "NLS" refers to an amino acid sequence that facilitates the uptake of proteins into the nucleus of a cell. For example, protein uptake can be facilitated by nuclear transport. NLSs are well known and will be obvious to those skilled in the art. For example, exemplary sequences of NLSs are described in PCT application number PCT / EP2000 / 011690 (publication number WO2021 / 038547), the contents of which are incorporated herein by reference with respect to exemplary NLSs. In some embodiments, the NLS is PKKKRKV (SEQ ID NO: 1), KRPAATKKAGQAKKKK (SEQ ID NO: 2), PAAKRVKLD (SEQ ID NO: 3), RQRRNELKRSP (SEQ ID NO: 4), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 5), RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 6), VSRKRPRP (SEQ ID NO: 7), PPKKARED (SEQ ID NO: 8), P The amino acid sequences may include, but are not limited to, QPKKKPL (SEQ ID NO: 9), SALIKKKKKMAP (SEQ ID NO: 10), DRLRR (SEQ ID NO: 11), PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 17). One or more NLSs can be selectively fused to gene-editing proteins such as Cas proteins or prime editor proteins. The NLS fused to a protein can be used to facilitate the translocation of such a protein to which it is linked to the nucleus, which is the desired site.
[0070] Proteins, peptides, and polypeptides As used herein, the terms “protein,” “peptide,” and “polypeptide” are interchangeable and refer to polymers of amino acid residues linked by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides having any size, structure, or function. Typically, these proteins, peptides, or polypeptides may have a length of at least three or more amino acids. In some embodiments, a protein, peptide, or polypeptide may refer to an individual protein or a combination of proteins. For example, protein, peptide, or polypeptide may be used as a term encompassing all meanings of individual proteins, fusion proteins formed by the fusion of two or more elements (in which case at least one of the two elements is a protein), and protein complexes formed by the compounding of two or more elements (in which case at least one of the two elements is a protein). In some embodiments, one or more amino acids in a protein, peptide, or polypeptide may be modified. In this case, modifications included in the protein, peptide, or polypeptide may be modifications caused by the addition of chemicals such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, and fatty acid groups, conjugation, functionalization, or linkers for other modifications. In some embodiments, the protein, peptide, or polypeptide may be a single molecule or a multimolecular complex. In some embodiments, the protein, peptide, or polypeptide may be a naturally occurring protein. In some embodiments, the protein, peptide, or polypeptide may be a protein fragment. In some embodiments, the protein, peptide, or polypeptide may be naturally occurring, recombinant, synthetic, or prepared by any combination thereof. Any protein provided herein may be prepared by any method known in the art.For example, any protein provided herein may be prepared by recombinant protein expression and purification, which is particularly suitable for fusion proteins containing peptide linkers. Inventions for recombinant protein expression and purification are widely known. See [Green, Michael R., and Joseph Sambrook. "Molecular cloning." A Laboratory Manual 4th (2012).] (Its entire contents are incorporated herein by reference).
[0071] functional equivalent The terms “functional equivalent” or “equivalent” refer to a secondary molecule or conceptualizable element that is functionally equivalent but is not necessarily structurally equivalent to a primary molecule or conceptualizable element. For example, “Cas9 equivalent” refers to a protein that has the same, substantially the same, or similar function as Cas9 but does not necessarily have the same amino acid sequence. Throughout this application, where a particular protein is referred to, it is intended that the particular protein referred to includes all of its functional equivalents. For example, where it is written “protein X,” the term protein X may be interpreted to include its functional equivalents. In this embodiment, “functional equivalent” or “equivalent” of protein X includes any paralogs, orthologues, fragments, homologs, naturally occurring, engineered, mutated, and synthetic versions of protein X having equivalent function. For example, where the term Cas protein is used, this term may be interpreted to include equivalents of the Cas protein (e.g., Cas nickase). In another example, where the term reverse transcriptase is used, this term may be interpreted to include equivalents of reverse transcriptase.
[0072] cyclic permutant As used herein, the term “circular permutation” refers to a polypeptide or protein that contains a circular permutation, which is a structural rearrangement of the protein in which the order of amino acids found in the protein’s amino acid sequence is altered. A circular permutation is a protein with modified N-terminus and / or C-terminus compared to its wild-type counterpart. For example, the wild-type C-terminal half of a protein becomes a new N-terminal half. For example, circular permutation (or circular permutation: CP) refers to a topological rearrangement of the primary sequence of a protein in which the sequence is split at different sites, and adjacent N-terminus and C-terminus are newly prepared, while these N-terminus and C-terminus are linked by a peptide linker. As a result, proteins with the same or similar three-dimensional (3D) shape overall, but with different linkages, can be prepared. For example, protein structures with improved or modified characteristics can be prepared, including reduced proteolytic sensitivity, improved catalytic activity, modified matrix or ligand binding, and / or improved thermal stability. Circular permutation proteins can occur naturally (e.g., concanavalin A and lectins). Additionally, cyclic substitutions can occur as a result of post-translational modifications or be manipulated using recombination techniques. A cyclic substitution of a particular protein may be included in the equivalent of that particular protein.
[0073] As an example of a circularly permuted Cas9, "circularly permuted Cas9" refers to any Cas9 protein or variant thereof resulting from a circularly permuted Cas9 having locally rearranged N-terminus and C-terminus. Such circularly permuted Cas9 proteins ("circularly permuted Cas9: CP-Cas9") or variants thereof have the ability to bind to DNA when complexed with guide RNA. See [Oakes, Benjamin L., Dana C. Nadler, and David F. Savage. "Protein engineering of Cas9 for enhanced function." Methods in enzymology. Vol. 546. Academic Press, 2014. 491-511.; and Oakes, Benjamin L., et al. "CRISPR-Cas9 circular permutants as programmable scaffolds for genome modification." Cell 176. 1-2 (2019): 254-267.] (each of which is incorporated herein by reference). The disclosures herein include novel CP-Cas9s, insofar as the resulting circulating replacement protein has the ability to bind to DNA when complexed with gRNA, or in consideration of any previously known CP-Cas9. Exemplary sequences of CP-Cas9 proteins are disclosed in WO2020191233A1 (Application No. PCT / US2020 / 023712), which is incorporated herein by reference in its entirety.
[0074] Fusion protein As used herein, the term “fusion protein” refers to a hybrid polypeptide comprising a domain or protein derived from at least two different types of elements (in this case, at least one of the elements being a protein). For example, a fusion protein may be a hybrid polypeptide comprising proteins derived from two different types of proteins. One type of protein may be located at the amino-terminus (N-terminus) or carboxy-terminus (C-terminus) of the fusion protein, thus forming an “amino-terminus fusion protein” or a “carboxy-terminus fusion protein,” respectively. In some embodiments, the term fusion protein may be used to refer to an element having a single molecular form in which two or more elements are covalently linked. In other embodiments, the term fusion protein may be used to refer to an element having a multimolecular complex form in which two or more elements are noncovalently linked.
[0075] Linker As used herein, the term “linker” may refer to a molecule that links two other molecules or parts. In a fusion protein, the linker linking two proteins may be an amino acid sequence. For example, Cas9 can be linked to reverse transcriptase via an amino acid linker sequence to form a fusion protein. Additionally, the linker linking two nucleotide sequences may be a nucleotide sequence. For example, in conventional guide RNA, crRNA is linked to tracrRNA via a linker, thus forming a single-stranded guide RNA. In other embodiments, the linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker may have a length of 1 to 200 amino acids, but is not limited thereto. In some embodiments, the linker may have a length of 1 to 500 nucleotides, but is not limited thereto. Furthermore, longer linkers may also be conceivable.
[0076] Dual-specific ligands As used herein, the terms “bispecific ligand” or “bispecific moiety” refer to a ligand that binds to two different types of ligand-binding domains. In certain embodiments, the ligand is a small molecule compound, peptide, or polypeptide. In other embodiments, the ligand-binding domain is a dimerizing domain that can be placed in a protein as a peptide tag. In various embodiments, two proteins, each independently containing the same or different dimerizing domains, can be induced to dimerize through the binding of each dimerizing domain to a bispecific ligand. As used herein, “bispecific ligand” may also mean “chemical inducer of dimerization (CID).”
[0077] Dimerization domain The term "dimerization domain" refers to a ligand-binding domain that binds to the binding site of a bispecific ligand. A first dimerization domain binds to a first binding site of a bispecific ligand, and a second dimerization domain binds to a second binding site of the same bispecific ligand. When the first dimerization domain fuses to a first protein and the second dimerization domain fuses to a second protein, the first and second proteins can dimerize in the presence of the bispecific ligand. In this case, the bispecific ligand has at least one portion that binds to the first dimerization domain and at least another portion that binds to the second dimerization domain. In some embodiments, a dimerization domain (such as the first dimerization domain) can be linked to a Cas protein. In some embodiments, a dimerization domain (such as the second dimerization domain) can be linked to a reverse transcriptase.
[0078] Nikkaze The term "nickase" refers to a Cas protein in which one of its two nuclease domains has been inactivated. Nickase can cleave only one strand of target DNA molecules.
[0079] Flap end nuclease As used herein, the term “flap endonucleases” refers to enzymes that catalyze the removal of 5' single-stranded DNA flaps. Such enzymes handle the removal of 5' flaps formed during cellular processes, including DNA replication. In some embodiments, a prime editing method may use endogenously or exogenously provided flap endonucleases to remove 5' flaps of endogenous DNA formed in a target region during prime editing. Flap endonucleases are known in the art and are disclosed in detail in [Patel, Nikesh, et al. "Flap endonucleases pass 5'-flaps through a flexible arch using a disorder-thread-order mechanism to confer specificity for free 5'-ends." Nucleic acids research 40.10(2012):4507-4519.; and Tsutakawa, Susan E., et al. "Human flap endonuclease structures, DNA double-base flipping, and a unified understanding of the FEN1 superfamily." Cell 145.2(2011):198-211.], each of which is incorporated herein by reference. An exemplary flap endonuclease may be flap structure-specific endonuclease 1 (FEN1). The sequence of FEN1 is disclosed in WO2020191233A1 (Application No. PCT / US2020 / 023712).
[0080] Effective amount As used herein, the term “effective amount” refers to the amount of a biologically active agent sufficient to produce a desired biological response. For example, in some embodiments, the effective amount of a prime editor protein may refer to the amount of protein sufficient to edit a nucleotide sequence in a target region, such as a genome. In some embodiments, the effective amount of a prime editor protein provided herein, such as a fusion protein comprising a nickasase Cas9 domain and reverse transcriptase, may refer to the amount of fusion protein sufficient to induce editing of a target region intended to be specifically bound and edited by the fusion protein. As will be understood by those skilled in the art, the effective amount of an agent, such as a fusion protein, nuclease, hybrid protein, protein dimer, complex of a protein (or protein dimer) and a polynucleotide, or polynucleotide, can vary depending on various factors, such as the desired biological response, the specific gene to be edited, the genome to be edited, the target region to be edited, the cell or tissue being targeted, and the agent used.
[0081] about As used herein, the term "about" means an approximation to any quantity, such as 30%, 25%, 20%, of a given quantity, level, value, number, frequency, percentage, dimension, size, volume, weight, or length. 15 It refers to a quantity, level, value, number, frequency, percentage, dimension, size, quantity, weight, or length that varies by %, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.
[0082] CRISPR / Cas system Overview of the CRISPR / Cas System CRISPR This "CRISPR" section is provided for the convenience of engineers, and the terms used in this section are not intended to limit the terms disclosed herein.
[0083] CRISPR is a family of DNA sequences (i.e., the CRISPR cluster) found in bacteria and archaea that represent snippets of previous infections by viruses that have invaded prokaryotes. These DNA snippets are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with CRISPR-associated RNA and CRISPR-associated protein (Cas protein) sequences, they constitute the prokaryotic immune defense system. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). Cas9 / crRNA / tracrRNA then endonuclease-cleaves linear or circular dsDNA targets complementary to the RNA. Specifically, target strands not complementary to the crRNA are first endonuclease-cleaved and then exonuclease-trimmed 3'-5'. DNA binding and cleavage typically require both proteins and both RNAs. However, single guide RNAs (sgRNAs, or simply gRNAs) have been developed, and such single-stranded RNAs are engineered so that both crRNA and tracrRNA aspects are combined into a single RNA species. See, for example, [Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Science 337.6096(2012):816-821.] (its entire contents are incorporated herein by reference). Cas9 recognizes short motifs (protospacer adjacent motifs: PAMs) in CRISPR repeat sequences and helps distinguish between self and non-self.Not only is CRISPR biology well-known, but the sequence and structure of Cas9 nuclease are also well-known to those skilled in the art (see, for example, [Ferretti, Joseph J., et al. "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Proceedings of the National Academy of Sciences 98.8(2001):4658-4663.; Deltcheva, Elitza, et al. "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Nature 471.7340(2011):602-607.; and Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." science 337.6096(2012):816-821.] (the full contents of each of these are incorporated herein by reference)]. Cas9 orthologs have been described in a variety of species, including but not limited to Streptococcus pyogenes (S. pyogenes) and Streptococcus thermophilus. Additional suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on the disclosure herein, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in [Chylinski, Krzysztof, Anais Le Rhun, and Emmanuelle Charpentier. "The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems." RNA biology 10.5(2013):726-737.], the entire contents of which are incorporated herein by reference.
[0084] CRISPR / Cas system and DNA molecule editing using the system The CRISPR / Cas system, developed from the aforementioned CRISPR, is a technology that uses Cas proteins derived from the cell's CRISPR system and guide nucleic acids to direct the Cas proteins to a target region, allowing for the editing of desired DNA molecules (such as the cell's genome) at a desired site. For example, the Cas protein forms a Cas / gRNA complex with guide RNA (gRNA). The Cas / gRNA complex is directed to the desired site by the guide RNA it contains. The Cas protein in the Cas / gRNA complex induces a double-strand break (DSB) or a nick (in the case of nickase) at the desired site. Using the CRISPR / Cas system, it may be possible to edit not only the cell's genome but also DNA molecules not located in the genome. Since the discovery of CRISPR, single-stranded guide RNAs (sgRNAs) linked to tracrRNA and crRNA have been developed for the CRISPR / Cas system as described above (see [Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." science 337.6096(2012):816-821.] (the full contents of which are incorporated herein by reference). In addition, various classes and / or types of Cas proteins have been developed, such as Cas9, Cas12a (cpf1), Cas12b (c2c1), Cas12e (CasX), Cas12k (c2c5), Cas14, Cas14a, Cas13a (c2c2), Cas13b (c2c6), Cas nickase (Cas9 nickase, etc.), and dead Cas. In some embodiments, the Cas protein is sometimes referred to as the CRISPR enzyme. For an understanding of the CRISPR / Cas system, see WO2018 / 231018 (International Publication Number) (its entire contents are incorporated herein by reference). For the convenience of the technician, the Cas proteins (or CRISPR enzymes) that may be used in the CRISPR / Cas system are further described below.
[0085] Cas protein Overview of Cas protein In relation to the CRISPR / Cas system, the term Cas protein can be used to refer to a protein that creates a nick or DSB in a desired region to complete editing, or a protein that helps induce editing. The term Cas protein can be used to include its equivalents. Typically, Cas proteins have nuclease activity that cleaves nucleic acids. For example, some Cas proteins can induce double-strand breaks (DSBs), which are sometimes called Cas nucleases. In another example, some Cas proteins can induce nicks, which are sometimes called Cas nickases. Some Cas proteins are modified to lack nuclease activity, and these are sometimes called dead Cas. In the CRISPR / Cas system, Cas proteins can be used interchangeably with CRISPR enzymes. A representative example of a Cas protein is Cas9.
[0086] As used herein, the term Cas protein is used to refer collectively to edited proteins or inactive Cas proteins that can produce double-strand breaks (DSBs) or nicks in a target region, as used in the CRISPR / Cas system. Examples of Cas proteins include, but are not limited to, Cas9, Cas9 variants, Cas9 nickase (nCas9), dead Cas9, Cpf1 (Cas12a) (Type V CRISPR-Cas system), C2c1 (Cas12b) (Type V CRISPR-Cas system), C2c2 (Cas13a) (Type VI CRISPR-Cas system), and C2c3 (Type V CRISPR-Cas system). Examples of additional Cas proteins are described in [Abudayyeh, Omar O., et al. "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353.6299(2016):aaf5573.], the full contents of which are incorporated herein by reference.
[0087] In one embodiment, the Cas protein is derived from Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Campylobacter jejuni, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptosporangium roseum, and Streptosporangium roseum. roseum), Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp. (Polaromonas Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp.(Synechococcus sp.), Acetohalobium arabaticum, Ammonifex degensii, Caldicellulosiruptor bescii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus Caldus), Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsonii, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp. sp.), Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp.The Cas proteins (e.g., Cas9 or Cpf1) may be derived from various microorganisms such as *Microcoleus chthonoplastes*, *Oscillatoria sp.*, *Petrotoga mobilis*, *Thermosipho africanus*, or *Acaryochloris marina*.
[0088] Below, we will illustrate this using the Cas9 protein, a representative example of Cas proteins.
[0089] Cas9 protein In the CRISPR / Cas9 system, a protein that has nuclease activity to cleave nucleic acids, or a protein with inactivated nuclease activity, is called a Cas9 protein. The term Cas9 protein is used to include its equivalents. Additionally, Cas9 proteins are sometimes called Cas9 nucleases, casn1 nucleases, or clustered regularly interspaced short palindromic repeat (CRISPR)-associated nucleases. Cas9 proteins correspond to Class 2 Type II of the CRISPR / Cas system classification, and examples include proteins derived from Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptomyces roseum, or Streptomyces roseum.The sequence and structure of the Cas9 protein are well known to those skilled in the art (see, for example, [Ferretti, Joseph J., et al. "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Proceedings of the National Academy of Sciences 98.8(2001):4658-4663.; Deltcheva, Elitza, et al. "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Nature 471.7340(2011):602-607.; and Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." science 337.6096(2012):816-821.] (the full contents of each of these are incorporated herein by reference)]). Additional Cas9 proteins and sequences are disclosed in [Chylinski, Krzysztof, Anais Le Rhun, and Emmanuelle Charpentier. "The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems." RNA biology 10.5(2013):726-737.], which is incorporated herein by reference in its entirety. For example, the DNA cleavage domain of Cas9 is known to contain two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the strand non-complementary. Inactivating either of these subdomains can silence the nuclease activity of the inactivated subdomain, and inactivating both of these subdomains can silence the overall nuclease activity of Cas9. For example, the H840A mutation provides Cas9 nickasase.For example, both the D10A and H840A mutations completely inactivate the nuclease activity of S. pyogenes Cas9 (see [Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." science 337.6096(2012):816-821.]). In some embodiments, proteins containing Cas9 fragments may be provided. For example, a protein may contain one or more of the following two Cas9 domains: the gRNA-binding domain of Cas9 and the DNA-cleaving domain of Cas9. In some embodiments, Cas9 variants may be provided. Cas9 variants are homologous to Cas9 or its fragments. For example, a Cas9 variant may be at least approximately 70% identical, at least approximately 80% identical, at least approximately 90% identical, at least approximately 95% identical, at least approximately 96% identical, at least approximately 97% identical, at least approximately 98% identical, at least approximately 99% identical, at least approximately 99.5% identical, at least approximately 99.6% identical, at least approximately 99.7% identical, at least approximately 99.8% identical, or at least approximately 99.9% identical to a wild-type Cas9 (such as SpCas9). In some embodiments, a Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more amino acid changes compared to wild-type Cas9 (such as SpCas9). In some embodiments, a Cas9 variant may include a Cas9 fragment (such as a gRNA-binding domain and / or a DNA-cleaving domain).In some embodiments, the Cas9 variant fragment may be at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, or at least about 99.9% identical to the corresponding wild-type Cas9 fragment. In some embodiments, the wild-type Cas9 fragment or Cas9 variant fragment may be at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% or higher of the amino acid length of the corresponding wild-type Cas9.
[0090] Guide RNA Overview of Guide RNA In the CRISPR / Cas system, the Cas protein associates with a guide nucleic acid to form a Cas / guide nucleic acid complex. Typically, in the CRISPR / Cas system, a guide RNA (gRNA) is used as the guide nucleic acid, and the Cas protein associates with the guide RNA to form a Cas / gRNA complex. The Cas / gRNA complex is sometimes called a ribonucleoprotein (RNP). The Cas / gRNA complex creates a nick or double-strand break (DSB) in a target region containing a sequence that corresponds to (e.g., complementary to) the spacer sequence of the guide RNA (gRNA), and the Cas protein induces the DSB or nick. The site where the DSB or nick is created may be near a PAM sequence in the genome.
[0091] Cas / gRNA targeting involves protospacer-adjacent motifs (PAMs) in the genome and spacer sequences in guide RNAs. Cas proteins (such as Cas9) directed to the target region by the PAMs and guide RNA spacer sequences then construct double-sided blocks (DSBs) in the target region.
[0092] In the CRISPR / Cas gene editing system, the RNA that directs the Cas protein to the target region and recognizes a specific sequence contained in the target DNA molecule is called guide RNA.
[0093] The functional structure of guide RNA can be broadly divided into: 1) the scaffold sequence portion and 2) the guide domain containing the guide sequence. The scaffold sequence portion, with which Cas proteins (such as the Cas9 protein) interact, is the region where the resulting complex is formed upon binding to the Cas protein. Typically, the scaffold sequence portion includes tracrRNA and crRNA repeat sequences, and the scaffold sequence is determined by the type of Cas protein used. The guide sequence is a region that can complementarily bind to a specific length of nucleotide sequence in the target nucleic acid (such as the cellular genome or target DNA molecule). The guide sequence may be artificially modified and is determined by the target nucleotide sequence of interest related to the desired gene editing.
[0094] In some embodiments, the guide RNA may be described as comprising crRNA and tracrRNA. The crRNA may include spacer and repeat sequences. A portion of the repeat sequence of the crRNA may interact with (e.g., complementarily bind to) the tracrRNA portion. As described above, a single-stranded guide RNA (sgRNA) in which tracrRNA and crRNA are ligated may be provided (see Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Science 337.6096(2012):816-821.] (its entire contents are incorporated herein by reference). In other words, the guide RNA may be provided as double-stranded or single-stranded.
[0095] In some embodiments, the sgRNA may be described as comprising a guide domain, a first complementary domain, a linker domain, and a second complementary domain. In this case, the sgRNA may include, but is not limited to, additional domains comprising one or more of the proximal and tail domains. In this case, the linker domain ligates the first and second complementary domains, and some or all of the first complementary domain forms a complementary bond with some or all of the second complementary domain. As a result, the first complementary domain, the linker domain (e.g., including a polynucleotide linker), and the second complementary domain form a secondary structure such as a loop structure (see [PCT application number PCT / KR2018 / 006803, publication number WO2018 / 231018]).
[0096] The term guide RNA also includes equivalent guide nucleic acid molecules, whether naturally occurring or not (e.g., engineered, recombinant, etc.), that associate with Cas9 equivalents, homologs, orthologues, or paralogs, enabling the localization of the Cas9 equivalent to a specific target nucleotide sequence. As described above, Cas9 equivalents may include other Cas proteins derived from any type of CRISPR system (such as types II, V, and VI), including Cpf1 (type V CRISPR-Cas system), C2c1 (type V CRISPR-Cas system), C2c2 (type VI CRISPR-Cas system), and C2c3 (type V CRISPR-Cas system). Additional Cas equivalents are described in [Abudayyeh, Omar O., et al. "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353.6299(2016):aaf5573.], the entire contents of which are incorporated herein by reference. Guide RNA used in conventional CRISPR / Cas systems may be referred to as “conventional” guide RNA to distinguish it from a modified guide RNA called prime editing guide RNA (pegRNA) invented for the prime editing methods and compositions described herein. Prime editing guide RNA (pegRNA) may have a form in which an elongation arm is attached to the 3' or 5' end of conventional guide RNA.
[0097] Guide RNA or pegRNA may include one or more of the following: spacers, gRNA cores, elongation arms (particularly in pegRNA), and transcription terminators. Furthermore, various structural elements may be additionally included, but are not limited to these. A spacer includes a spacer sequence, which refers to a sequence in the guide RNA or pegRNA that binds to the region containing the protospacer sequence of the target region. The gRNA core, sometimes called the gRNA scaffold or backbone sequence, refers to a sequence in the gRNA or pegRNA responsible for binding to Cas9 or its equivalent. The gRNA core does not include the spacer or target sequence used to guide Cas9 to the target region (target DNA). Elongation arms (particularly in pegRNA) are elements of pegRNA that include a primer binding site (PBS) and a DNA synthesis template sequence for positioning a single-stranded DNA flap containing the desired gene modification by a polymerase (such as reverse transcriptase). Elongation arms may be located at the 3' or 5' end of the pegRNA and are designed to position the desired gene modification. The elongation arm of pegRNA is sometimes called the elongation region. In some embodiments, the guide RNA or pegRNA may further contain a transcription stop sequence at the 3' of the molecule.
[0098] Guide RNA guide sequence The guide RNA may include a guide domain containing a guide sequence. The guide sequence may be used interchangeably with the spacer sequence. The guide domain may be used interchangeably with the spacer. The guide sequence is a portion that may be artificially designed and is determined by the target nucleotide sequence of interest. In some embodiments, the guide sequence may be designed to target a sequence adjacent to the PAM sequence located on the desired DNA molecule to be edited. As described above, localization of the Cas / gRNA complex to the target site (such as an on-target site) may be achieved. The structure of the guide nucleic acid may vary depending on the type of CRISPR. For example, the guide RNA used in the CRISPR / Cas9 gene editing system may have a 5'-[guide domain]-[scaffold]-3' structure.
[0099] In one embodiment, the guide sequence may have a length of 5 nt to 40 nt. In one embodiment, the guide sequence included in the guide domain of the guide RNA may have a length of 10 nt to 30 nt. In one embodiment, the guide sequence may have a length of 15 nt to 25 nt. In one embodiment, the guide sequence may have a length of 18 nt to 22 nt. In one embodiment, the guide sequence may have a length of 20 nt. In one embodiment, the target sequence that acts as a sequence in the genome that forms a complementary bond with the guide sequence (including all target sequences present in the spacer-binding and spacer-unbinding strands) may have a length of 5 nt to 40 nt or 5 bp to 40 bp. In one embodiment, the target sequence that acts as a sequence in the genome that forms a complementary bond with the guide sequence may have a length of 10 nt to 30 nt or 10 bp to 30 bp. In one embodiment, the target sequence may have a length of 15 nt to 25 nt or 15 bp to 25 bp. In one embodiment, the target sequence may have a length of 18 nt to 22 nt or 18 bp to 22 bp. In one embodiment, the target sequence may have a length of 20 nt or 20 bp.
[0100] PAM For a conventional CRISPR / Cas system to cleave a target DNA molecule, two requirements must be met. First, a base sequence (nucleotide sequence) of a certain length that can be recognized by a Cas protein (such as the Cas9 protein) must be present in the target gene or target nucleic acid. In this case, the base sequence (nucleotide sequence) of a certain length that can be recognized by the Cas9 protein is called a protospacer adjacent motif (PAM) sequence. The PAM sequence is a unique sequence determined by the Cas9 protein. Second, a sequence that can bind complementaryly to the spacer sequence contained in the guide RNA must be present near the PAM sequence of a certain length. In this case, the PAM sequence may be used to include all sequences present in both the spacer-unbound and spacer-bound strands.
[0101] As described above, the Cas / gRNA complex of the CRISPR / Cas system is directed to the target region by the guide sequence of the gRNA and the PAM sequence of the target DNA molecule (e.g., the genome of a cell). In the target DNA molecule, the PAM sequence may be located on a strand of the guide RNA to which the guide sequence does not bind, rather than on the strand to which the guide sequence binds. The PAM sequence can be determined independently depending on the type of Cas protein used. In one embodiment, the PAM sequence may be any one selected from the following (disclosed in the 5'→3' direction): NGG (SEQ ID NO: 19); NNNNRYAC (SEQ ID NO: 20); NNAGAAW (SEQ ID NO: 21); NNNNGATT (SEQ ID NO: 22); NNGRR(T) (SEQ ID NO: 23); TTN (SEQ ID NO: 24); and NNNVRYAC (SEQ ID NO: 25). Each N may independently be A, T, C, or G. Each R may independently be A or G. Each Y may independently be C or T. Each W may independently be A or T. For example, when using SpCas9 as the Cas protein, the PAM sequence may be NGG (SEQ ID NO: 19). For example, when using Streptococcus thermophilus Cas9 (StCas9) as the Cas protein, the PAM sequence may be NNAGAAW (SEQ ID NO: 21). For example, when using Neisseria meningitides Cas9 (NmCas9), the PAM sequence may be NNNNGATT (SEQ ID NO: 22). For example, when using Campylobacter jejuni Cas9 (CjCas9), the PAM may be NNNVRYAC (SEQ ID NO: 25). In one embodiment, the PAM sequence may be ligated to the 3' end of a target sequence present on the spacer-unbound strand (in this case, the target sequence present on the spacer-unbound strand refers to a sequence that does not bind to the guide RNA). In one embodiment, the PAM sequence may be located at the 3' end of a target sequence present on the spacer-unbound strand. The target sequence present on the spacer-unbound strand refers to a sequence that does not bind to the guide sequence of the guide RNA.The target sequences present in the non-spacer-binding strand are complementary to the target sequences present in the spacer-binding strand.
[0102] The site where the DSB or nick is created may be near the PAM sequence of the genome. In one embodiment, the site where the DSB or nick is created may be in the range of -0 to -20 or +0 to +20 relative to the 5' or 3' end of the PAM sequence present in the spacer-unbound strand. In one embodiment, the site where the DSB or nick is created may be in the range of -1 to -5 or +1 to +5 relative to the PAM sequence in the spacer-unbound strand. For example, in a CRISPR / Cas system using SpCas9, it is well known that SpCas9 cleaves between the third and fourth nucleotides located upstream of the PAM sequence.
[0103] Conventional CRISPR / Cas system genome editing process For the convenience of engineers, a simplified genome editing process using a conventional CRISPR / Cas system is disclosed using the following example. In this case, a conventional CRISPR / Cas system refers to a system that can edit DNA molecules using Cas proteins and conventional gRNAs.
[0104] For example, an environment can be provided in which the desired DNA molecule to be edited can come into contact with the Cas / gRNA complex. For intracellular genome editing, the Cas protein or nucleic acid encoding the Cas protein, and the guide RNA or nucleic acid encoding the guide RNA are introduced into the cell, thereby creating an environment in which the Cas protein and guide RNA can come into contact with the cell's genomic DNA. In an environment where the Cas protein and guide RNA can come into contact with the cell's genomic DNA, the Cas protein and guide RNA can form a Cas / gRNA complex. Even in the absence of genomic DNA, the Cas / gRNA complex can be formed if both the Cas protein and gRNA are in a suitable environment. The Cas / gRNA complex is directed towards a target region containing a pre-designed target sequence, involving the PAM sequence of the genome and the guide sequence of the gRNA contained in the Cas / gRNA complex. The Cas / gRNA complex directed towards the target region creates a double-stroke backbone (DSB) in the target region (e.g., in the case of Cas9). The DNA from which the DSB has been created (cleaved) is then repaired through the DNA repair process, thereby completing gene editing of the target region or target site. The two main pathways for repairing double-stroke bonds (DSBs) in DNA are homology-directed repair (HDR) and nonhomologous end joining (NHEJ). HDR is a naturally occurring DNA repair system that lies between the two pathways and can be used to modify the genomes of various organisms, including humans. HDR-mediated repair can be used, but is not limited to, inserting a desired sequence into a target region or site, or inducing specific point mutations. HDR-mediated repair can be carried out using HDR, a DNA repair system, and an HDR template (such as a donor template that may be supplied from outside the cell). NHEJ refers to the DSB repair process of DNA, which, unlike HDR, joins the cleaved ends without an HDR template; that is, this repair process does not require an HDR template. NHEJ may also be a DNA repair mechanism that can be selected primarily to induce indels.An indel (insertion / deletion) can refer to a variation in the nucleotide sequence of a nucleic acid before gene editing, in which several nucleotides are deleted and arbitrary nucleotides are inserted, and / or a combination of such insertions and deletions. The occurrence of several indels in a target gene can lead to the inactivation of the corresponding gene. The DNA repair mechanisms HDR and NHEJ are disclosed in detail in [Sander, Jeffry D., and J. Keith Joung. "CRISPR-Cas systems for editing, regulating and targeting genomes." Nature biotechnology 32.4(2014):347-355.], the entire contents of which are incorporated herein by reference.
[0105] The conventional CRISPR / Cas system, which forms the basis of the prime editing system, is described in detail above for the convenience of the engineer. This application relates to a novel system for predicting off-target events that may occur in the DNA editing process using the prime editing system. Prior to describing the off-target prediction system in the prime editing system provided by this application, the prime editing system on which the off-target prediction system is based, and the process of editing DNA molecules using the prime editing system are described in detail below.
[0106] Prime editing system Overview of the Prime Editing System Prime editing, developed by David R. Liu et al., is a technique that uses specialized guide RNA, including Cas proteins, polymerases (such as reverse transcriptase), and DNA synthesis templates, to edit DNA molecules (such as genomes) and incorporate or insert desired edits into target regions of the DNA molecules. A description of prime editing and various embodiments are disclosed in [Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.; Chen, Peter J., et al. "Enhanced prime editing systems by manipulating cellular determinants of editing outcomes." Cell 184.22(2021):5635-5652.; and PCT application number PCT / US2020 / 023712, publication number WO2020191233A1], the entire contents of each of these are incorporated herein by reference.
[0107] In prime editing, the genome is edited using (1) a prime editor protein containing a Cas protein and polymerase (such as reverse transcriptase), and (2) prime editing guide RNA (pegRNA) to introduce the desired edit to the target region of the target DNA molecule. Various embodiments of prime editing are disclosed in PCT application number PCT / US2020 / 023712 (publication number WO2020191233A1), the entire contents of which are incorporated herein by reference.
[0108] Prime editing, a highly versatile and precise genome editing method developed by David R. Liu et al., uses a prime editor protein containing the Cas protein to directly write new genetic information into target regions of DNA molecules (such as genomes). Prime editing is a novel platform genome editing method. Prime editing primarily uses the Cas protein, polymerase, and pegRNA, where the pegRNA has an extension arm attached to a conventional guide RNA. In this case, the extension arm contains an extension region. The extension region contains an editing template that serves as a template for the desired edit so that the desired edit is inserted into the target region. In this case, the insertion of the desired edit into the target region is carried out through a number of processes, including polymerization by a polymerase (such as reverse transcriptase) linked to the Cas protein. Polymerization by polymerase is performed on the spacer-unbound strand, using the DNA synthesis template contained in the extension region of the pegRNA as the polymerization template.
[0109] For example, in Prime Edit version 2 (PE2), a nick is created on the spacer-unbound strand (induced and / or created by the Cas protein contained in the PE2 Prime Editor protein), and then reverse transcriptase polymerization (reverse transcription) occurs on the DNA synthesis template strand in the 5'→3' direction from the site where the nick was created relative to the spacer-unbound strand. Reverse transcription is performed using the DNA synthesis template contained in the elongation region as the template for reverse transcription. In such a polymerization process, a sequence complementary to part or all of the DNA synthesis template is encoded at the site on the spacer-unbound strand where the nick was created. The thus encoded sequence forms a 3' DNA flap. The 3' DNA flap contains an edit, which has a DNA sequence complementary to the edit template contained in the DNA synthesis template. Subsequently, the 5' DNA flap is removed via a 5' DNA flap cleavage process (which may involve, for example, the 5' DNA flap endonuclease FEN1). Additionally, the desired edit is incorporated into the desired site through the processes of ligation of the 3' DNA flap and repair and / or replication of the cellular DNA. The process of editing DNA molecules with Prime Editing Version 2 (PE2) is described in detail in [Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.], the entire contents of which are incorporated herein by reference.
[0110] The term "editing" as used in relation to prime editing may refer to edits incorporated into the DNA molecule as a result of the prime editing system. For example, "editing" may be used to refer to edits incorporated into the spacer-unbound strand, edits incorporated into the spacer-bound strand, and / or edits incorporated into the double helix. This is because edits placed on the 3' flap are ultimately placed on the spacer-unbound and spacer-bound strands through processes including ligation of the 3' DNA flap and repair and / or replication of cellular DNA as described above. Editing may include one or more nucleotide insertions, one or more nucleotide deletions, and one or more nucleotide substitutions of other nucleotides, either individually or in combination.
[0111] For example, editing may involve the insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotides to be inserted may be located consecutively or not in the nucleic acid. For example, editing may include the deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotide to be deleted may be located consecutively or not in the nucleic acid. For example, editing may include substitutions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotides to be substituted may be located consecutively or not in the nucleic acid. In another example, editing may include the above insertions and substitutions. In another example, editing may include the above deletions and substitutions. In yet another example, editing may include the above insertions, deletions and substitutions.[Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.] The publication that first disclosed prime editing by David R. Liu et al. described the scope of prime editing as "all four transition point mutations; all eight transposition point mutations; insertions (1 bp ~ ≥ 44 bp); deletions (1 bp ~ ≥ 80 bp); and combinations of the above," indicating that there are diverse forms of editing that can be placed in DNA molecules through prime editing. Furthermore, prime editing technology is still under development and improvement, and therefore the scope of prime editing is not limited to what is disclosed in the above-mentioned literature. [Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.] (Its entire contents are incorporated herein by reference) describes prime editing as a highly versatile and precise genome editing method that directly "writes" new genetic information into specific DNA regions. From this perspective, descriptions herein regarding genetic information that can be inserted or placed into DNA through prime editing should not be interpreted restrictively.
[0112] In some cases, prime editing can be considered a “surveillance and replacement” genome editing technique. This is because the prime editor (or prime editor complex) used to perform prime editing can explore and locate the desired target region to be edited while replacing the corresponding endogenous DNA strand of the target region with a substitute strand containing the desired edit. PCT application number PCT / US2020 / 023712 (publication number WO2020191233A1) (its entire contents are incorporated herein by reference) discloses that the prime editor described in the above document is not limited to reverse transcriptase, and such reverse transcriptase is merely one type of DNA polymerase that may be used in prime editing. Therefore, whenever reverse transcriptase is mentioned, those skilled in the art will understand that any suitable DNA polymerase may be used instead of reverse transcriptase. Similarly, it will be well understood that not only Cas9, nCas9, etc., but also proteins or domains functionally equivalent to Cas9 may be used in prime editing.
[0113] A guide RNA (i.e., pegRNA) specialized for prime editing is complexed with a Cas protein (e.g., a fusion protein containing a Cas protein) and ultimately undergoes a prime editing process to place the desired edit at a target site in a target region of a DNA molecule (such as a genome). The pegRNA contains an editing template for delivering the desired information to the target DNA. A substitute strand containing a sequence corresponding to the editing template is prepared from the editing template and used to replace the corresponding endogenous DNA strand. The prime editing mechanism may include a step of nicking the target region in a single strand of DNA, exposing a 3'-hydroxyl group, and delivering information from the pegRNA to the target DNA. Subsequently, the prime editing mechanism includes a step of delivering the desired information into the target region using the exposed 3'-hydroxyl group via a DNA polymerization process based on a sequence capable of delivering the desired information to the pegRNA. In various embodiments, the elongation region providing the polymerization template for the editing-containing substitute strand may be formed from RNA or DNA. In the case of RNA elongation regions, the polymerase used for prime editing may be an RNA-dependent DNA polymerase (such as reverse transcriptase). In the case of DNA elongation regions, the polymerase used for prime editing may be a DNA-dependent DNA polymerase. The newly synthesized strand by prime editing (i.e., the replacement DNA strand containing the desired edit) may be homologous to the genomic target sequence, except that it contains the desired nucleotide modifications. The newly synthesized DNA strand is sometimes called a single-stranded DNA flap (such as a 3' single-stranded DNA flap), which replaces the corresponding endogenous strand. In various embodiments, prime editing works by contacting the target DNA molecule with a Cas protein (in this case, the Cas protein is contained in the prime editor protein) complexed with prime editing guide RNA (pegRNA). An example of editing a DNA molecule (such as a genome) by prime editing can be described as follows: After contacting the nCas9 (which may be contained in, for example, the prime editor protein) / pegRNA complex with the DNA molecule, the pegRNA guides nCas9 to bind to the target region.A nick is introduced into one strand of the DNA strand in the target region (nCas9 introduces the nick), thus preparing an available 3' end on one strand of the DNA strand. The available 3' end is located in the target region. In certain embodiments, the nick may be made into a strand that does not hybridize with a portion of the pegRNA sequence, i.e., a spacer-unbound strand. In other specific embodiments, the nick may be made into a strand that hybridizes with a portion of the pegRNA sequence, i.e., a spacer-bound strand. The region located at the 3' end of the DNA strand formed by Cas9 nickase nicking (the region located upstream of the nick site) interacts with a portion of the pegRNA elongation region for reverse transcription priming. In certain embodiments, the 3' end of the DNA strand hybridizes to a primer-binding site (PBS) or reverse transcriptase priming sequence contained in the pegRNA elongation region. A single strand of DNA is synthesized by reverse transcriptase (which may be contained in, for example, a prime-edit fusion protein) in the direction from the 3' end of the primed region toward the 5' end of the pegRNA. In other words, a single strand of DNA is synthesized in the 5'→3' direction relative to a spacer-unbound strand (PAM-containing sequence) hybridized to the primer-binding site. The single strand of DNA thus synthesized contains the desired nucleotide modification (one or more base modifications, one or more insertions, one or more deletions, or a combination thereof). The synthesized single strand of DNA is sometimes called a 3' single-stranded DNA flap. When the 3' single strand enters the endogenous DNA, the formed (unedited) 5' endogenous DNA flap is removed. The removal of the 5' endogenous DNA flap can be achieved via a 5' flap cleavage process. The 3' single-stranded DNA flap that has entered the endogenous DNA is ligated. Then, DNA repair takes place. As a result, the desired edit is fully incorporated into the target region.
[0114] The objective of a prime editing system can be achieved by factors including, for example, prime editor proteins and pegRNA.
[0115] The following describes the prime editor proteins and pegRNAs used in prime editing.
[0116] Prime Editor Protein Overview of Prime Editor Proteins In some embodiments, a prime editor protein (or prime edit construct) means a construct in the form of a complex or fusion protein comprising a Cas protein and a polymerase. Prime editor proteins may also be referred to as prime edit proteins, prime edit constructs, prime edit enzymes, prime editor enzymes, and prime edit fusion proteins. Prime editor proteins may include structures represented as [Cas]-[P] or [P]-[Cas], where "P" refers to any polymerase (such as reverse transcriptase) or an element derived therefrom, and "Cas" refers to a Cas protein (such as Cas9 nickase or an SpCas9 variant including wild-type SpCas9) or an element derived therefrom. The "[]-[" or "-" indicating that the Cas protein and polymerase are linked may refer to any linker or other element, or linkage, that has the function of covalently or noncovalently linking the Cas protein and polymerase.
[0117] As described above, the prime editor protein comprises a Cas protein (such as Cas9 nickase) and a reverse transcriptase (or DNA polymerase). The prime editor protein may be in the form of a fusion protein consisting of one molecule, or in the form of a complex consisting of two or more molecules, but is not particularly limited. Prime editing can be performed on a target region by the prime editor protein in the presence of pegRNA. The prime editor protein forms a complex with pegRNA, in which case the complex is sometimes called a prime editor protein / pegRNA complex. In some embodiments, the prime editor protein may be called a prime editing protein.
[0118] In some embodiments, the term “prime editing system” may refer to a prime editor protein and pegRNA, or to the editing of DNA molecules performed using a prime editor protein and pegRNA.
[0119] As described above, the term “prime editing system” can be used as a comprehensive concept to describe anything related to prime editing. In some embodiments, the prime editing system may further include other elements or their use in addition to the prime editor protein and pegRNA. For example, the prime editing system may further include a conventional guide RNA or its use that can direct the nicking of a second site in the unedited strand.
[0120] In some embodiments, the prime editor protein is (i) Cas protein; and (ii) polymerase Includes.
[0121] The following describes the Cas protein and polymerase contained in the prime editor protein.
[0122] Prime Editor Protein Element 1 - Cas Protein Prime editor proteins include Cas proteins and polymerases. Prime editor proteins may include Cas proteins, which are described in detail in the "CRISPR / Cas System" section. The term Cas9 protein is used to include its equivalents. Cas proteins are sometimes called CRISPR enzymes, nucleic acid programmable DNA binding proteins (napDNAbp), or CRISPR proteins.
[0123] In some embodiments, the Cas protein is Cas12a, Cas12b1 (C2c1), Cas12c (C2c3), Cas12e (CasX), Cas12d (CasY), Cas12g, Cas12h, Cas12i, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, The Cas protein may be, but is not limited to, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cas13a (C2c2), Cas13b, Cas13c, Cas13d, Cas14, xCas9, cyclic displacement Cas9, Argonaut (Ago) domain, or fragments, homologs, or variants thereof. In some embodiments, the Cas protein may be a Cas protein having nickase activity. The Cas protein having nickase activity may be, but is not limited to, Cas9 nickase or Cas12 nickase (e.g., Cas12a nickase, Cas12b1 nickase, etc.). In some embodiments, the Cas protein may be a Cas protein having nuclease activity. In some embodiments, the Cas protein may contain one or more amino acid substitutions or amino acid variations in the HNH domain and / or RuvC domain.
[0124] For example, a variant may contain an amino acid sequence having approximately 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% sequence identity compared to the amino acid sequence of the wild-type Cas protein or the parental Cas protein. For example, a variant may include one or more insertions, one or more deletions, one or more substitutions, or a combination thereof, compared to the amino acid sequence of the wild-type Cas protein or the parental Cas protein. For example, the Cas protein may be Cas9 derived from Streptococcus pyogenes (SpCas9), Cas9 derived from Campylobacter jejuni (CjCas9), Cas9 derived from Staphylococcus aureus (SaCas9), or a variant thereof. For example, the Cas protein may be SpyMac, iSpymac, GeoCas9, xCas9, cyclic substitution Cas9, or a variant thereof. For example, a SpCas9 variant may include variations in amino acid residues, such as one or more insertions, one or more deletions, one or more substitutions, or combinations thereof, compared to the amino acid sequence of wild-type SpCas9. For example, a SpCas9 variant containing the H840A substitution provides a Cas protein with nickase activity. For example, a SpCas9 variant containing the D10A substitution provides a Cas protein with nickase activity. For example, a SpCas9 variant may include R221K and N394K substitutions. For example, a SpCas9 variant may have a form in which one or more amino acid residues selected from D10, R221, L244, N394, H840, K1211, and L1245 of wild-type SpCas9 are substituted with other amino acid residues.For example, a SpCas9 variant may include one or more selected from D10A, R221K, L244Q, N394K, H840A, K1211Q, and L1245V. In some embodiments, the Cas protein is a SpCas9 variant having nickase activity containing H840A; a SpCas9 variant having nickase activity containing R221K, N394K, and H840A (see Chen, Peter J., et al. "Enhanced prime editing systems by manipulating cellular determinants of editing outcomes." Cell 184.22(2021):5635-5652.); or a wild-type SpCas9 variant having nuclease activity (i.e., inducing DSBs) (see Adikusuma, Fatwa, et al. "Optimized nickase-and nuclease-based prime editing in human and mouse cells." Nucleic acids research See 49.18(2021):10785-10795.); or SpCas9 variants having nuclease activity including R221K and N394K, but are not limited to these. In some embodiments, the Cas protein may be codon-optimized. In some embodiments, the prime editor protein may include a PAM-less Cas protein.
[0125] Various examples of Cas proteins that may be included in prime editor proteins are described in detail in [U.S. Patent Application No. 17 / 219,672].
[0126] In some embodiments, wild-type SpCas9 may contain the following amino acid sequence of SEQ ID NO: 28:
[0127] In some embodiments, wild-type SpCas9 variants, including the H840A variation, may contain the following amino acid sequence of SEQ ID NO: 29:
[0128] In some embodiments, wild-type SpCas9 variants, including the R221K and N394K variations, may contain the following amino acid sequence of SEQ ID NO: 30:
[0129] In some embodiments, wild-type SpCas9 variants, including the R221K, N394K, and H840A variations, may include the following amino acid sequence of SEQ ID NO: 31:
[0130] Prime Editor Protein Element 2 Polymerase Overview of polymerases used in prime editing Prime editor proteins include Cas proteins and polymerases. Polymerase refers to an enzyme or protein that synthesizes nucleotide chains and may be used in relation to the prime editing systems or prime editing-based systems described herein. A polymerase is a “template-dependent polymerase” (i.e., a polymerase that synthesizes nucleotide chains based on the order of nucleotide bases in a template chain). A polymerase may also be a “template-independent” polymerase. Polymerases may also be classified as “DNA polymerases” or “RNA polymerases.”
[0131] In various embodiments, the prime editing system or prime editor protein may include a DNA polymerase that synthesizes DNA strands.
[0132] In some embodiments, the DNA polymerase may be a DNA-dependent DNA polymerase, in which case the pegRNA may contain a DNA template that acts as a template for polymerization by the DNA-dependent DNA polymerase. In this case, the pegRNA may be called a hybrid pegRNA or chimera, containing an RNA portion (a guide RNA component including a spacer and a gRNA core) and a DNA portion (a DNA template).
[0133] In various embodiments, the DNA polymerase may be an "RNA-dependent DNA polymerase." In this case, the pegRNA may contain an RNA template that acts as a template for polymerization by the RNA-dependent DNA polymerase. In other words, the pegRNA may consist of RNA components and may contain an RNA elongation region.
[0134] Polymerase can also refer to an enzyme that catalyzes the polymerization of nucleotides. Typically, polymerization by polymerase begins at the 3' end of a primer annealed to a polynucleotide template sequence (such as a primer sequence annealed to the primer-binding site of pegRNA in prime editing) and proceeds toward the 5' end of the template strand. DNA polymerase can catalyze the polymerization of deoxynucleotides. As used herein, the term polymerase may be used to include the meanings of enzymes, proteins, variants thereof, and fragments thereof that catalyze and / or perform nucleotide polymerization. In this case, a polymerase fragment refers to any portion of wild-type or mutant (variant) DNA polymerase that contains an amino acid sequence shorter than the full length of wild-type polymerase and has the ability to catalyze and / or perform deoxynucleotide polymerization under at least one condition. Such fragments may exist as separate entities or constitute a larger polypeptide, such as a fusion protein.
[0135] Example of polymerase: reverse transcriptase For example, a polymerase, which is one of the elements used in prime editing, may be a reverse transcriptase (RT). Reverse transcriptase refers to a class of polymerases characterized as RNA-dependent DNA polymerases. All known reverse transcriptases require primers to synthesize a DNA transcript from an RNA template. Where used herein, the term reverse transcriptase may be used to include the meanings of its variants and fragments. For example, a variant may contain an amino acid sequence having approximately 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9% sequence identity compared to the amino acid sequence of wild-type reverse transcriptase or parental reverse transcriptase. For example, a variant may include one or more insertions, one or more deletions, one or more substitutions, or a combination thereof, compared to the amino acid sequence of wild-type reverse transcriptase or parental reverse transcriptase.
[0136] Reverse transcriptase can originate from many different sources. Examples of reverse transcriptase sources include, but are not limited to, Moloney's mouse leukemia virus (M-MLV or MLVRT), human T-cell leukemia virus type 1 (HTLV-1), bovine leukemia virus (BLV), Roussarcoma virus (RSV), human immunodeficiency virus (HIV), yeasts such as Saccharomyces, Neurospora, and Drosophila, primates, and rodents.
[0137] Examples of reverse transcriptases include reverse transcriptase from avian myeloblastosis virus (AMV), reverse transcriptase from Moloney mouse leukemia virus (M-MLV) (see [GERARD, GARY F., et al. "Influence on stability in Escherichia coli of the carboxy-terminal structure of cloned Moloney murine leukemia virus reverse transcriptase." Dna 5.4(1986):271-279.; and Kotewicz, Michael L., et al. "Cloning and overexpression of Moloney murine leukemia virus reverse transcriptase in Escherichia coli." Gene 35.3(1985):249-258.]), and RNase. Examples of reverse transcriptases that substantially lack H activity include, but are not limited to, M-MLV reverse transcriptase (see application number US07 / 671,156, publication number US5244797A), human immunodeficiency virus (HIV) reverse transcriptase, avian sarcoma leukemia virus (ASLV) reverse transcriptase, Rous sarcoma virus (RSV) reverse transcriptase, avian erythroblastosis virus (AEV) helper virus MCAV reverse transcriptase, avian myelocytosis virus MC29 helper virus MCAV reverse transcriptase, avian reticuloendotheliopathy virus (REV-T) helper virus REV-A reverse transcriptase, avian sarcoma virus UR2 helper virus UR2AV reverse transcriptase, avian sarcoma virus Y73 helper virus YAV reverse transcriptase, Rous-associated virus (RAV) reverse transcriptase, myeloblastosis-associated virus (MAV) reverse transcriptase, their variants, or their fragments. In some embodiments, the reverse transcriptase may be a retroviral reverse transcriptase. In some embodiments, the reverse transcriptase may be an error-prone reverse transcriptase. An error-prone reverse transcriptase (or any polymerase in the broad sense) refers to a reverse transcriptase derived from another reverse transcriptase that has a lower error rate than the naturally occurring or wild-type M-MLV reverse transcriptase.Reverse transcriptases prone to errors may have a higher error rate than the wild-type reverse transcriptase being compared to them. For example, 6.7 × 10⁻⁶. -5 , 7.14×10 -5 , 7.7×10 -5 , 9.1×10 -5 , or 1 × 10 -4 The error rate can be obtained. [See Bebenek, K., et al. "Error-prone polymerization by HIV-1 reverse transcriptase. Contribution of template-primer misalignment, miscoding, and termination probability to mutational hot spots." Journal of Biological Chemistry 268.14(1993):10324-10334.; and Sebastian-Martin, Alba, Veronica Barrioluengo, and Luis Menendez-Arias. "Transcriptional inaccuracy threshold attenuates differences in RNA-dependent DNA synthesis fidelity between retroviral reverse transcriptases." Scientific Reports 8.1(2018):1-13.], the entire contents of each of these are incorporated herein by reference.
[0138] In some embodiments, the reverse transcriptase may be M-MLV reverse transcriptase. The term M-MLV reverse transcriptase may be used to include its variants and fragments. Examples of M-MLV reverse transcriptase include wild-type M-MLV reverse transcriptase, M-MLV reverse transcriptase variants, wild-type M-MLV reverse transcriptase fragments, or fragments of wild-type M-MLV reverse transcriptase variants. For example, an M-MLV reverse transcriptase variant may have a form in which one or more amino acid residues selected from P51, S67, E69, L139, T197, D200, H204, F209, E302, E302, T306, F309, W313, T330, L345, L435, N454, D524, E562, D583, H594, L603, E607, and D653 of wild-type M-MLV reverse transcriptase or other wild-type reverse transcriptase are substituted with another amino acid residue. The amino acid sequence of wild-type M-MLV reverse transcriptase is disclosed in SEQ ID NO: 26. For example, an M-MLV reverse transcriptase variant may contain one or more amino acid variations selected from P51L, S67K, E69K, L139P, T197A, D200N, H204R, F209N, E302K, E302R, T306K, F309N, W313F, T330P, L345G, L435G, N454K, D524G, E562Q, D583N, H594Q, L603W, E607K, and D653N (where the base sequence for the amino acid variations is the amino acid sequence of wild-type M-MLV reverse transcriptase SEQ ID NO: 26). In certain embodiments, the reverse transcriptase may be an M-MLV reverse transcriptase variant containing the amino acid variations D200N, T306K, W313F, T330P, and L603W (e.g., an M-MLV reverse transcriptase quintuple mutant). In certain embodiments, the reverse transcriptase may be a terminal-cleaved M-MLV reverse transcriptase. In this case, the terminal-cleaved M-MLV reverse transcriptase may contain four mutations (D200N, T306K, W313F, and T330P). In this case, the L603W mutation present in the above M-MLV reverse transcriptase quintuple mutant is no longer present due to terminal cleavage. In some embodiments, the polymerase or reverse transcriptase may be codon-optimized.
[0139] Reverse transcriptase (RT) genes (or the genetic information contained therein) can be obtained from many different sources. For example, the genes can be obtained from eukaryotic cells infected with retroviruses or from various plasmids containing part or all of a retroviral genome. In addition, messenger RNA-like RNA containing the RT gene can be obtained from retroviruses. Various examples of reverse transcriptases that may be contained in prime editor proteins are described in detail in [U.S. Patent Application No. 17 / 219,672].
[0140] In some embodiments, the wild-type M-MLV reverse transcriptase may contain the following amino acid sequence of SEQ ID NO: 26: (Sequence ID 26).
[0141] In some embodiments, wild-type M-MLV reverse transcriptases, including the D200N, T306K, W313F, T330P, and L603W variations, may contain the following amino acid sequence of SEQ ID NO: 27: (Sequence ID 27).
[0142] Elements that may be additionally included in prime editor proteins The prime editor protein comprises a Cas protein and a polymerase (such as reverse transcriptase). In some embodiments, the prime editor protein may further include additional elements in addition to the two elements, such as one or more linkers (such as linkers for linking the elements contained in the prime editor protein) and one or more nuclear localization sequences (NLS).
[0143] Prime editor proteins may contain one or more linkers. For example, a linker may be used to link a Cas protein to another structure contained in the prime editor protein. The linker may be any linker known in the art. For example, a linker may be used to link a polymerase to another structure contained in the prime editor protein. For example, a linker may be used to link an NLS to another structure contained in the prime editor protein. For example, a linker may be used to link a Cas protein and a polymerase. For example, a linker may be used to link a linker to another linker selected independently. In some embodiments, the linker may be a covalent bond, an organic molecule, a group, a polymer, or a chemical moiety. In some embodiments, each linker may be selected independently. The linker may have a length of 3 to 100 or more amino acids. For example, the linker may be approximately 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 amino acids in length, or it may be a length within the range set by any two values selected from the above values. In some embodiments, the linker may include the following amino acid sequences: one or more G, one or more XP (where X is any amino acid), one or more EAAAK (SEQ ID NO: 35), one or more GGS (SEQ ID NO: 36), one or more SGGS (SEQ ID NO: 37), or one or more GGGGS (SEQ ID NO: 38). In some embodiments, the linker may include, but is not limited to, the amino acid sequences SGSETPGTSESATPES (SEQ ID NO: 39) or SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 40).In some embodiments, the linker may be an XTEN linker (e.g., linker XTEN16). As described above, the prime editor protein may include one or more linkers, each of which may be independently selected or determined. Various examples of linkers are described in detail in [U.S. Patent Application No. 17 / 219,672].
[0144] The prime editor protein may contain one or more NLSs. In some embodiments, the prime editor protein may contain two or more NLSs. If the prime editor protein contains multiple NLSs, each NLS may be independently selected or determined. The NLS may be any NLS known in the art. The NLS may be any later discovered NLS for nuclear localization. The NLS may be any naturally occurring NLS or any non-naturally occurring NLS (e.g., having one or more mutations).In some embodiments, the NLS may include, but is not limited to, the following: SV40 virus large T antigen NLS having the amino acid sequence PKKKRKV (SEQ ID NO: 1); binocular SV40 NLS having the amino acid sequence KRTADGSEFESPKKKRKVE (SEQ ID NO: 18) (or binocular SV40 NLS having one amino acid deletion in the portion other than PKKKRKV); NLS from nucleoplasmin (such as binocular nucleoplasmin NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 2)); c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 3) or RQRRNELKRSP (SEQ ID NO: 4); hRNPA1 M9 having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 5) NLS; IBB domain sequence of RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 6) from importin α; myoma T protein sequences of VSRKRPRP (SEQ ID NO: 7) and PPKKARED (SEQ ID NO: 8); human p53 sequence of PQPKKKPL (SEQ ID NO: 9); mouse c-abl IV sequence of SALIKKKKKKMAP (SEQ ID NO: 10); influenza virus NS sequences of DRLRR (SEQ ID NO: 11) and PKQKKRK (SEQ ID NO: 12); hepatitis virus delta antigen sequence of RKLKKKIKKL (SEQ ID NO: 13); mouse Mx1 protein sequence of REKKKFLKRR (SEQ ID NO: 14); human poly(ADP-ribose) polymerase sequence of KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15); or NLS sequence derived from the steroid hormone receptor (human) glucocorticoid sequence of RKCLQAGMNLEARKTKK (SEQ ID NO: 16). In some embodiments, NLS can be codon-optimized.
[0145] Various examples of NLS are described in detail in [U.S. Patent Application No. 17 / 219,672].
[0146] Prime Editing Guide RNA (pegRNA) Overview of pegRNA As used herein, the terms “prime editing guide RNA,” “pegRNA,” or “elongation guide RNA” refer to a form of guide RNA specialized to include one or more additional sequences so that the prime editing methods and compositions disclosed herein can be carried out. pegRNA is used in conjunction with a prime editor protein in a prime editing system. As described herein, pegRNA includes an elongation arm or elongation region. The elongation arm may, but is not limited to, a single-stranded RNA sequence and / or a DNA sequence. As described above, guide RNA used in conventional CRISPR / Cas systems (i.e., guide RNA that does not include the elongation arm of pegRNA) may be called conventional guide RNA and thus distinguished from pegRNA. For example, the elongation arm may be formed at the 3' end of conventional guide RNA. In another example, the elongation arm may be formed at the 5' end of conventional guide RNA. In some embodiments, pegRNA may include a spacer region, a gRNA core, and an elongation arm formed at the 3' or 5' end of conventional guide RNA.
[0147] Extension arm The term "elongation arm" refers to a portion of the pegRNA nucleotide sequence that provides various functions, including a DNA synthesis template (e.g., including an editing template) and a primer binding site (PBS) for polymerases (such as reverse transcriptase). In pegRNA, the elongation arm may be described as an elongation region. In some embodiments, the elongation arm may be located at the 3' end of the guide RNA. In some embodiments, the elongation arm located at the 3' end of the guide RNA may be called the 3' elongation arm. In other embodiments, the elongation arm may be located at the 5' end of the guide RNA. In some embodiments, the elongation arm located at the 5' end of the guide RNA may be called the 5' elongation arm. In some embodiments, the elongation arm may include a homology arm. In some embodiments, the elongation arm may include an editing template. In some embodiments, the elongation arm may include a primer binding site. In various embodiments, the elongation arm (such as the 3' elongation arm) may include the following elements in the 5'→3' direction: a DNA synthesis template and a primer binding site. In other words, based on the entire pegRNA, the pegRNA may include the following elements in the 5'→3' direction: a spacer, a gRNA core, a DNA synthesis template, and a primer binding site. The DNA synthesis template may include a homology region and an editing template. In various embodiments, the elongation arm may include the following elements in the 5'→3' direction: a homology region, an editing template, and a primer binding site. In other words, based on the entire pegRNA, the pegRNA may include the following elements in the 5'→3' direction: a spacer, a gRNA core, a homology region, an editing template, and a primer binding site. In some embodiments, the 5' elongation arm may include the following elements in the 5'→3' direction: a DNA synthesis template and a primer binding site.
[0148] The polymerization activity of reverse transcriptase, an example of a polymerase, is located in the 5'→3' direction relative to the strand that will ultimately bind to the template strand. When a primer is annealed to the primer-binding site (PBS), the reverse transcriptase polymerizes a single strand of DNA using the complementary template strand (DNA synthesis template) as the reverse transcription template. Various embodiments of the extension arms used in prime editing are described in detail in [U.S. Patent Application No. 17 / 219,672].
[0149] Typically, the elongation arm of a pegRNA may be described as comprising two regions, for example, a primer-binding site (PBS) and a DNA synthesis template (such as a reverse transcription template). For example, in PE2, the primer-binding site binds to a primer sequence formed from an endogenous DNA strand at a nicking target site induced by a prime editor protein, thereby exposing the 3' end of the strand to be nicked. As described herein, binding of the primer sequence to the primer-binding site in the elongation arm of a pegRNA generates a double-stranded region with the exposed 3' end (i.e., the 3' end of the primer sequence), which then provides a matrix that allows reverse transcriptase to polymerize a single strand of DNA from the exposed 3' end along the length of the DNA synthesis template. The sequence of the resulting single-stranded DNA product is complementary to the DNA synthesis template. Polymerization continues toward the 5' end of the DNA synthesis template (or elongation arm) until polymerization is terminated. Accordingly, the DNA synthesis template is encoded by the polymerase of the prime editor protein into a resulting single-stranded DNA product (i.e., a 3' single-stranded DNA flap containing the desired gene editing information). As a result, a 3' single-stranded DNA flap (e.g., complementary to the DNA synthesis template) is formed, replacing the endogenous DNA strand corresponding to the target region located immediately downstream of the PE-induced nick site. Polymerization of the DNA synthesis template may, but is not limited to, continue toward the 5' end of the elongation arm until polymerization is terminated. Polymerization may be terminated by a variety of means, including, but not limited to, (a) reaching the 5' end of pegRNA, (b) reaching an impassable RNA secondary structure (such as a hairpin or stem / loop), or (c) reaching a replication termination signal such as a specific nucleotide sequence that blocks or inhibits polymerase, or a nucleic acid phase signal such as superhelical DNA or RNA. However, the termination of polymerization is not limited to these. Several prime editing-related documents have reported that sequences homologous to parts of the pegRNA gRNA core are found in the 3' DNA flap or editing site.Based on this, it will be understood by those skilled in the art that the above embodiments are merely illustrative and that the termination of polymerization is not limited to the above embodiments.
[0150] Primer binding site (PBS) In a prime editing system, information present in the DNA synthesis template contained in pegRNA is delivered to the endogenous strand through polymerization by polymerase. Polymerization by polymerase requires the binding of a primer to the template strand, and DNA polymerization becomes possible through primer binding or annealing. In a prime editing system, a portion of the region where a DSB or nick is created, induced by the Cas protein, is used as a primer. For example, in the explanation based on PE2, a portion of the region located upstream of the nick on the spacer-unbound strand, induced by the Cas protein of the prime editor protein, is used as a primer. In this case, the region designed to bind complementaryly to the sequence of the region located upstream of the nick is called the primer binding site, and the primer binding site is located in the elongation region of pegRNA. The prime editing process in PE2 is further explained below. Once the primer binding site and the region of endogenous DNA (such as the genome) used as a primer are bound, reverse transcriptase performs reverse transcription using the primer as a template for reverse transcription. In this case, it will be apparent to those skilled in the art that reverse transcription is performed in the 3'→5' direction relative to the reverse transcription template strand (i.e., pegRNA). Once reverse transcription is performed, a sequence complementary to the DNA template sequence is incorporated into the 3' flap of the genomic DNA, and information about the DNA template is delivered to the 3' flap by reverse transcription. Subsequently, through processes including 5' flap removal and cellular DNA repair and / or replication, the information about the DNA template is finally delivered to another desired strand of the DNA to be edited. The desired outcome of prime editing is to deliver or place information about the DNA template on the first strand (in this case, the first strand is the spacer-unbound strand) and / or the second strand (in this case, the second strand is the spacer-bound strand) of the desired site to be edited. In other words, as a result of exemplary PE2 prime editing, a DNA sequence complementary to the DNA template sequence is present at the desired site on the first strand, and the same DNA sequence as the DNA template sequence is present at the desired site on the second strand.
[0151] In some embodiments, the primer binding site of pegRNA may be designed as a sequence complementary to the sequence of a region located upstream of the site where a nick or DSB is created in the DNA molecule (such as genomic DNA). In some embodiments, the primer binding site may be designed as a sequence complementary to the sequence of a region located upstream of the site where a nick or DSB is created in the spacer-unbound strand of the DNA molecule. In other words, the sequence of a region located upstream of the site where a nick or DBS is created in the spacer-unbound strand of the DNA molecule can function as a primer in the prime editing process. As described above, in the example of PE2, the sequence located in the 5' direction of the nick functions as a primer, and the nick end of the DNA molecule is exposed to reverse transcriptase through the binding of the primer to the primer binding site.
[0152] In some embodiments, the primer may have a length of 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, or 50nt or greater, or a length within a range formed by two values selected from the above values. In a particular embodiment, the primer may have a length of 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, or 25nt, or it may have a length within a range formed by two values selected from the above values.
[0153] In some embodiments, the primer binding site may have a length of 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, or 50nt or greater, or a length within the range formed by two values selected from the above values. In certain embodiments, the primer binding site may have a length of 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, or 25nt, or a length within the range formed by two values selected from the above values. The length of the primer binding site may be appropriately selected depending on the purpose and is not particularly limited.
[0154] DNA synthesis template As used herein, the term “DNA synthesis template” in prime editing refers to a region or portion that is used as a template strand by the polymerase of the prime editor protein to encode a 3' single-stranded DNA flap containing the desired edit. Furthermore, the term “DNA synthesis template” refers to a region or portion contained within the elongation region of pegRNA, which is replaced at the target site with the corresponding endogenous DNA strand through the prime editing mechanism. Various embodiments of the DNA synthesis template and the elongation region of pegRNA are described in detail in [U.S. Patent Application No. 17 / 219,672], which is incorporated herein by reference in its entirety.
[0155] The extension region containing the DNA synthesis template may consist of DNA, RNA, or a DNA / RNA hybrid. In the case of RNA, the polymerase of the prime editor protein may be an RNA-dependent DNA polymerase (such as reverse transcriptase). The DNA synthesis template may be called a DNA polymerization template or reverse transcription template (RT template), in which case the RT template is intended for use with reverse transcriptase in the prime editing system. In the case of DNA, the polymerase of the prime editor may be a DNA-dependent DNA polymerase. In various embodiments, the DNA synthesis template (such as an RT template) may include an "editing template" and a "homologous region".
[0156] In some embodiments, the DNA synthesis template may include part or all of an optional 5' end modification region e2, in addition to the editing template and homology region. Depending on the nature of region e2 (e.g., whether it includes secondary structures such as a hairpin, T-loop, or stem / loop), the polymerase may not encode the e2 region at all, or it may encode part or all of it. In some embodiments, for a 3' elongation arm, the DNA synthesis template may include part of the elongation arm covering from the 5' end of the primer binding site (PBS) to the 3' end of the gRNA core. In other embodiments, for a 5' elongation arm, the DNA synthesis template may include part of the elongation arm covering from the 5' end of the pegRNA molecule to the 3' end of the primer binding site. Preferably, the DNA synthesis template excludes the primer binding site (PBS) of pegRNA having a 3' elongation arm or a 5' elongation arm.
[0157] In certain embodiments described herein, the DNA synthesis template may be referred to as a “reverse transcription template (RT template)” which includes an editing template and homology arms. The RT template may refer to a portion of the sequence of pegRNA elongation arms used as a template in DNA synthesis. The term “RT template” may be used interchangeably with the DNA synthesis template.
[0158] In transpriming, the primer-binding site (PBS) and DNA synthesis template can be engineered using individual molecules called transpriming RNA templates (tPERT) (see U.S. Patent Application No. 17 / 219,672).
[0159] DNA synthesis template element 1 - editing template The term "editing template" refers to a portion of an extension arm that encodes a desired edit in a single-stranded 3' DNA flap synthesized by a polymerase, such as DNA-dependent DNA polymerase or RNA-dependent DNA polymerase (e.g., reverse transcriptase). In other words, the editing template may be complementary to the desired edit. In some embodiments, a DNA synthesis template may include an editing template and a homology arm. In some embodiments, an RT template may include an editing template and a homology arm. The term "RT template" is equivalent to "DNA synthesis template." However, as used herein, an RT template is based on the use of a prime editor protein having a polymerase, i.e., reverse transcriptase, while a DNA synthesis template is more broadly based on the use of a prime editor protein having any polymerase.
[0160] The desired edit to be placed in the target region of the DNA molecule to be edited (such as a genome) may include one or a combination of the following: insertion of one or more nucleotides, deletion of one or more nucleotides, and substitution of one or more nucleotides with other nucleotides. For example, the edit may include insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotides to be inserted may be located consecutively or not in the nucleic acid. For example, editing may include the deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotide to be deleted may be located consecutively or not in the nucleic acid. For example, editing may include substitutions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotides to be substituted may or may not be located consecutively in the nucleic acid. In another example, editing may include the above insertions and substitutions. In another example, editing may include the above deletions and substitutions. In another example, editing may include the above insertions and deletions. In another example, editing may include the above insertions, deletions, and substitutions.
[0161] DNA synthesis template element 2 - homology arm (or homology region) The term “homologous arm” refers to a portion of the elongation arm that is incorporated into the target DNA region through the replacement of the endogenous strand. For example, in PE2 prime editing, the term “homologous arm” may refer to a portion of the elongation arm that encodes a single-stranded DNA flap portion encoded by reverse transcriptase. For example, in the PE2 system, the portion of the single-stranded DNA flap encoded by the homologous arm is complementary to the unedited strand of the target DNA (such as the spacer-bound strand). In other words, in PE2, the sequence of the homologous arm has a sequence complementary to the corresponding sequence located on the spacer-unbound strand of the target DNA and has a substantially identical sequence to the corresponding DNA sequence located on the spacer-bound strand. The homologous arm replaces the endogenous strand, facilitating the annealing of the single-stranded DNA flap and thus helping to set up edits on the target DNA molecule. By definition, the homologous arm is encoded by the polymerase of the prime editor disclosed herein and is therefore part of the DNA synthesis template.
[0162] Additional elements that may be included in pegRNA and engineered pegRNA (epegRNA) Engineered pegRNA (epegRNA) is described with reference to [Nelson, James W., et al. "Engineered pegRNAs improve prime editing efficiency." Nature biotechnology 40.3(2022):402-410.], the entire contents of which are incorporated herein by reference. EpegRNA, a type of pegRNA, may be used to refer to an improved pegRNA. Specifically, epegRNA refers to pegRNA having an RNA motif attached to its 3' or 5' end. In some embodiments, epegRNA may be pegRNA having an RNA motif (or an engineered RNA motif) attached to its 3' end. EpegRNA may, for example, include the following elements in the 5'→3' direction: a spacer, a gRNA core, a DNA synthesis template, a primer binding site, and an RNA motif.
[0163] David R. Liu et al. developed engineered pegRNAs (epegRNAs) by adding RNA motifs to the 3' end of pegRNAs to improve their stability and prevent degradation of the 3' extension region. Specifically, in the aforementioned literature, David R. Liu et al. disclosed epegRNAs in which a stable pseudoknot was additionally incorporated to the 3' end of conventional pegRNAs. Examples of pseudoknots include, but are not limited to, evopreQ1 (a modified prequeosine 1-1 riboswitch aptamer) and mpknot (a frameshift pseudoknot from Moloney mouse leukemia virus), as described in [Nelson, James W., et al. "Engineered pegRNAs improve prime editing efficiency." Nature biotechnology 40.3(2022):402-410.].
[0164] epegRNA can be used regardless of the type of prime editor protein used. For example, epegRNA can be used with a prime editor protein containing a prime edit version 2 (PE2) SpCas9 nickase. In another example, epegRNA can be used with a PE-nuclease-containing Cas9 having nuclease activity (i.e., DSB activity) to edit DNA molecules (such as genomes). As used herein, the term pegRNA includes embodiments of epegRNA, and any description of pegRNA should be interpreted as including content related to epegRNA unless otherwise specified.
[0165] In some embodiments, pegRNA may further include a 3' engineering region at its 3' end. PegRNA containing a 3' engineering region is sometimes called epegRNA. In other words, epegRNA may further include a 3' engineering region in addition to the elements of pegRNA. In some embodiments, the 3' engineering region may include an RNA protection motif. In certain embodiments, the RNA protection motif may include an RNA sequence. In certain embodiments, the RNA protection motif may include a DNA sequence. In certain embodiments, the RNA protection motif may include a DNA / RNA hybrid sequence. In certain embodiments, the RNA protection motif may include, but is not limited to, evopreQ1 or mpknot, and may include any other structure to prevent RNA degradation and enhance stability.
[0166] In some embodiments, the 3' engineering region may include an RNA protection motif and a linker for ligating the RNA protection motif. The linker functions to ligate the RNA protection motif and the primer binding site in the epegRNA. In some embodiments, the linker for ligating the RNA protection motif may include an RNA sequence. In some embodiments, the linker for ligating the RNA protection motif may include a DNA sequence. In some embodiments, the linker for ligating the RNA protection motif may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nt or greater in length, or in length within a range set by any two values selected from the above values. In some embodiments, the linker for ligating the RNA protection motif may be designed to avoid base-pair interactions between the linker and PBS, or between the linker and the pegRNA spacer. In some embodiments, the sequence of the linker for ligating the RNA protection motif may be designed considering the sequence of the target region of the target DNA molecule.
[0167] Below, we will illustrate various prime editing versions developed based on prime editor proteins and pegRNA, which are the fundamental elements of prime editing. Prime editing is not limited to the versions illustrated below.
[0168] Prime edited version example Summary of Prime Edited Version Example Based on the core mechanism of prime editing described above, various prime editing versions have been developed. For the convenience of technicians in the art, examples of prime editing versions will be described. The method for detecting off-targets in prime editing provided herein can, but is not limited to, the use of the prime editor protein, many types of pegRNA including epegRNA, and / or additional elements such as dnMLH1 in many prime editing versions exemplified below. Furthermore, the method for detecting off-targets in prime editing provided herein can be applied to the prime editing versions exemplified below and to new prime editing versions that may be developed later. Therefore, it should be noted that the prime editing versions exemplified below do not limit the scope of application of the method provided herein.
[0169] Prime Edit Version 1 (PE1) Prime Edit Version 1 (PE1) describes a Prime Edit system version that includes the use of the following elements: Prime editor proteins including SpCas9(H840A) and wild-type Moloney mouse leukemia virus reverse transcriptase (MMLV RT); and pegRNA.
[0170] In other words, the PE1 prime editor protein contains a Cas protein with nickase activity and wild-type MMLV RT. The PE1 prime editor protein has the form of a fusion protein in which the Cas protein and reverse transcriptase are linked via a linker.
[0171] The PE1 prime editor protein and pegRNA form a complex, thereby directing or executing the editing of DNA molecules in a target region (such as genome editing).
[0172] PE1 is described in detail in [Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.].
[0173] Prime Edit Version 2 (PE2) Prime Edit Version 2 (PE2) refers to a version of the Prime Edit system that includes the use of the following elements: Prime editor proteins including SpCas9(H840A) and MMLV RT(D200N+L603W+T330P+T306K+W313F); and pegRNA.
[0174] In other words, the PE2 prime editor protein contains a Cas protein with nickase activity and an MMLV RT quintuple mutant. The PE2 prime editor protein has the form of a fusion protein in which the Cas protein and reverse transcriptase are linked via a linker. Specifically, the PE2 prime editor protein has the following structure: [bpNLS(SV40)]-[SpCas9 H840A]-[SGGSX2-XTEN16-SGGSX2]-[MMLV RT quintuplet mutant]-[bpNLS(SV40)].
[0175] In this case, bpNLS(SV40) refers to the binocular SV40 NLS. MMLV RT quintuplet variant refers to an MMLV RT variant that includes the amino acid variations D200N, L603W, T330P, T306K, and W313F compared to wild-type MMLV RT.
[0176] The PE2 prime editing system is described in detail in [Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.; and Chen, Peter J., et al. "Enhanced prime editing systems by manipulating cellular determinants of editing outcomes." Cell 184.22(2021):5635-5652.]. In some embodiments, the PE2 prime editor protein may contain the amino acid sequence of SEQ ID NO: 32.
[0177] The amino acid sequence of sequence number 32 is as follows:
[0178] Prime Editing Version 3 (PE3) The PE3 prime editing system refers to a prime editing version developed to enhance prime editing efficiency by nicking the strand that has not been edited (i.e., the strand bound to the spacer of the pegRNA) using a second-strand nicking guide RNA. The second-strand guide RNA can be designed in the form of a conventional gRNA (such as sgRNA) to nick the strand that has not been edited in the vicinity of the editing site or target site. In some embodiments, PE3 may include the use of a separate Cas9 nickase in addition to the prime editing protein.
[0179] PE3b refers to PE3, but the second-strand nicking guide RNA is designed for temporal control such that the second-strand nick is not introduced until the desired edit is installed. This is achieved by designing a gRNA with a spacer sequence that matches only the edited strand, not the original allele. PE3 and PE3b are described in detail in [[Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785 (2019): 149-157.]]
[0180] Prime Editing Version 4 (PE4) Prime Editing Version 4 (PE4) includes the use of the same mechanism as PE2, but further includes the use of a plasmid encoding dominant negative MLH1 (dnMLH1) or the use of dnMLH1. For example, PE4 may be recognizable as including the use of the following elements: The PE2 prime editing protein; pegRNA; and Dominant-negative MLH1 (dnMLH1). [Chen, Peter J., et al. "Enhanced prime editing systems by manipulating cellular determinants of editing outcomes." Cell 184.22(2021):5635-5652.] describes how dominant-negative MLH1 can knock out endogenous MLH1 through inhibition, thereby reducing the measles-mumps-rubella (MMR) response in cells and potentially increasing prime editing efficiency.
[0181] Prime Edit Version 5 (PE5) Prime Edit Version 5 (PE5) includes the use of the same mechanism as PE3, but further includes the use of a plasmid encoding dominant-negative MLH1 (dnMLH1) or the use of dnMLH1 itself.
[0182] PE5 is described in detail in [Chen, Peter J., et al. "Enhanced prime editing systems by manipulating cellular determinants of editing outcomes." Cell 184.22(2021):5635-5652.].
[0183] PEmax PEmax is an improved prime-editing version developed to enhance editing efficiency. The PEmax prime-editor protein contains SpCas9 variants and MMLV RT variants. Specifically, the PEmax prime-editor protein has the following structure: [bpNLS(SV40)]-[SpCas9 R221K N394K H840A]-[SGGSX2-bpNLS(SV40)-SGGSX2]-[MMLV RT quintuplet variant (codon optimized)]-[bpNLS(SV40)]-[NLS(c-Myc)].
[0184] In this case, bpNLS(SV40) refers to the binocular SV40 NLS. MMLV RT quintuple mutant (codon optimized) refers to a human codon-optimized MMLV RT variant that includes the amino acid variations D200N, L603W, T330P, T306K, and W313F compared to wild-type MMLV RT. "SpCas9 R221K N394K H840A" refers to an SpCas9 variant that includes the amino acid variations R221K, N394K, and H840A compared to wild-type SpCas9. NLS(c-Myc) refers to c-Myc NLS. PEmax is described in detail in [Chen, Peter J., et al. "Enhanced prime editing systems by manipulating cellular determinants of editing outcomes." Cell 184.22(2021):5635-5652.]. Furthermore, the above-mentioned literature discloses various versions of prime editor proteins, including PE2* prime editor protein, CMP-PE-V1 prime editor protein, and CMP-PEmax prime editor protein, all of which can be used in the off-target prediction system in prime editing provided herein.
[0185] Nuclease-based Prime Edit Nuclease-based prime editing is a version of prime editing that uses a Cas protein with nuclease activity (i.e., DSB activity) instead of Cas9(H840A) nickase (such as wild-type SpCas9 or a non-nickase SpCas9 variant). The prime editor protein for nuclease-based prime editing is sometimes called a PE nuclease. Unlike PE3, which is designed to cause a nick on the strand that binds to the spacer of pegRNA, the use of two types of gRNA is not mandatory. Through a prime editing protein containing one type of pegRNA and a Cas nuclease (non-nickase), a DSB is made at the desired site, and thus editing is induced. Nuclease-based prime editing is described in detail in [Adikusuma, Fatwa, et al. "Optimized nickase-and nuclease-based prime editing in human and mouse cells." Nucleic acids research 49.18(2021):10785-10795.], the entire content of which is incorporated herein by reference. An example of a PE nuclease is PE2-nuclease. PE2-nuclease has the following structure: [bpNLS(SV40)]-[SpCas9(WT)]-[SGGSx2-XTEN16-SGGSx2]-[MMLV RT]-[bpNLS(SV40)].
[0186] In some embodiments, the PE2-nuclease may contain the amino acid sequence of SEQ ID NO: 33.
[0187] The amino acid sequence of sequence number 33 is as follows:
[0188] PEmax-nuclease PEmax-nuclease is a nuclease-based prime editor protein developed based on PEmax prime editor protein (i.e., a type of PE-nuclease), and is a prime editor protein containing a Cas protein that has nuclease activity other than niccasing activity (i.e., DSB activity). PEmax-nuclease has the following structure: [bpNLS(SV40)]-[SpCas9 R221K N394K]-[SGGSX2-bpNLS(SV40)-SGGSX2]-[MMLV RT quintuplet variant (codon optimized)]-[bpNLS(SV40)]-[NLS(c-Myc)].
[0189] In some embodiments, the PEmax nuclease may contain the amino acid sequence of SEQ ID NO: 34.
[0190] The amino acid sequence of sequence number 34 is as follows:
[0191] Use of epegRNA As described above, epegRNA is an improved pegRNA version. The pegRNA used in the above prime editing system may be epegRNA or other pegRNAs, but is not particularly limited.
[0192] Genome editing process using the Prime Editing System For the convenience of engineers in this field, the process of editing a cell's genome using a prime editing system will be explained using PE2 as an example. An example of the process of editing a cell's genome using a prime editing system within a cell is as follows: The PE2 prime editor protein and pegRNA form a complex, and then the complex is brought into contact with the cell's genome. The spacer of the pegRNA binds to the sequence of the corresponding target region. A nick is created on the strand of genomic DNA that is not bound to the spacer. The nick is created between the 3rd and 4th nucleotides located upstream of the 5' end of the PAM sequence. The sequence located upstream of the nick site functions as a primer that forms a complementary binding with the primer binding site of the pegRNA, thereby exposing the 3' end of the cleaved strand to the reverse transcription process. Based on the primer that forms a complementary binding with the primer binding site, the reverse transcriptase performs the reverse transcription process to form a 3' DNA flap. The reverse transcription template in this reverse transcription process is the RT template of the pegRNA. Through cell-specific mechanisms such as 5' flap removal, 3' flap ligation, and DNA mismatch repair, information about the 3' flap is placed within the genomic DNA. As a result of prime editing, information about the pegRNA RT template is delivered to desired sites on both strands of genomic DNA. The RT template contains the desired editing template (i.e., the editing template), and the information contained in the editing template is ultimately delivered to the target site in the genomic DNA.
[0193] The following describes in detail a method for predicting or confirming off-target effects in prime editing provided herein, a method developed to target prime editing so as to be widely available or applicable in confirming the confirmation of off-target effects that may occur in prime editing described above or to be developed in the future. The following method for predicting or confirming off-target effects in prime editing may use, but is not limited to, the prime editor proteins used in many of the prime editing versions described above. Furthermore, additional elements used in the prime editing versions described above may also be used in the method for predicting or confirming off-target effects in prime editing of this application. It will be apparent to those skilled in the art that prime editor proteins, pegRNAs, and / or prime editing systems developed based on the technical features of prime editing characterized by the use of Cas proteins and polymerases may be used in the method for predicting off-target effects of this application.
[0194] Off-target prediction method provided in this application Off-target In the field of DNA editing (such as gene editing or genome editing), off-target refers to gene modifications that occur at unintended sites. Gene modifications caused by off-target effects can be nonspecific. Genome editing tools developed to date include conventional CRISPR / Cas systems, base editing systems, prime editing systems, transcription activator-like effector nucleases (TALENs), meganucleases, and zinc finger nucleases. Such genome editing tools or systems may be designed so that editing of the target region can occur through specific mechanisms that enable binding to a predetermined sequence (such as the sequence of the target region). For example, in the CRISPR / Cas gene editing system, a guide RNA (gRNA) directs the Cas / gRNA complex to move to the intended target site. PAM sequences in the genome may also be involved in the movement to the target site. However, the Cas / gRNA complex is still highly likely to bind to sequences at unintended sites rather than the sequence of the target region. As described above, when a Cas / gRNA complex binds to a sequence at an unintended site, creating a nick or DSB there, unintended gene modification occurs. Off-target effects induce unintended gene modifications such as unintended point mutations, deletions, insertions, inversions, and translocations. Similarly, even in the process of editing DNA molecules (such as genomic DNA) by prime editing, off-target problems exist despite the fact that at least the PAM sequence and the pegRNA spacer sequence are involved in targeting.
[0195] The binding of genome editing tools to off-target regions is known to occur from a partial but sufficient match with target sequences in the off-target regions. Regarding the off-target binding mechanism, reference can be made to the known literature [Lin, Yanni, et al. "CRISPR / Cas9 systems have off-target activity with insertions or deletions between target DNA and guide RNA sequences." Nucleic acids research 42.11 (2014): 7473-7485.].
[0196] The off-target binding mechanism is described as being grouped into two main forms: base mismatch tolerance and bulge mismatch. For example, the off-target region can include, but is not limited to, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches with the guide RNA sequence. For example, the off-target region can include, but is not limited to, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches with the sequence of the target site corresponding to the sequence of each region of the pegRNA. In other words, in prime editing, there may be mismatches in the off-target region in one or more of the spacer region of the pegRNA, the PBS of the pegRNA, the DNA synthesis template (such as the homology arm) of the pegRNA, and the region corresponding to the PAM sequence.
[0197] Off-target issues can cause disruption of critical coding regions, potentially leading to serious problems such as cancer. Furthermore, off-target issues can lead to confusion of variables in biological research, and are even more likely to result in irreproducible outcomes (see [Eid, Ayman, and Magdy M. Mahfouz. "Genome editing: the road of CRISPR / Cas9 from bench to clinic." Experimental & Molecular Medicine 48.10(2016):e265-e265.] (its entire content is incorporated herein by reference)).
[0198] As described above, such off-target issues still exist not only with CRISPR / Cas gene editing systems but also with base editing and prime editing systems developed based on CRISPR / Cas gene editing systems. In this specification, off-target may be used as a concept opposite to on-target and may refer to gene modification at unintended sites.
[0199] The need for off-target prediction methods suitable for prime editing Overview of the Need for Off-Target Prediction Methods Suitable for Prime Editing As described above, off-target reactions can cause severe side effects in various ways (including irreversible and / or undetectable side effects). Accordingly, identifying potential off-target reactions when using systems for editing DNA molecules (such as genome editing systems) is extremely important in the research and development of therapeutic drugs. In addition, identifying genuine off-target reactions that occur with designed editing systems (such as CRISPR / Cas systems or prime editing systems) is costly and time-consuming. For this reason, research and development is underway on various methods that can identify off-target candidates, i.e., predict off-target reactions. However, existing methods developed prior to the filing date of this application for predicting potential off-target reactions in gene editing processes (such as genome editing processes using genome editing systems) were developed targeting conventional CRISPR / Cas systems or base editing. Off-target prediction methods targeting prime editing, i.e., off-target prediction methods developed targeting genome editing by prime editing, have yet to be developed.Despite its unique editing mechanism, which differs from conventional CRISPR / Cas systems, prime editing still relies on off-target prediction systems developed for conventional CRISPR / Cas systems to predict potential off-target events during the prime editing DNA process. ([Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.; Kim, Do Yon, et al. "Unbiased investigation of specificities of prime editing systems in human cells." Nucleic acids research 48.18(2020):10576-10589.; Bae, Sangsu, Jeongbin Park, and Jin-Soo Kim. "Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases." Bioinformatics) See 30.10(2014):1473-1475.; and Jin, Shuai, et al. "Genome-wide specificity of prime editors in plants." Nature Biotechnology 39.10(2021):1292-1299.] (The full contents of each of these are incorporated herein by reference). Applying such conventional off-target prediction systems to prime editors has many drawbacks. The need for off-target prediction methods suitable for prime editors will be explained in more detail below.
[0200] Off-target prediction system used for predicting off-target effects in genome editing using conventional CRISPR / Cas systems. As described above, various methods have been developed to predict off-target effects in genome editing using the CRISPR / Cas system. Conventional off-target prediction and / or confirmation methods (e.g., systems, platforms, etc.) can be classified into three categories based on their mechanism of action (MOA): cell-based off-target prediction systems, in vitro off-target prediction systems, and in silico off-target prediction systems. Examples of prediction systems included in each category are as follows.
[0201] - Cell-based off-target prediction systems: GUIDE-seq, GUIDE-tag, BLISS, BLESS, DISCOVER-seq, DNA cleavage capture via integrase-deficient lentiviral vectors, HTGTS, CReVIS-seq, ITR-seq, TAG-seq, INDUCE-seq, etc.
[0202] - In vitro off-target prediction systems: Digenome-seq, DIG-seq, CHANGE-seq, CIRCLE-seq, SITE-seq, etc.
[0203] - In silico off-target prediction systems: Cas-OFFinder, CRISPOR, CHOPCHOP, etc.
[0204] Each of the off-target prediction systems described above has different advantages and disadvantages. Typically, two or three systems are used in combination to predict genome-wide off-target activity during CRISPR-based genome editing.
[0205] Application of CRISPR / Cas-based off-target prediction systems to base editing. The system described above was expected to be used to predict off-target activity of base editor systems such as cytidine base editors and adenine base editors developed using Cas proteins. However, the system developed to predict off-target activity that may occur in genome editing using CRISPR / Cas systems was not suitable for application to base editing systems with different operating mechanisms. A more suitable off-target prediction system for base editing was needed, and accordingly, more appropriate and sophisticated off-target activity prediction systems or methods, such as One-seq (cell-based), CBE Digenome-seq (in vitro), and ABE Digenome-seq (in vitro), were developed.
[0206] Traditional off-target prediction methods used in prime editing The initial step in the genome editing mechanism by prime editing, such as prime editing version 2 (PE2), is Cas9-induced nicking of the spacer-unbound strand. Therefore, off-target activity in PE2 was predicted to be similar to that of Cas9 or Cas9 nickase (nCas9). Accordingly, we attempted to predict off-target activity in prime editing by using systems for predicting and / or confirming off-target activity in CRISPR / Cas genome editing, such as GUIDE-seq, nDignome-seq, and CAS-OFFinder (in silico). However, experiments disclosed herein have confirmed that using such methods to predict and / or confirm off-target activity in conventional CRISPR / Cas genome editing is not suitable for predicting off-target activity in prime editing.
[0207] Demand for off-target prediction systems suitable for prime editing Genome editing using prime editor proteins and pegRNAs is performed based on a different mechanism than that of conventional CRISPR / Cas system genome editing. Additionally, unlike conventional CRISPR / Cas systems, prime editing involves numerous factors (e.g., primer binding sites, reverse transcription templates, reverse transcriptases, etc.) in addition to the guide sequence, and is carried out through a process involving numerous enzymes (flap endonucleases, exonucleases, ligases, etc.). Although prime editing was developed based on the conventional CRISPR / Cas system, the genome editing mechanism of prime editing differs in various ways from that of conventional CRISPR / Cas systems. Therefore, conventional off-target prediction methods developed for genome editing using conventional CRISPR / Cas systems are not suitable for off-target prediction in prime editing. Furthermore, because it involves multiple processes involving numerous factors as described above, developing in vitro-based off-target analysis methods that can faithfully mimic such complex intracellular processes is difficult. For this reason, conventional off-target prediction methods are not applicable to prime editing, or are expected to produce inaccurate results.
[0208] In fact, the inventors of this application have experimentally confirmed that not only mismatches in the spacer region of pegRNA, but also primer binding sites, homology arms, and / or editing templates affect off-target effects in prime editing (see the sections "Editing Patterns at Validated Off-Target Sites" and "Region-Specific Mismatch Analysis" in the experimental examples of this application).
[0209] No methods have yet been reported for predicting off-target activity developed to target prime editing, taking into account the mechanism of prime editing. In other words, there is still no reliable off-target prediction method to identify off-target candidates in prime editing.
[0210] Overview of the off-target prediction method provided in this application This application provides a novel off-target prediction method suitable for prime editing. The inventors of this application have confirmed that the application of conventional off-target prediction systems targeting CRISPR / Cas systems results in inaccurate predictions (false positives and / or false negatives). Therefore, the inventors of this application have developed a novel method or system for predicting off-target outcomes in prime editing. Focusing on the ability or effect of prime editing to insert (place or write) a desired sequence at a desired site, the inventors of this application have developed a novel system or method for off-target prediction suitable for prime editing using a novel prime editing guide RNA (pegRNA) including a tag template for tag insertion. Furthermore, the inventors of this application have confirmed that the predictive reliability and / or accuracy of the newly developed off-target prediction system for prime editing is superior to that of conventional off-target prediction systems developed targeting conventional CRISPR / Cas genome editing systems.
[0211] The off-target prediction system provided in this application, which is developed to target prime editing (i.e., developed to be suitable for prime editing), is sometimes referred to as tagmentation for prime editor sequencing (TAPE-seq). Furthermore, a novel pegRNA used in TAPE-seq, which contains a tag template for placing a tag on the genome, is sometimes referred to as tagmentation pegRNA (tpegRNA).
[0212] This application provides a method or system for predicting off-target events that may occur in the process of editing DNA molecules using a prime editing system. This application provides a method for predicting off-target events that may occur in the genome editing process using a prime editing system. The method for predicting off-target events may also be referred to as, for example, a method for identifying off-target candidates, a method for identifying off-target information, or a method for identifying candidate off-target sites. In addition, any description relating to a method or system for predicting off-target events that may occur in the process of editing DNA molecules (such as genomes), or for identifying off-target information, may be used without limitation. As used herein, the term “off-target” includes the concept of an off-target region. For example, an off-target region or site may be described as an off-target. As used herein, “off-target prediction” may mean identifying off-target candidates. As used herein, “off-target prediction” may mean identifying off-target candidate regions. As used herein, the descriptions of “off-target,” “off-target prediction,” and “off-target candidates” should not be interpreted restrictively. In other words, methods for predicting off-target effects in prime editing may be described as any of the following, but are not limited to these, and can be used interchangeably as long as they relate to the prediction or confirmation of off-target effects that may occur in prime editing: prediction of off-target effects that may occur in prime editing; confirmation (or screening) of off-target candidates in (or that may occur in prime editing); confirmation (or screening) of off-target effects in (or that may occur in prime editing); confirmation of information regarding off-target effects in (or that may occur in prime editing); confirmation of areas where off-target effects may occur; confirmation of off-target sites; etc.
[0213] Regarding off-target prediction, the terms false positive and / or false negative may be used. Identifying regions other than true off-targets as off-target candidates may be expressed as a false positive result. A high false positive rate may be associated with a low validation rate. In this case, true off-target refers to validated off-targets, and is used to refer to off-targets that actually occur, in addition to off-target candidates found by the prediction system. For example, off-targets that occur when editing a cell's genome with a prime editing system are sometimes called true off-targets. In contrast, regions associated with off-targets found by an off-target prediction system may be called "off-target candidates," "predicted off-targets," etc., and can be distinguished from true off-targets. Off-target candidates found by an off-target prediction system may or may not be true off-targets. For example, true off-targets can be found by validating each off-target candidate. It is important that off-target prediction systems exhibit a low false positive rate, because if many off-target candidates are derived from the off-target prediction system, it becomes difficult to find true off-targets.
[0214] In another embodiment, the group of off-target candidates identified by the off-target prediction system may not include all true off-targets. This case is associated with a higher rate of missed detections. For example, if true off-target regions are not identified as off-target candidates, the rate of missed detections will be higher.
[0215] As described above, the system for predicting off-target events that occur during the DNA molecule editing process in prime editing according to this application is characterized by tagmentation based on the prime editing mechanism using tpegRNA. Below, we will describe in detail the tools for predicting off-target events in this application (prime editor protein and tpegRNA, etc.).
[0216] A tool for predicting off-target effects in prime editing. An overview of tools for predicting off-target events in prime editing (elements used in TAPE-seq) The present invention's method for predicting off-target editing requires two elements: Prime Editor protein; and Tagmentation pegRNA (tpegRNA) containing tag templates.
[0217] The tool for predicting off-targets in prime editing according to the present invention may include at least a prime editor protein and tpegRNA.
[0218] The method for predicting off-target activity in this application may be referred to as TAPE-seq. TAPE-seq, a method designed based on the prime editing mechanism and developed to target prime editing for off-target prediction, may utilize the prime editing mechanism. Accordingly, the method for predicting off-target activity provided herein includes the use of prime editor proteins used in prime editing. In other words, the off-target prediction system of this application may utilize various prime editor proteins as described above. Prime editor proteins used in the off-target prediction system in prime editing of this application include Cas proteins and polymerases (such as reverse transcriptase). However, such description does not require the use of the same type of prime editor protein as the prime editor protein in the specific prime editing system targeted for off-target prediction (e.g., the specific prime system targeted by off-target prediction by TAPE-seq). The off-target prediction system of this application may use the same or different type of prime editor protein as the prime editor protein in the prime editing system targeted for off-target prediction.
[0219] Similarly, in the off-target prediction system of this application, the use of the same type of pegRNA as the specific prime editing system targeted for off-target prediction is not necessarily required. In the off-target prediction system of this application, a pegRNA-based tpegRNA of the same type as the pegRNA used in the specific prime editing system targeted for off-target prediction can be used. Alternatively, a tpegRNA based on a different type of pegRNA other than typical pegRNA (such as epegRNA) may be used.
[0220] For example, even if the first prime editing system designated for confirming off-target information through an off-target prediction system is a PE2 prime editing system, a prime editor protein with nuclease activity (such as PE2-nuclease or PEmax-nuclease) can be used in TAPE-seq performed to confirm off-target information in the first prime editing system. In another example, if the first prime editing system designated for confirming off-target information through an off-target prediction system is a PE2 prime editing system, a PE2 prime editor protein can be used in TAPE-seq. Similarly, even if the first prime editing system targeted for off-target prediction is a PE2 prime editing system, an engineered tpegRNA (etpegRNA) can be used in TAPE-seq. In yet another example, if the first prime editing system targeted for off-target prediction is a PE2 prime editing system, a tagmentation pegRNA (tpegRNA) other than an engineered tpegRNA (etpegRNA) can be used.
[0221] Prime Editor Protein The off-target prediction system for prime editing described herein includes the use of a prime editor protein. The prime editor protein includes a Cas protein and a polymerase (such as reverse transcriptase). The prime editor protein is described in detail in the “Prime Editing System” section of this application. Examples of prime editor proteins that may be used in the off-target prediction system of this application include, but are not limited to, the prime editor proteins described above. It will be understood by those skilled in the art that fusion proteins or complexes for prime editing (or inventions that inherit the inventive concept of prime editing) developed for prime editing after the filing date of this application may also be used in the off-target prediction system of this application.
[0222] Similarly, examples of tpegRNAs that can be used in the off-target prediction system of this application include, but are not limited to, various embodiments of tpegRNAs developed based on the above-mentioned pegRNAs. It will be understood by those skilled in the art that tpegRNAs based on pegRNAs for prime editing developed for prime editing after the filing date of this application (or inventions that inherit the inventive concept of prime editing) can also be used in the off-target prediction system of this application.
[0223] In one embodiment, the prime editor protein used in the off-target prediction system in prime editing of the present invention may include Cas proteins and polymerases. In one embodiment, the Cas proteins include Cas12a, Cas12b1 (C2c1), Cas12c (C2c3), Cas12e (CasX), Cas12d (CasY), Cas12g, Cas12h, Cas12i, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Cs The Cas protein may be, but is not limited to, m4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cas13a (C2c2), Cas13b, Cas13c, Cas13d, Cas14, xCas9, cyclic displacement Cas9, Argonaut (Ago) domain, or fragments, homologs, or variants thereof. In certain embodiments, the Cas protein may have nickase activity. In certain embodiments, the Cas protein may be nCas9. In certain embodiments, the Cas protein may be SpCas9 nickase. In certain embodiments, the Cas protein may have nuclease activity. In certain embodiments, the Cas protein may be a Cas protein having nuclease activity. In certain embodiments, the Cas protein may be a Cas9 variant having nuclease activity. In certain embodiments, the Cas protein may be SpCas9 or a variant thereof. For example, a SpCas9 variant may have a form in which one or more amino acid residues selected from D10, R221, L244, N394, H840, K1211, and L1245 of wild-type SpCas9 are substituted with other amino acid residues.In certain embodiments, the Cas protein may include an amino acid sequence containing the H840A variation of the amino acid sequence of wild-type SpCas9 (SEQ ID NO: 28). In certain embodiments, the Cas protein may include an amino acid sequence containing the R221K and N394K amino acid variations in the amino acid sequence of wild-type SpCas9 (SEQ ID NO: 28). In certain embodiments, the Cas protein may include an amino acid sequence containing the R221K and N394K amino acid variations in the amino acid sequence of wild-type SpCas9 (SEQ ID NO: 28). In certain embodiments, the Cas protein may include the amino acid sequence of SEQ ID NO: 29, 30, or 31.
[0224] In certain embodiments, the polymerase may be a reverse transcriptase. In certain embodiments, the reverse transcriptase may be wild-type M-MLV reverse transcriptase. In certain embodiments, the reverse transcriptase may be a wild-type M-MLV reverse transcriptase variant. In certain embodiments, the wild-type M-MLV reverse transcriptase variant may include an amino acid sequence in which one or more amino acid variations selected from D200N, T306K, W313F, T330P, and L603W are included in the amino acid sequence of wild-type M-MLV reverse transcriptase (SEQ ID NO: 26). In certain embodiments, the wild-type M-MLV reverse transcriptase variant may include the amino acid variations D200N, T306K, W313F, T330P, and L603W based on the amino acid sequence of wild-type M-MLV reverse transcriptase SEQ ID NO: 26. In certain embodiments, wild-type M-MLV reverse transcriptase variants may include the amino acid variations D200N, T306K, W313F, and T330P, based on the amino acid sequence of SEQ ID NO: 26 of wild-type M-MLV reverse transcriptase. In certain embodiments, the reverse transcriptase may include the amino acid sequence of SEQ ID NO: 26 or 27.
[0225] As described above, prime editor proteins may further contain additional elements such as one or more linkers and / or one or more NLSs.
[0226] Examples of prime editor proteins that can be used in the off-target prediction system of this application include prime editor proteins in the above-mentioned prime-edited versions (e.g., PE1-PE5, PEmax, nuclease-based prime editing, PEmax-nuclease, etc.). In some embodiments, the prime editor protein may be the PE2 prime editor protein, PE2-nuclease, PEmax prime editor protein, or PEmax-nuclease. In certain embodiments, the prime editor protein may be the PEmax-nuclease.
[0227] Tagmentation pegRNA (tpegRNA) Overview of tpegRNA Tagmentation pegRNA (tpegRNA) is a guide nucleic acid developed from pegRNA designed to insert a tag sequence into a DNA molecule and is used in the off-target prediction method (i.e., an off-target prediction method in prime editing) provided herein. tpegRNA developed from pegRNA is sometimes referred to as a type of pegRNA. The tpegRNA provided herein contains a tag template and can be used to deliver information contained in the tag template (such as a tag sequence) to a DNA molecule (such as a genome) based on a prime editing mechanism.
[0228] In some embodiments, tpegRNA may be a single-stranded nucleic acid molecule (such as single-stranded RNA). In some embodiments, tpegRNA may be a nucleic acid complex composed of two or more strands (such as a complex of single-stranded RNA and double-stranded RNA). When tpegRNA is formed to contain two strands, portions of the sequences of the two strands may form complementary bonds with the gRNA core region, thus forming double-stranded tpegRNA. In certain embodiments, tpegRNA may be a single-stranded RNA molecule.
[0229] Several embodiments of this application provide tpegRNA. The elements contained in tpegRNA will be disclosed below.
[0230] tpegRNA includes a spacer, a gRNA core, and an elongation region. As described above, the pegRNA used for prime editing has a morphology in which an elongation arm is added to the 3' or 5' end of a conventional gRNA. Typically, pegRNA has a morphology in which an elongation arm is added to the 3' end of a conventional gRNA. Similarly, tpegRNA has a morphology in which an elongation arm is added to the 3' or 5' end of a conventional gRNA, and the elongation arm may include an elongation region.
[0231] In some embodiments, tpegRNA may have a form in which an elongation arm is attached to the 3' end of a conventional gRNA. In some embodiments, the spacer, gRNA core, and elongation region may be located in the 5'→3' direction in tpegRNA. In some embodiments, tpegRNA may further include, but is not limited to, one or more independently selected additional functional elements (e.g., linkers, transcription terminators, RNA protection motifs, etc.) in one or more regions selected from between the 5' end and the spacer, between the spacer and the gRNA core, between the gRNA core and the elongation region, and between the elongation region and the 3' end. In other words, in tpegRNA, such independently selected additional functional elements may or may not be present between each of the above elements, and are not particularly limited. In some embodiments, the elongation region of tpegRNA includes a tag template. In some embodiments, the tag template may be described separately from the DNA synthesis template (e.g., RT template). For example, the elongation region of tpegRNA may be described as including a primer-binding site (PBS), a tag template, and a DNA synthesis template. In this case, the tag template and the DNA synthesis template are described separately, meaning that the tag template is described separately from the DNA synthesis template of conventional pegRNA. In another embodiment, the tag template may be encoded in a DNA molecule edited by the reverse transcriptase of a prime editor protein and therefore described as one element of the DNA synthesis template. For example, the elongation region of tpegRNA may be described as including a primer-binding site and a DNA synthesis template (in this case, the DNA synthesis template includes the tag template). In the following description, the tag template and the DNA synthesis template will be described separately. Unless otherwise stated, tpegRNA will be recognized as including a tag template.
[0232] Furthermore, the tpegRNA elongation region may further include one or more independently selected additional functional regions in addition to PBS, tag templates, and DNA synthesis templates.
[0233] For example, the elongation region of tpegRNA may further include a 3' engineering region containing an RNA protection motif. When the elongation region of tpegRNA further includes a 3' engineering region containing an RNA protection motif, the tpegRNA is sometimes called engineered tpegRNA (etpegRNA). For example, the RNA protection motif may include the sequence CGCGGUUCUAUCUAGUUACGCGUUAAACCAACUAGAA (SEQ ID NO: 41). In some embodiments, the 3' engineering region may further include a linker for ligating the RNA protection motif, in addition to the RNA protection motif. In this case, the linker for ligating the RNA protection motif may function to ligate the RNA protection motif and PBS. In this specification, the term tpegRNA is used as a concept including embodiments of etpegRNA, and descriptions of tpegRNA should be interpreted as including content related to etpegRNA unless otherwise specified. Specific embodiments limited to the use of etpegRNA will be described in the context of etpegRNA.
[0234] In the initial stage, the 3' engineering area is 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, 50nt, 51nt, 52nt, 53nt, 54nt, 55nt, 56nt, 57nt, 58 The 3' engineering region may have lengths of nt, 59nt, 60nt, 61nt, 62nt, 63nt, 64nt, 65nt, 66nt, 67nt, 68nt, 69nt, 70nt, 71nt, 72nt, 73nt, 74nt, 75nt, 76nt, 77nt, 78nt, 79nt, 80nt, 81nt, 82nt, 83nt, 84nt, 85nt, 86nt, 87nt, 88nt, 89nt, 90nt, 91nt, 92nt, 93nt, 94nt, 95nt, 96nt, 97nt, 98nt, 99nt, or 100nt or greater, or it may have lengths within a range set by two values selected from the above values, but is not limited thereto. In a particular embodiment, the 3' engineering region may have lengths of 10 to 70nt. In a particular embodiment, the 3' engineering region may have lengths of 20 to 60nt.
[0235] In some embodiments, tpegRNA is approximately 30nt, 40nt, 50nt, 60nt, 70nt, 80nt, 90nt, 100nt, 110nt, 120nt, 130nt, 140nt, 150nt, 160nt, 170nt, 180nt, 190nt, 200nt, 210nt, 220nt, 230nt, 240nt, 250nt, 260nt, 270nt, 280nt, 290nt, 300nt, 310nt, 320nt, 330nt, 34 The tpegRNA may have lengths of 0nt, 350nt, 360nt, 370nt, 380nt, 390nt, 400nt, 410nt, 420nt, 430nt, 440nt, 450nt, 460nt, 470nt, 480nt, 490nt, 500nt, 520nt, 540nt, 560nt, 580nt, or 600nt or greater, or it may have lengths within a range set by two values selected from the above values, but is not limited to these. In certain embodiments, the tpegRNA may have lengths of 100-300nt or 100-400nt.
[0236] It should be noted that, unlike typical pegRNA (pegRNA without a tag template), the tpegRNA of this application includes a tag template for inserting a tag sequence into a DNA molecule. For the convenience of those skilled in the art, examples of conventional gRNA, pegRNA, and tpegRNA are disclosed in Figure 1. The examples of gRNA, pegRNA, and tpegRNA disclosed in Figure 1 are illustrated based on the essential elements contained in each guide RNA, and it will be apparent to those skilled in the art that additional elements may be included between or at the ends of each element.
[0237] Below, we will explain each element of tpegRNA in detail.
[0238] Conventional gRNA portion-spacer As described above, tpegRNA may include a spacer, a gRNA core, and an extension region. In this case, the spacer and gRNA core are elements derived from conventional gRNA. The spacer and gRNA core are adequately described in the "CRISPR / Cas System" and "Prime Editing System" sections herein. The spacer includes a spacer sequence. The spacer sequence may be designed according to a target sequence, without limitation. In this case, a region of the PAM sequence may be considered. The spacer sequence may be designed as a sequence complementary to the target sequence in the spacer-bound strand of genomic DNA. The spacer sequence may be the same sequence (or substantially the same sequence or corresponding sequence) as the target sequence in the spacer-unbound strand of genomic DNA. The spacer sequence may be an RNA sequence, a DNA sequence, or an RNA / DNA hybrid sequence. Typically, the spacer sequence is an RNA sequence. Similar to conventional gRNA, the spacer sequence is involved in directing the Cas protein (the Cas protein included in the Prime Editor) to the target region. In other words, the spacer sequence and the target sequence form a complementary bond, so the prime editor protein / tpegRNA complex is located in the target region, and the prime editor protein creates a nick or DSB in the target region.
[0239] In some embodiments, the spacer array may have lengths of approximately 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, or 50nt or greater, or may have lengths within a range of two values selected from the above values, but is not limited thereto. In certain embodiments, the spacer array may have lengths of 10 to 35nt. In certain embodiments, the spacer array may have a length of 13 to 30 nt. In certain embodiments, the spacer array may have a length of 15 to 25 nt.
[0240] Conventional gRNA portion - gRNA core As described above, tpegRNA may include a spacer, a gRNA core, and an elongation region. In this case, the spacer and gRNA core are elements derived from conventional gRNA. The gRNA core is the portion that interacts with the Cas protein and binds to the Cas protein to form a complex. The gRNA core is sometimes called the scaffold region. The gRNA core or scaffold may be designed differently depending on the type of Cas protein used, for example, depending on the type of microorganism from which the Cas protein originates and the type of CRISPR system.
[0241] In one embodiment, the gRNA core may include a scaffold sequence. The scaffold sequence may be an RNA sequence, a DNA sequence, or an RNA / DNA hybrid sequence. Some sequences of the gRNA core may interact with other sequences of the gRNA core, thus forming structures such as stems / loops or hairpins.
[0242] In some embodiments, the scaffold sequence is approximately 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 4 The scaffold array may have lengths of 9nt, 50nt, 55nt, 60nt, 65nt, 70nt, 75nt, 80nt, 85nt, 90nt, 95nt, 100nt, 110nt, 120nt, 130nt, 140nt, 150nt, 160nt, 170nt, 180nt, 190nt, 200nt, 210nt, 220nt, 230nt, 240nt, 250nt, 260nt, 270nt, 280nt, 290nt, or 300nt or greater, or it may have lengths within a range set by two values selected from the above values. In certain embodiments, the length of the scaffold array may be 30 to 200nt, but is not limited to these. In certain embodiments, the length of the scaffold array may be 50 to 150nt. In certain embodiments, the length of the scaffold array may be 60 to 100nt.
[0243] Overview of the tpegRNA elongation region As described above, tpegRNA may contain an elongation region (or elongation arm). The elongation region of tpegRNA is characterized by containing a tag template. The elongation region of pegRNA may be located at the 3' or 5' end of the conventional gRNA. For example, tpegRNA may have the following structure in the 5'→3' direction: "[conventional gRNA portion]-[elongation region]" or "[elongation region]-[conventional gRNA portion]". The [conventional gRNA portion] may include the spacer and scaffold (gRNA core) described above. Preferably, the elongation region is located at the 3' end of the conventional gRNA portion. For example, tpegRNA may include a spacer, a gRNA core, and an elongation region. In some embodiments, the spacer, gRNA core, and elongation region may be located in the 5'→3' direction in tpegRNA.
[0244] In some embodiments, the elongation region, spacer, and gRNA core may be located in the 5'→3' direction in the tpegRNA. In some embodiments, the elongation region of the tpegRNA may include an RNA sequence, a DNA sequence, or a DNA / RNA hybrid sequence. Preferably, the elongation region includes an RNA sequence, but is not limited to one.
[0245] The elongation region of tpegRNA is characterized by the inclusion of a tag template. In other words, the elongation region includes a primer binding site (PBS), a tag template, and a DNA synthesis template (such as an RT template). The elongation region may further include one or more independently selected additional elements (e.g., linkers, RNA protection motifs, etc.) between the above elements or at the terminal elements.
[0246] Possible additional elements In some embodiments, tpegRNA may include one or more independently selected additional elements in addition to the elongation region, gRNA core, and spacer. The additional elements may be, but are not limited to, any one of the following: a linker, a polyU tail, a polyA tail, and an RNA protection motif. For example, tpegRNA may include a U-rich, A-rich, or AU-rich sequence at its 3' end. In certain embodiments, tpegRNA may include a (U)n sequence at its 3' end, where n may be an integer from 3 to 20. In certain embodiments, tpegRNA may include a (U)7 sequence at its 3' end.
[0247] tpegRNA extension region (1) Overview of the tpegRNA elongation region (1) As described above, tpegRNA includes an elongation region. The elongation region may include the tag template and primer binding sites, which are described in detail with respect to the pegRNA.
[0248] In some embodiments, the elongation region of tpegRNA may be described as comprising a first region containing a DNA synthesis template, a second region containing a tag template, and a third region containing a primer binding site. In this case, part or all of the first region may be the DNA synthesis template. In this case, part or all of the second region may be the tag template. In this case, part or all of the third region may be the primer binding site. The elements included in the elongation region will be described in detail below.
[0249] Tag template The elongation region of tpegRNA may include a tag template. A tag template refers to a portion of the elongation region complementary to the tag sequence to be placed on the spacer-unbound strand of a DNA molecule or single-stranded DNA flap (such as a 3' DNA flap) synthesized by a polymerase such as reverse transcriptase. The tag template may also be complementary to the tag sequence to be placed on the spacer-unbound strand of the DNA molecule or on the DNA flap (such as a 3' DNA flap). The off-target prediction method of this application can achieve the objective of off-target prediction in prime editing by confirming information about the tag, including the tag sequence to be placed on the DNA molecule and a sequence complementary to the tag sequence (for example, information about the site where the tag sequence is inserted, whether or not the tag sequence or a sequence complementary to the tag sequence exists, the chromosome in which the tag sequence is inserted, etc.). Examples of tag sequences corresponding to tpegRNA tag templates can be found in [Tsai, Shengdar Q., et al. "GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases." Nature biotechnology 33.2(2015):187-197.], the entire contents of which are incorporated herein by reference.
[0250] The tag template for tpegRNA and the tag sequence to be inserted into DNA are not particularly limited and can be appropriately selected depending on the intended use of tpegRNA. For example, the tag template sequence may include sequences such as AUACCGUUAUUAACAUAUGACAACUCAAUUAAAC (SEQ ID NO: 42), GUUAUUAACAUAUGACAACUCAAUUAAAC (SEQ ID NO: 43), UAUGACAACUCAAUUAAAC (SEQ ID NO: 44), AUUAACAUAUGAC (SEQ ID NO: 45), GACAACUCA (SEQ ID NO: 46), or CUCAAUUA (SEQ ID NO: 47). For example, the tag sequence may include sequences such as GTTTAATTGAGTTGTCATATGTTAATAACGGTAT (SEQ ID NO: 48), GTTTAATTGAGTTGTCATATGTTAATAAC (SEQ ID NO: 49), or GTTTAATTGAGTTGTCATA (SEQ ID NO: 50).
[0251] In some embodiments, the tag template may be an RNA sequence, a DNA sequence, or an RNA / DNA hybrid sequence. Preferably, the tag template is an RNA sequence.
[0252] In some embodiments, the tag template may have a length of 1nt to 500nt. In some embodiments, the tag template may have lengths of 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, 50nt, 51nt, 52nt, 53nt, 54nt, 55nt, It may have lengths of 56nt, 57nt, 58nt, 59nt, 60nt, 61nt, 62nt, 63nt, 64nt, 65nt, 66nt, 67nt, 68nt, 69nt, 70nt, 71nt, 72nt, 73nt, 74nt, 75nt, 76nt, 77nt, 78nt, 79nt, 80nt, 81nt, 82nt, 83nt, 84nt, 85nt, 86nt, 87nt, 88nt, 89nt, 90nt, 91nt, 92nt, 93nt, 94nt, 95nt, 96nt, 97nt, 98nt, 99nt, or 100nt or greater, or it may have lengths within the range set by two values selected from the above values. In certain embodiments, the tag template may have a length of 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, or 50nt or greater, or a length within a range set by two values selected from the above values. In certain embodiments, the tag template may have a length of 10 to 70nt.In certain embodiments, the tag template may have a length of 10 to 50 nt. In certain embodiments, the tag template may have a length of 15 to 40 nt. In certain embodiments, the tag template may have a length of 25 to 40 nt. In certain embodiments, the tag template may have a length of 30 to 40 nt. In certain embodiments, the tag template may have a length of 19, 24, 29, or 34 nt.
[0253] The length of the tag template may be appropriately designed according to the purpose of the disclosure disclosed herein, which is to analyze off-target effects in prime editing through the tag sequence to be placed. If the tag template is too short, it may be difficult to obtain information about the tag sequence inserted into the DNA molecule. If the tag template is too long, the tag sequence may not be placed in the DNA molecule, and the objective of off-target prediction may not be achieved.
[0254] Furthermore, tag templates can be designed without limitation according to the intended use of tpegRNA. In the off-target prediction method of this application, the tag template serves as the basis for the tag sequence to be inserted into genomic DNA. In other words, the tag template is used as a reverse transcription template, and thus the tag sequence is placed in genomic DNA. Using the tag sequence placed in genomic DNA in this manner, or a sequence complementary to the tag sequence, tagging sites in genomic DNA can be identified. Using the tagging sites, regions where off-target effects may occur (candidate off-target regions or off-target candidates, etc.) can be identified. When designing the tag sequence or tag template for tpegRNA used in off-target prediction, it may be considered whether the same sequence exists in genomic DNA. This is because if the same sequence as the tag sequence or tag template sequence exists in genomic DNA, for example, it may affect the off-target prediction results. In another example, even if the same sequence exists, if the location of the same sequence is known in advance, the results for the corresponding location can be excluded from the off-target prediction results. As described above, the tag template sequence or tag sequence can be designed according to the purpose or plan for using tpegRNA.
[0255] Primer binding site (PBS) The elongation region of tpegRNA may contain a primer-binding site (PBS). The PBS of tpegRNA may play the same or similar role as the primer-binding site of pegRNA in prime editing. The polymerization activity of the polymerase (such as reverse transcriptase) of the prime editing protein may be located in the 5'→3' direction relative to the strand bound to the template strand. When a primer (such as a region in the spacer-unbound strand) anneals to the primer-binding site, the polymerase (such as reverse transcriptase) can polymerize a single strand of DNA using the template strand as a template. For example, when using a prime editing protein of Prime Editing Version 2, the primer-binding site (PBS) of tpegRNA binds to a primer sequence formed from the endogenous DNA strand at the nicking target site resulting from the prime editing protein, thereby exposing the 3' end of the strand to be nicked. The binding of the primer-binding site and primer sequence in the elongation region of tpegRNA provides a matrix that allows the reverse transcriptase to polymerize a single strand of DNA. The primer binding site may have a sequence complementary to the primer sequence located upstream (towards the 5' direction) of the cleavage site (caused by a nick or DSB) in the spacer-unbound strand. In some embodiments, the primer sequence may be part of a sequence located in the range of -0 to -200 relative to the cleavage site. In certain embodiments, the primer sequence may be part of a sequence located in the range of -0 to -50 relative to the cleavage site. In certain embodiments, the primer sequence may be part of a sequence located in the range of -0 to -30 relative to the cleavage site. In certain embodiments, the primer sequence may be part of a sequence located in the range of -0 to -20 relative to the cleavage site. In this case, - refers to the 5' direction, and the number, such as 30, refers to the number of nucleotides. For example, -30 refers to the 30th nucleotide located upstream from the cleavage site. However, 0 refers to the cleavage site.
[0256] In some embodiments, the primer binding site may be an RNA sequence, a DNA sequence, or an RNA / DNA hybrid sequence. Preferably, the primer binding site is an RNA sequence.
[0257] In some embodiments, the primer binding site or primer may have a length of 1 nt to 500 nt. In some embodiments, the primer binding site or primer may have lengths of 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, 31 nt, 32 nt, 33 nt, 34 nt, 35 nt, 36 nt, 37 nt, 38 nt, 39 nt, 40 nt, 41 nt, 42 nt, 43 nt, 44 nt, 45 nt, 46 nt, 47 nt, 48 nt, 49 nt, 50 nt, 51 nt, 52 nt, 53 nt, 54 nt, 55 nt, 56 The length may be nt, 57nt, 58nt, 59nt, 60nt, 61nt, 62nt, 63nt, 64nt, 65nt, 66nt, 67nt, 68nt, 69nt, 70nt, 71nt, 72nt, 73nt, 74nt, 75nt, 76nt, 77nt, 78nt, 79nt, 80nt, 81nt, 82nt, 83nt, 84nt, 85nt, 86nt, 87nt, 88nt, 89nt, 90nt, 91nt, 92nt, 93nt, 94nt, 95nt, 96nt, 97nt, 98nt, 99nt, or 100nt or greater, or it may be a length within the range set by two values selected from the above values, but is not limited to these. In certain embodiments, the primer binding site or primer may have lengths of 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, or 50nt or greater, or may have lengths within a range set by two values selected from the above values.In certain embodiments, the primer binding site or primer may have a length of 3 to 30 nt. In certain embodiments, the primer binding site or primer may have a length of 5 to 20 nt. In certain embodiments, the primer binding site or primer may have a length of 5 to 15 nt.
[0258] DNA synthesis template The elongation region of tpegRNA may contain a DNA synthesis template. The DNA synthesis template may also be a reverse transcription template (RT template). The DNA synthesis template of tpegRNA may perform the same or similar role as the DNA synthesis template of pegRNA. The DNA synthesis template of tpegRNA may optionally contain an editing template. A typical pegRNA used in prime editing necessarily contains an editing template because prime editing is intended to perform editing. On the other hand, the tpegRNA used in the off-target prediction system of this application is primarily intended for tag placement rather than editing, and therefore selectively contains an editing template. In other words, in some embodiments, the DNA synthesis template may or may not contain an editing template. Preferably, the DNA synthesis template contains an editing template, but is not limited thereto.
[0259] In some embodiments, the DNA synthesis template may be an RNA sequence, a DNA sequence, or a DNA / RNA hybrid sequence. Preferably, the DNA synthesis template (such as an RT template) is an RNA sequence.
[0260] In some embodiments, the DNA synthesis template sequence may correspond to a portion of the sequence located in the region of the spacer-unbound strand between +0 and +500 relative to the cleavage site (resulting from a nick or DSB). In this case, + refers to the 3' direction, and the number, such as 500, refers to the order of the nucleotides from the cleavage site. For example, 1 refers to the first nucleotide located downstream from the cleavage site. For example, 500 refers to the 500th nucleotide located downstream from the cleavage site. However, 0 refers to the cleavage site. In some embodiments, the DNA synthesis template sequence may correspond to a portion of the sequence in the region of the spacer-unbound strand between <+100, <+90, <+80, <+70, <+60, <+50, <+40, <+30, <+20, or <+10 relative to the cleavage site (resulting from a nick or DSB). For example, the sequences of DNA synthesis templates other than the editing template may be sequences complementary to a portion of the sequences in the <+100, <+90, <+80, <+70, <+60, <+50, <+40, <+30, <+20, or <+10 regions relative to the cleavage site of the spacer-unbound strand, and / or may be substantially the same as the above portion of the sequence of the spacer-unbound strand.
[0261] In some embodiments, the DNA synthesis template may have a length of 1 nt to 500 nt. In some embodiments, the DNA synthesis template may have lengths of 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46 nt, 47nt, 48nt, 49nt, 50nt, 51nt, 52nt, 53nt, 54nt, 55nt, 56nt, 57nt, 58nt, 59nt, 60nt, 61nt, 62nt , 63nt, 64nt, 65nt, 66nt, 67nt, 68nt, 69nt, 70nt, 71nt, 72nt, 73nt, 74nt, 75nt, 76nt, 77nt, 78nt, 79nt, 80nt, 81nt, 82nt, 83nt, 84nt, 85nt, 86nt, 87nt, 88nt, 89nt, 90nt, 91nt, 92nt, 93nt, 94nt, 95nt, 96nt, 97nt, 98nt, 99nt, 100nt, 110nt, 120nt, 130nt, 140nt, 150nt, 160nt, 170nt, 180nt, 190nt, or 200nt or greater, or may have a length within the range set by two values selected from the above values, but is not limited to these. In certain embodiments, the DNA synthesis template may have a length of 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, or 40nt. In certain embodiments, the DNA synthesis template may have a length of 3 to 40nt.In certain embodiments, the DNA synthesis template may have a length of 5 to 30 nt. In certain embodiments, the DNA synthesis template may have a length of 7 to 30 nt.
[0262] In some embodiments, the DNA synthesis template may include an editing template and homology regions (or homology arms). In some embodiments, the DNA synthesis template may include homology regions. The homology regions included in the DNA synthesis template will be described below.
[0263] The homology region is the region of the pegRNA used for prime editing that corresponds to the homology arm or homology region mentioned above.
[0264] In some embodiments, the homology region is complementary to a portion of the sequence of the spacer-unbound strand of the target DNA. In some embodiments, the homology region has a sequence homologous to a portion of the sequence of the spacer-bound strand of the target DNA.
[0265] The homology region sequence is complementary to a portion of the sequence of a region located downstream (towards the 3' direction) of the resulting cleavage site (caused by a DSB or nick) in the spacer-unbound strand of the DNA molecule. For example, in prime-edited version 2, the homology region may have a sequence complementary to a sequence located downstream of the site where the nick is formed in the spacer-unbound strand. From another perspective, in prime-edited version 2, the homology region may have a sequence homologous to a portion of the sequence located upstream of the region corresponding to the site where the nick is formed in the spacer-bound strand.
[0266] On the other hand, homology regions replace the sequence of the intrinsic strand of the DNA molecule, promoting the annealing of single-stranded DNA flaps (such as 3' DNA flaps), and thus facilitating the placement of edit and / or tag sequences into the DNA molecule. Homology regions can be described as being encoded by the polymerase (such as reverse transcriptase) of prime editing proteins, and thus as part of the DNA synthesis template.
[0267] In some embodiments, the homology region may include an RNA sequence, a DNA sequence, or a DNA / RNA hybrid sequence. Preferably, the homology region includes an RNA sequence.
[0268] In some embodiments, the homology region may have a length of 1nt to 500nt. In some embodiments, the homology region may be 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, 50nt, 51nt, 52nt, 53nt, 54nt, 55nt, 5 It may have lengths of 6nt, 57nt, 58nt, 59nt, 60nt, 61nt, 62nt, 63nt, 64nt, 65nt, 66nt, 67nt, 68nt, 69nt, 70nt, 71nt, 72nt, 73nt, 74nt, 75nt, 76nt, 77nt, 78nt, 79nt, 80nt, 81nt, 82nt, 83nt, 84nt, 85nt, 86nt, 87nt, 88nt, 89nt, 90nt, 91nt, 92nt, 93nt, 94nt, 95nt, 96nt, 97nt, 98nt, 99nt, or 100nt or greater, or it may have lengths within the range set by two values selected from the above values. In certain embodiments, the homology region may have a length of 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, or 40nt. In certain embodiments, the homology region may have a length of 3 to 40nt. In certain embodiments, the homology region may have a length of 5 to 30nt. In certain embodiments, the homology region may have a length of 7 to 30nt.
[0269] In some embodiments, the DNA synthesis template may include an editing template. The editing template refers to a portion of the extension region that encodes the edit to be placed on a spacer-unbound strand or single-stranded DNA flap (such as a 3' DNA flap) synthesized by a polymerase (such as reverse transcriptase).
[0270] The editing template may be complementary to the edit to be placed on the spacer-unbound strand of a DNA molecule or DNA flap (such as a 3' DNA flap). For example, as a result of prime editing, the edit to be placed on the spacer-unbound strand is located downstream of the resulting cleavage site.
[0271] In some embodiments, the RT template may include an editing template, homology regions, etc. In this case, the RT template is equivalent to the DNA synthesis template. However, the RT template as used herein is based on the use of a prime editing protein having a polymerase, i.e., reverse transcriptase, while the DNA synthesis template is more broadly based on the use of a prime editing protein having any polymerase.
[0272] For example, the tpegRNA editing template may have the same sequence as the editing template corresponding to the desired edit for encoding the desired edit in a DNA molecule (in this case, the desired edit may be a desired edit pre-designed in a prime edit subject to off-target analysis through the off-target prediction system of the present invention). For example, the tpegRNA editing template may have a sequence complementary to the sequence of the desired edit to be placed in a DNA molecule (such as a genome) or DNA flap (such as a 3' DNA flap). In another example, the tpegRNA editing template may have a different sequence from the editing template corresponding to the desired edit to be encoded in a DNA molecule. In yet another example, the tpegRNA editing template may have a sequence different from some or all of the sequences complementary to the sequence of the desired edit to be placed in a DNA molecule (such as a genome) or DNA flap (such as a 3' DNA flap). In some embodiments, two types of tpegRNA may be used for off-target prediction in a prime edit, in which case the editing template sequences contained in each tpegRNA may differ from some or all of the editing template sequences of the desired edit.
[0273] In some embodiments, one type of tpegRNA may be used for off-target prediction in prime editing, in which case the sequence of the editing template contained in the tpegRNA may have the same sequence as the editing template corresponding to the desired edit. In some embodiments, one type of tpegRNA may be used for TAPE-seq, in which case the sequence of the editing template contained in the tpegRNA may have a sequence that is different from some or all of the editing template corresponding to the desired edit.
[0274] As described above, prime editing technology is a system designed to insert a desired sequence into a desired location (i.e., a system designed to "write" a desired sequence), and the editing is not particularly limited. For example, the length of the edit may be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 52, 54, 56, 58, or 60 nt (or bp) or greater.
[0275] In some embodiments, the edits to be placed on the DNA molecule to be edited, compared to the original sequence located in the region corresponding to the edit on the DNA molecule to be edited (i.e., the sequence before editing), may include the insertion of one or more nucleotides, the deletion of one or more nucleotides, the substitution of one or more nucleotides with other nucleotides, or a combination thereof. Furthermore, the edits to be placed on the DNA molecule to be edited may have a region designed to insert the same sequence as a portion of the sequence of the endogenous DNA strand to be replaced. For example, editing may involve the insertion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotides to be inserted may be located consecutively or not in the nucleic acid. For example, editing may include the deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotide to be deleted may be located consecutively or not in the nucleic acid. For example, editing may include substitutions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 or more nucleotides, in which case the nucleotides to be substituted may or may not be located consecutively in the nucleic acid. In another example, editing may include the above insertions and substitutions. In another example, editing may include the above deletions and substitutions. In another example, editing may include the above insertions and deletions. In another example, editing may include the above insertions, deletions, and substitutions.One or more of the above insertions, deletions, and substitutions may occur in the region of the "edited DNA molecule" corresponding to the site where the edit is to be performed.
[0276] Furthermore, the edits to be placed on the target DNA molecule may include a region designed to insert the same sequence as a portion of the endogenous DNA strand to be replaced. This region, present in the editing template that encodes the edit, is sometimes called the "homologous region of the editing template." An editing template may contain one or more homologous regions. In other words, an editing template may contain one or more homologous regions.
[0277] For the convenience of engineers in the relevant technical field, examples of possible structures for an editing template will be provided. Editing templates can be designed for any purpose without limitation, and therefore, their possible embodiments should not be interpreted as being limited to the following examples. For example, an editing template may have the following structure: [First homology region of the editing template]-[Nucleotide for G to T substitution]-[Second homology region of the editing template]-[Nucleotide for A to T substitution]-[Third homology region of the editing template]. In another example, an editing template may have the following structure: [First homology region of the editing template]-[Nucleotide for A to C substitution]-[Second homology region of the editing template]. In yet another example, an editing template may have the following structure: [First homology region of the editing template]-[Nucleotide for TAA insertion]. In yet another example, an editing template may have the following structure: [First homology region of the editing template]-[Nucleotide for TGG insertion]-[Second homology region of the editing template]-[Nucleotide for A to G substitution]. In a further example, an editing template may have the following structure: [nucleotide for AGG insertion]-[first homology region of the editing template].
[0278] In some embodiments, the site where editing occurs may be within a range of +0 to +100 relative to the cleavage site of the spacer unbound chain. In certain embodiments, the site where editing occurs may be within a range of +0 to +60. In certain embodiments, the site where editing occurs may be within a range of +1 to +30. In certain embodiments, the site where editing occurs may be within a range of +0 to +20. In certain embodiments, the site where editing occurs may be within a range of +0 to +10. In some embodiments, when inserting tags, the site where editing occurs may be located downstream of the placed tag sequence. For example, editing may occur within a range of +10 to +50 relative to the cleavage site.
[0279] In some embodiments, the editing template may consist of RNA. In certain embodiments, the editing template may consist of DNA. In certain embodiments, the editing template may consist of an RNA / DNA hybrid. In certain embodiments, the editing template may consist of RNA.
[0280] In some embodiments, the editing template may have a length of 1nt to 200nt. In some embodiments, the editing template may have a length of 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, 36nt, 37nt, 38nt, 39nt, 40nt, 41nt, 42nt, 43nt, 44nt, 45nt, 46nt, 47nt, 48nt, 49nt, or 50nt or greater. In a particular embodiment, the editing template may have a length of 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, or 20nt or greater.
[0281] Relative positional relationship of the elongation region The following describes the relative positions of each of the above elements in the extension region. The tag template, PBS, and DNA synthesis may be directly linked to other elements (for example, through covalent bonds), or they may be linked via additional elements such as linkers.
[0282] In the case of a 3' elongated region (i.e., a tpegRNA having an elongated region attached to the 3' end of a conventional gRNA), the following positional relationships may be observed.
[0283] In some embodiments, these elements may be positioned in the 5'→3' direction of the tpegRNA in the extension region in the order of DNA synthesis template, tag template, and primer binding site. In this case, in the 5'→3' explanation of the cleavage site resulting from the spacer-unbound strand in the DNA molecule (such as genomic DNA), the tag sequence delivered by the tag template may be located at the first nucleotide, and the sequence delivered by the DNA synthesis template may be located at the second nucleotide. In other words, the positional relationship between the tag sequence placed on the DNA molecule and the sequence complementary to the DNA synthesis template in the spacer-unbound strand may have the following structure: v-[tag sequence]-[sequence complementary to the DNA synthesis template]. In this case, v refers to the resulting cleavage site.
[0284] In other embodiments, these elements may be located in the 5'→3' extension region of tpegRNA in the order of tag template, DNA synthesis template, and primer binding site. In this case, in the DNA molecule (such as genomic DNA), the 5'→3' explanation of the resulting cleavage site in the spacer-unbound strand may involve the sequence delivered by the DNA synthesis template being located at the first nucleotide, and the tag sequence delivered by the tag template being located at the second nucleotide. In other words, the positional relationship between the tag sequence placed on the DNA molecule and the sequence complementary to the DNA synthesis template in the spacer-unbound strand may have the following structure: v-[sequence complementary to the DNA synthesis template]-[tag sequence]. In this case, v refers to the resulting cleavage site.
[0285] Preferably, the DNA synthesis template, tag template, and primer binding site may be located in such order in the 5'→3' direction of the tpegRNA, but are not limited to this arrangement.
[0286] In the case of a 5' elongation region (i.e., a tpegRNA having an elongation region attached to the 5' end of a conventional gRNA), the following positional relationships may be observed. In some embodiments, the tag template, DNA synthesis template, and primer binding site may be located in the 5'→3' direction of the tpegRNA in the order described.
[0287] In some embodiments, the tag template may be located between the DNA synthesis template and the primer binding site. In some embodiments, the tag template may be located between the gRNA core and the DNA synthesis template. In some embodiments, the tag template may be located between the spacer and the DNA synthesis template. In some embodiments, the DNA synthesis template may be located between the tag template and the primer binding site. In some embodiments, the DNA synthesis template may be located between the tag template and the gRNA core. In some embodiments, the DNA synthesis template may be located between the tag template and the spacer. An exemplary embodiment of tpegRNA described as comprising a DNA synthesis template, a tag template, and a primer binding site is disclosed in Figure 2.
[0288] tpegRNA extension region (2) In some embodiments, tpegRNA may be described as comprising a homology region, an editing template, a tag template, and a primer binding site, thereby enabling the description of cases where the tag template is located between the editing template and the homology region. In some embodiments, tpegRNA may comprise a first region containing a homology region, a second region containing an editing template, a third region containing a tag template, and a fourth region containing a primer binding site. In this case, part or all of the first region may be a homology region. In this case, part or all of the second region may be an editing template. In this case, part or all of the third region may be a tag template. In this case, part or all of the fourth region may be a primer binding site.
[0289] The positional relationships between each element based on the primer binding site, tag template, and DNA synthesis template are described in detail in the previous section. Therefore, the positional relationships between the homology region, editing template, and tag template will be described below. As described above, the tag template is placed in the genomic DNA by polymerase and can be described as part of the DNA synthesis template. In some embodiments, including the following embodiments in the section "tpegRNA extension region (2)" of this specification, the tag template may be described as being included in the DNA synthesis template, and this will not mislead those skilled in the art. We will illustrate tpegRNAs that include a 3' extension region. In some embodiments, the tag template may be located downstream of the editing template, i.e., between the primer binding site and the editing template. In some embodiments, the tag template may be located downstream of the homology region, i.e., between the homology region and the primer binding site. In some embodiments, the tag template may be located between the editing template and the homology region. In some embodiments, the tag template may be located upstream of the homology region, i.e., between the homology region and the gRNA core. In some embodiments, the tag template may be located upstream of the editing template, i.e., between the editing template and the gRNA core. An exemplary embodiment of tpegRNA described as comprising a homology region, editing template, tag template, and primer binding site is disclosed in Figure 3.
[0290] Engineered tpegRNA Some embodiments of this application provide engineered tpegRNA (etpegRNA). pegRNA, epegRNA, and etpegRNA developed from tpegRNA are sometimes referred to as tpegRNA. In other words, the term “tpegRNA” in this application should be understood to include embodiments of etpegRNA. etpegRNA refers to pegRNA in which the elongation region of tpegRNA further includes a 3' engineering region which is an element of epegRNA. In other words, etpegRNA includes an elongation region containing a tag template, a DNA synthesis template, a primer binding site, and a 3' engineering region. In some embodiments, the 3' engineering region may include an RNA protection motif. In some embodiments, in addition to the RNA protection motif, the 3' engineering region may further include a linker for ligating the RNA protection motif. For example, each of the above elements of etpegRNA may be located in the elongation region in the 5'→3' direction in the order of DNA synthesis template, tag template, primer binding site, and 3' engineering region.
[0291] Unlike typical pegRNA (pegRNA without a tag template), tpegRNA contains a tag template for inserting a tag sequence into a DNA molecule.
[0292] Examples of tools for predicting off-target effects in prime editing The tool for predicting off-target in prime editing according to this application includes, as described above, at least the following two elements: Prime editor protein; and tpegRNA.
[0293] In some embodiments, the tool for predicting off-target reactions in prime editing may include additional elements. For example, one or more of the following may be included in the tool for predicting off-target reactions in prime editing: dominant-negative MLH1 (dnMLH1), Cas protein, guide RNA (such as conventional sgRNA), additional prime editing protein, pegRNA, and additional tpegRNA (such as tpegRNA containing an editing template with a different sequence than that of the tpegRNA used), but are not limited to these. Those skilled in the art will be able to improve or optimize the off-target prediction system for prime editing of this application by using appropriate additional elements.
[0294] The present invention provides an off-target prediction method for prime editing, designed based on the prime editing mechanism, which is a method for confirming or analyzing off-target information in prime editing. The prime editing mechanism is characterized by including a step of placing a desired edit on a target DNA molecule using pegRNA containing a DNA synthesis template (such as an RT template) used as a template in polymerization processes (such as reverse transcription). The present invention provides an off-target prediction method for prime editing, based on the characteristic mechanism of prime editing, which confirms or analyzes off-target information in prime editing by inserting a tag sequence into a target DNA molecule and confirming information about the inserted tag sequence. Therefore, the present invention provides an off-target prediction method that uses the characteristic mechanism of prime editing described above in the process of inserting a tag sequence.
[0295] Below, we will disclose an example of the mechanism for inserting a tag into the DNA molecule to be edited in the off-target prediction method of this application. This disclosure is made for the convenience of those skilled in the art who are reading this specification, and the scope of this specification should not be limited by the following description.
[0296] Below, we will disclose an example of a mechanism for inserting a tag into a DNA molecule using Prime Edit version 2 tpegRNA and Prime Editor protein.
[0297] Prime editing proteins (including nCas9 and reverse transcriptases MMLV_RT(D200N)(T330P)(L603W)(T306K)(W313F)) and tpegRNA form a complex. The gRNA core of tpegRNA is sometimes called the gRNA scaffold or backbone sequence, and refers to the sequence in gRNA, pegRNA, or tpegRNA that is responsible for binding to Cas9 or its equivalent. tpegRNA can bind to the Cas protein contained in the prime editing protein via the gRNA core.
[0298] The prime editor protein / tpegRNA complex localizes to potential off-target sites based on the spacer sequence and the PAM sequence. The tpegRNA spacer sequence can form complementary bindings with the target (on-target or off-target) sequence of the complementary DNA molecule. In this case, the complementary binding may or may not contain mismatches, or it may contain one or more mismatches. The mismatches may be one or more of the bulge mismatches and base mismatches known to cause off-target effects. Furthermore, off-target effects can occur from mismatches between parts of the sequence in the elongation region and the sequence of the genomic DNA. Moreover, the sites where the prime editor protein / tpegRNA complex localizes may not be limited to the PAM sequence. When compared to an on-target sequence, predicted off-target sequences (such as off-target candidates) may include one or more mismatches selected from one or more PAM mismatches, one or more spacer mismatches (i.e., mismatches present in the protospacer, which is the sequence corresponding to the spacer sequence), one or more PBS mismatches (i.e., mismatches present in the primer sequence, which is the sequence corresponding to the PBS sequence), and one or more DNA synthesis template mismatches (i.e., mismatches present in the sequence corresponding to the DNA synthesis template).
[0299] The Cas protein (nCas9 in PE2) of the prime editor protein creates a nick between -3 and -4 nucleotides relative to the 5' position of the PAM sequence (5'-NGG-3') located upstream of the spacer-unbound PAM sequence. As a result, the tag sequence can be inserted into a window of 1 to 100 nucleotides downstream of the nick site. The tag sequence can be inserted into a region of the PAM sequence ranging from approximately -4 to +100 nucleotides. Figure 4 shows an example of a DNA molecule with a resulting off-target nick and the prime editor protein / tpegRNA complex that induces the nick.
[0300] Upstream of the site where the nick is created, PBS anneals to a region that functions as a primer (a portion of the spacer-unbound strand of the DNA molecule, sometimes called a primer). Figure 5 shows the annealing of PBS to the primer.
[0301] After annealing, reverse transcription is performed using reverse transcriptase, with the tag template and DNA synthesis template as templates. Reverse transcription is performed in the 5'→3' direction relative to the polymerized nucleotide chain, or in other words, in the 5'→3' direction relative to the spacer-unbound chain. A sequence complementary to the tag template (tag sequence) is added to the endogenous DNA strand by reverse transcription, followed by the addition of a sequence complementary to the DNA synthesis template to the endogenous DNA strand.
[0302] Figure 6 shows the tag sequences and edits added to the endogenous DNA strand by reverse transcription. The tag sequences and sequences corresponding to the DNA synthesis template (edits, homology regions, and complementary sequences, etc.) added to the endogenous DNA strand construct the 3' DNA flap. The 5' flap is removed, and the tag sequences and edits are finally incorporated into the DNA molecule through the repair system.
[0303] Through the process described above, the tag sequence is inserted into a region where editing can be performed by prime editing. Therefore, the tag sequence can be inserted into sites where off-target editing can occur as well as on-target editing. Accordingly, by confirming the presence and / or location of the tag sequence, the site and / or possibility of off-target editing can be verified. Subsequently, the tag sequence is analyzed using methods that allow for specific analysis of the tag sequence, such as tag-specific amplification or sequencing. Through the analysis of the tag sequence, information about the tag sequence can be obtained, such as the type of DNA molecule in which the tag sequence is inserted (e.g., chromosome type), the site where the tag sequence is inserted (e.g., the site of the DNA molecule in which the tag sequence is inserted), and the insertion rate of the tag sequence at each site. Based on the information about the tag sequence, information about potential off-target editing during prime editing can be obtained.
[0304] The scenarios for inserting a tag sequence into a target DNA molecule (such as genomic DNA) are not particularly limited. In some embodiments, tag insertion may not disrupt the rest of the prime editing pattern. In this case, if the tag sequence is removed from the product resulting from the prime editing, the product resulting from the prime editing without the tag sequence may be the same as the prime editing pattern induced by pegRNA without the tag template. For example, the tag sequence may be placed with edits at one or more off-target candidate sites and / or on-target sites. In some embodiments, tag insertion may disrupt the rest of the prime editing pattern. For example, the tag sequence may be placed without edits at one or more off-target candidate sites and / or on-target sites. In another example, edits may be placed without a tag sequence at one or more off-target candidate sites and / or on-target sites. In certain embodiments, the tag sequence may be placed with edits at one or more off-target candidate sites and / or on-target sites. The off-target prediction system of the present application includes a process of contacting the genomic DNA of a cell with a prime editor protein and tpegRNA, and then analyzing the genomic DNA. The process of the off-target prediction system described in this application will be described in detail below.
[0305] Contact between cellular genomic DNA and prime editor proteins and tpegRNA Overview of contact with genomic DNA The present invention relates to the identification of off-target prediction in prime editing, which may occur during the DNA editing process by prime editing. In other words, as a result of the present invention's off-target prediction method for prime editing, information regarding potential off-target candidates that may occur during the DNA editing process by prime editing can be derived. For example, through the present invention's off-target prediction method, it is possible to derive whether or not off-target candidates exist, the off-target candidate site, and the off-target candidate score related to true off-targets. In order to obtain information regarding off-targets that occur during the DNA editing process, the prime editor protein and tpegRNA must be in contact with the target DNA. Once contact with the target DNA is achieved, a tag insertion mechanism, including a DNA cleavage process, can be executed. The target DNA may be, for example, the genomic DNA of a cell. As described above, the present invention's off-target prediction method can be classified as one of the cell-based off-target prediction methods, and contact between the genomic DNA of a cell and the prime editor protein and tpegRNA can be performed intracellularly.
[0306] The cells used in the off-target prediction method in prime editing are not particularly limited. In some embodiments, the cells may be animal cells or plant cells. In some embodiments, the cells may be human cells or cells from non-human animals (e.g., mouse, rat, monkey, chimpanzee, dog, cat, cow, pig, horse, sheep, etc.), but are not particularly limited. In some embodiments, the cells used in the off-target prediction method of the present invention may be patient-derived cells. In some embodiments, the cells used in the off-target prediction method of the present invention may be cells from cell lines (e.g., human, mouse, monkey, rat cell lines). In certain embodiments, the cells may be human cells or human cell lines. Examples of cell line-derived cells include, but are not limited to, 3T3 cells, A549 cells, HeLa cells, HEK 293 cells, K562 cells, Huh7 cells, Jurkat cells, OK cells, Ptk2 cells, or Vero cells.
[0307] One embodiment of the off-target prediction system of the present invention may include a step of contacting the genomic DNA of a cell with a prime editor protein and tpegRNA (or a prime editor protein / tpegRNA complex). This step of contacting the genomic DNA with the prime editor protein and tpegRNA can be performed intracellularly or in the nucleus of a cell, but is not particularly limited. To contact the genomic DNA with the prime editor protein and tpegRNA, cells containing the prime editor protein and tpegRNA must be prepared. The following describes in detail cells containing the prime editor protein and tpegRNA and the method of preparing them.
[0308] Cells including tools for predicting off-target effects in prime editing In some embodiments, the off-target prediction method of the present invention may comprise a step of preparing cells that include a tool for predicting off-target effects in prime editing.
[0309] Some embodiments of the present invention provide cells that include tools for predicting off-target behavior in prime editing.
[0310] Tools for predicting off-target effects in prime editing include a prime editor protein and tpegRNA. In some embodiments, tools for predicting off-target effects in prime editing may include additional elements. For example, one or more of the following may be further included in tools for predicting off-target effects in prime editing: dominant-negative MLH1 (dnMLH1), Cas protein, guide RNA (such as conventional sgRNA), additional prime editing protein, pegRNA, and additional tpegRNA (such as tpegRNA containing an editing template with a different sequence than that of the tpegRNA used).
[0311] Method for preparing cells, including an off-target prediction tool for prime editing. The preparation of cells containing tools for predicting off-target effects in prime editing can be achieved by introducing each element of the prime editing tool into the cells (e.g., by electroporation) or by introducing nucleic acids encoding each element of the prime editing tool into the cells. The process for preparing cells containing tools for predicting off-target effects in prime editing will be described in detail below.
[0312] In some embodiments, the step of preparing cells containing a tool for predicting off-targets in prime editing may include the step of contacting cells with a prime editor protein or a nucleic acid encoding the prime editor protein, and tpegRNA or a nucleic acid encoding the tpegRNA.
[0313] In some embodiments, the step of preparing cells containing a tool for predicting off-target effects in prime editing may include the step of introducing a prime editor protein or a nucleic acid encoding the prime editor protein, and tpegRNA or a nucleic acid encoding the tpegRNA into cells. Cells in which the prime editor protein or the nucleic acid encoding the prime editor protein and tpegRNA or a nucleic acid encoding the tpegRNA are in contact, or cells into which the prime editor protein or the nucleic acid encoding the prime editor protein and tpegRNA or a nucleic acid encoding the tpegRNA have been introduced, may be called cells for analysis.
[0314] Contact between cells and each element of the tools for predicting off-target effects in prime editing may occur simultaneously (e.g., with a single composition or using an all-in-one vector) or at different times. For example, these tools may be introduced into cells by contacting cells with a composition containing a prime editor protein or a nucleic acid encoding the prime editor protein and tpegRNA or a nucleic acid encoding the tpegRNA. In another example, these tools may be introduced into cells by contacting cells with a first composition containing a prime editor protein or a nucleic acid encoding the prime editor protein, and then subsequently (or next) contacting the cells with a second composition containing tpegRNA or a nucleic acid encoding the tpegRNA. As described above, the process for introducing tools for predicting off-target effects in prime editing into cells is not particularly limited.
[0315] In some embodiments, the prime editor protein or the nucleic acid encoding the prime editor protein, and / or tpegRNA or the nucleic acid encoding the tpegRNA may be introduced into cells in vector or non-vector form.
[0316] In some embodiments, the prime editor protein may be a fusion protein consisting of a single molecule, or it may be in the form of a complex containing two or more molecules. For example, if the prime editor protein is a fusion protein in single-molecule form, the prime editor protein or the nucleic acid encoding the prime editor protein can be introduced into the cell. In another example, if the prime editor protein is in the form of a complex containing two or more molecules, each element constituting the prime editor protein or each nucleic acid encoding each element can be introduced into or delivered into the cell simultaneously (e.g., in the form of an assembled complex or by being encoded into a single vector) or separately (e.g., in the form of separate elements, by being encoded into separate vectors, or at appropriate time intervals).
[0317] In some embodiments, the prime editor protein or the nucleic acid encoding the prime editor protein and the tpegRNA or the nucleic acid encoding the tpegRNA may be introduced into the cell simultaneously (e.g., in the form of an assembled complex or by being encoded into a single vector) or separately (e.g., in the form of separate elements, by being encoded into separate vectors, or at appropriate time intervals). In some embodiments, the prime editor protein may be delivered or introduced into the cell in the form of a protein. In some embodiments, the prime editor protein may be delivered or introduced into the cell in the form of the nucleic acid encoding the prime editor protein. In some embodiments, the tpegRNA may be delivered or introduced into the cell in the form of RNA. In some embodiments, the tpegRNA may be delivered or introduced into the cell in the form of the nucleic acid encoding the tpegRNA.
[0318] In some embodiments, the prime editor protein or the nucleic acid encoding the prime editor protein (such as DNA encoding the prime editor protein) and / or tpegRNA or the nucleic acid encoding the tpegRNA (such as DNA encoding the tpegRNA) may be introduced into cells in the form of liposomes, plasmids, viral vectors, nanoparticles, or protein-translation domains (PTDs).
[0319] In some embodiments, a prime editor protein or a nucleic acid encoding the prime editor protein, and / or tpegRNA or a nucleic acid encoding the tpegRNA may be delivered or introduced into cells by any one of the following methods: electroporation, lipofection, microinjection, gene gun, viromosome, liposome, immunoliposome, and lipid-mediated transfection.
[0320] In some embodiments, nucleic acids encoding the prime editor protein (such as DNA, RNA, or a DNA or RNA hybrid encoding the prime editor protein) and / or nucleic acids encoding tpegRNA (such as DNA, RNA, or a DNA or RNA hybrid encoding tpegRNA) can be delivered or introduced into cells by methods known in the art. Alternatively, nucleic acids encoding the prime editor protein and / or nucleic acids encoding tpegRNA can be delivered into a target by vector, non-vector, or a combination thereof.
[0321] The vector may be a viral vector or a non-viral vector (e.g., plasmid). The non-vector may be naked DNA, a DNA complex, or mRNA.
[0322] Vector-based implementation In some embodiments, the prime editor protein or the nucleic acid encoding the prime editor protein and / or the tpegRNA and the nucleic acid encoding the tpegRNA may be introduced into the cell in the form of a vector, in other words, they may be delivered or introduced into a target by the vector.
[0323] In some embodiments, the vector may contain nucleic acids encoding the prime editor protein and / or nucleic acids encoding tpegRNA. In some embodiments, the nucleic acids encoding the prime editor protein may be contained in a single vector or may be divided and contained in multiple vectors. For example, the nucleic acids encoding the prime editor protein may be introduced or delivered into a cell by one, two, three, four, five or more vectors. In some embodiments, the nucleic acids encoding tpegRNA may be contained in a single vector or may be divided and contained in multiple vectors. For example, the nucleic acids encoding tpegRNA may be introduced or delivered into a cell by one, two, three, four, five or more vectors. In some embodiments, the nucleic acids encoding the prime editor protein and the nucleic acids encoding tpegRNA may be contained in a single vector or may be divided and contained in multiple vectors. For example, nucleic acids encoding prime editor proteins and nucleic acids encoding tpegRNA can be introduced or delivered into cells by one, two, three, four, five or more vectors.
[0324] In some embodiments, the vector may include one or more regulatory / controlling elements.
[0325] In this case, the regulatory / control element may be one or more selected from promoters, enhancers, introns, polyadenylation signals, Kozak consensus sequences, internal ribosome entry sites (IRESs), nuclear localization signals (NLSs) or nucleic acids encoding them, poly(A), splice acceptors, and 2A sequences. The promoter may be a promoter recognized by RNA polymerase II. The promoter may be a promoter recognized by RNA polymerase III. The promoter may be an inducible promoter. The promoter may be a target-specific promoter. The promoter may be a viral or nonviral promoter. As the promoter, an appropriate promoter may be selected depending on the regulatory region.
[0326] In some embodiments, the vector may be a viral vector or a recombinant viral vector. The virus may be a DNA virus or an RNA virus.
[0327] In this case, the DNA virus may be a double-stranded DNA (dsDNA) virus or a single-stranded DNA (ssDNA) virus.
[0328] In this case, the RNA virus may be a single-stranded RNA (ssRNA) virus. The virus may be, but is not limited to, a retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, or herpes simplex virus. The AAV vector may be, but is not limited to, any one selected from, for example, AAV1, AAV2, AAV5, AAV6, AAV8, AAV9, AAVrh.10, AAVrh.74, and AAVhu.37. Examples of AAV vectors used in research or clinical practice are disclosed in detail in [Wang, Dan, Phillip WL Tai, and Guangping Gao. "Adeno-associated virus vector as a platform for gene therapy delivery." Nature reviews Drug discovery 18.5(2019):358-378.], the full contents of which are incorporated herein by reference. Typically, viruses can infect a host (such as a cell) and introduce nucleic acids that encode the genetic information on the virus into the host, or insert nucleic acids that encode the genetic information into the host genome. Viruses with such characteristics can be used to introduce nucleic acids encoding a target protein or target sequence into a target (such as a cell). Furthermore, the target protein and target sequence can be expressed within the host.
[0329] Non-vector-based implementation In one embodiment, a prime editor protein or a nucleic acid encoding the prime editor protein and / or tpegRNA and a nucleic acid encoding the tpegRNA can be introduced into a cell by non-vector-based introduction.
[0330] In some embodiments, one or more of the prime editor protein or the nucleic acid encoding the prime editor protein and / or tpegRNA and the nucleic acid encoding the tpegRNA may be introduced into cells by non-vector-based introduction.
[0331] In some embodiments, one or more of the prime editor protein or the nucleic acid encoding the prime editor protein and / or tpegRNA and the nucleic acid encoding the tpegRNA may be introduced into the cell by one or more non-vectors. For example, Prime Editor protein or the Prime The nucleic acid encoding the editor protein, and tpegRNA or the nucleic acid encoding said tpegRNA, can be introduced or delivered into the cell by one, two, three, four, five or more non-vectors.
[0332] Non-vectors are Prime Editor protein or the Prime The non-vector may include nucleic acids encoding an editor protein and tpegRNA or nucleic acids encoding said tpegRNA. The non-vector may be naked DNA, a DNA complex, mRNA, or a mixture thereof. The non-vector may be delivered or introduced into the target by electroporation, gene gun, sonication, magnetofection, transient cell compression or squeezing (disclosed in Lee, et al, (2012) Nano Lett., 12, 6322-6327), lipid-mediated transfection, dendrimers, nanoparticles, calcium phosphate, silica, silicates (ormosil), or a combination thereof. For example, delivery by electroporation can be carried out by mixing cells with nucleic acids encoding the desired elements in a cartridge, chamber, or cuvette, and then applying electrical stimulation of a defined duration and amplitude. In another example, the non-vector may be delivered using nanoparticles. The nanoparticles may be inorganic nanoparticles (e.g., magnetic nanoparticles, silica, etc.) or organic nanoparticles (e.g., polyethylene glycol (PEG) coated lipids, etc.). The outer surface of the nanoparticles may be conjugated with a positively charged polymer (e.g., polyethyleneimine, polylysine, polycerin, etc.) to enable adhesion.
[0333] Delivery or introduction in the form of peptides, polypeptides, proteins, or RNA. In one embodiment, the prime editor protein and / or tpegRNA may be delivered or introduced into a target by methods known in the art. The peptide, polypeptide, protein, or RNA may be delivered or introduced into a cell by electroporation, microinjection, transient cell compression or squeezing (disclosed in Lee, et al, (2012) Nano Lett., 12, 6322-6327), lipid-mediated transfection, nanoparticles, liposomes, peptide-mediated delivery, or a combination thereof.
[0334] As described above, cells containing prime editor protein and tpegRNA are obtained. The prime editor protein and tpegRNA (or prime editor protein / tpegRNA complex) in the cells can come into contact with the cell's genomic DNA. The results that can be achieved by the contact between the cell's genomic DNA and the prime editor protein and tpegRNA are described in detail below.
[0335] Results of contacting genomic DNA with prime editor protein and tpegRNA (tagmentation) As a result of contacting genomic DNA with a prime editor protein and tpegRNA, a tag sequence and a sequence complementary to the tag sequence may be inserted into the genomic DNA. In other words, a tag may be placed in the genomic DNA. Such a process of placing a tag in genomic DNA is sometimes called tagmentation. As a result of contact, the tag may be placed in off-target candidate regions and / or on-target regions. After contacting genomic DNA with a prime editor protein and tpegRNA, the resulting genomic DNA is sometimes called the genomic DNA to be analyzed. In some embodiments, the genomic DNA to be analyzed may not contain a tag. This applies when there are no off-target candidates, or when tag sequence placement in the genomic DNA fails. In some embodiments, the genomic DNA to be analyzed may contain a tag. Genomic DNA to be analyzed containing a tag is sometimes called tagged DNA (or tagged DNA). The tag is located in off-target candidate sites (i.e., candidate off-target regions) and / or on-target regions. By analyzing the tag inserted into the genomic DNA, it may be possible to identify potential candidate off-target regions that could be true off-targets. For example, the genomic DNA to be analyzed may contain one or more tags. By analyzing the presence or absence of each tag, each tagging site, etc., one or more off-target candidates can be identified. For example, the off-target prediction method of this application can be performed on a population of cells. The genomic DNA to be analyzed from some cells in the cell population may contain one or more tags. The genomic DNA to be analyzed from some cells in the cell population may not contain tags. By analyzing the genomic DNA of each cell in the cell population, it is possible to identify one or more off-target candidates. When a tag is inserted into an off-target candidate region, the tagmentation rate for each candidate off-target region can be obtained. Furthermore, by inserting tags into on-target regions as well, the tagmentation rate for on-target regions can be obtained.The tagging rate may be, for example, approximately 0.001%, 0.01%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, or 100%, or may fall within the range of two values selected from the above values.
[0336] Manipulated cells containing manipulated genomic DNA In some embodiments, manipulated cells containing a manipulated genome may be provided. For example, if the genomic DNA to be analyzed contains a tag, i.e., if the tag has been successfully placed on the genomic DNA to be analyzed, the genomic DNA to be analyzed may be referred to as a manipulated genome. For example, if the genomic DNA to be analyzed contains edits, i.e., if edits have been successfully placed on the genomic DNA to be analyzed, the genomic DNA to be analyzed may be referred to as a manipulated genome. In some embodiments, the manipulated genome DNA may contain one or more of the tags and edits. In some embodiments, a population of cells containing manipulated cells may be provided.
[0337] Analysis of DNA to be analyzed Overview of DNA analysis The off-target prediction system in prime editing of this application includes a step of analyzing the DNA to be analyzed. When the off-target prediction system in prime editing of this application is performed on cells, the DNA to be analyzed may be the genomic DNA to be analyzed. The analysis of the DNA to be analyzed will be explained using the analysis of the genomic DNA to be analyzed as an example. The genomic DNA to be analyzed may be the DNA of one or more genomes. The analysis of the genomic DNA to be analyzed may include, but is not particularly limited to, a step of analyzing the DNA of one or more genomes. By analyzing the genomic DNA to be analyzed, it becomes possible to obtain tagmentation information about the genomic DNA. For example, tagmentation information is not particularly limited, but may include whether or not a tag sequence is included in the genomic DNA to be analyzed, the location of each tag sequence (tagging site, etc.) for one or more tag sequences in the genomic DNA, the tagmentation rate at one or more tagging sites, etc. Information on off-target candidates can be obtained based on the tagmentation information. For example, information on off-target candidates may include, but is not particularly limited to, information on one or more off-target candidates, the off-target score of one or more off-target candidates, etc.
[0338] Analysis method Tagmentation information can be obtained by analyzing the target genomic DNA. The target genomic DNA may be manipulated genomic DNA. The off-target prediction system of this application is characterized by confirming information about sites where off-target events may occur based on tag sequences incorporated into the manipulated genome. Information about one or more tag sequences contained in the manipulated genome can be confirmed by methods known in the art or methods under development, but is not particularly limited. Information about tag sequences may include, but is not limited to, one or more of the following: whether or not each tag sequence is inserted, the chromosome in which each tag sequence is inserted, the site in which each tag sequence is inserted (e.g., a site within a chromosome), the insertion rate of the tag sequences, and the insertion rate for each site in which the tag sequences are inserted. For example, information about tag sequences can be confirmed by tag sequence analysis methods including tag-specific amplification and sequencing, but is not particularly limited. For information on the analysis methods of tag sequences, please refer to [Tsai, Shengdar Q., et al. "GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases." Nature biotechnology 33.2(2015):187-197.; Kim, Daesik, et al. "Digenome-seq: genome-wide profiling of CRISPR-Cas9 off-target effects in human cells." Nature methods 12.3 (2015): 237-243.; and Kim, Do Yon, et al. "Unbiased investigation of specificities of prime editing systems in human cells." Nucleic acids research 48.18(2020):10576-10589.], etc., and the entire contents of each of these are incorporated herein by reference.
[0339] In some embodiments, the analysis of the target genomic DNA may be tag-specific analysis (e.g., analysis to locate sites where tag sequences exist). In some embodiments, the method for analyzing the target genomic DNA may include tag-specific amplification. In some embodiments, the method for analyzing the target genomic DNA may include sequencing. In some embodiments, the analysis of the target genomic DNA may include tag-specific amplification and sequencing.
[0340] In some embodiments, the analysis of the target genomic DNA can be performed by DNA analysis methods well known to those skilled in the art. In some embodiments, the analysis of the target genomic DNA may be performed by a process that includes one or more PCR-based analyses selected from [Cameron, Peter, et al. "Mapping the genomic landscape of CRISPR-Cas9 cleavage." Nature methods 14.6(2017):600-606.] and sequencing ([Metzker, Michael L. "Sequencing technologies—the next generation." Nature reviews genetics 11.1(2010):31-46.; and Kumar, Kishore R., Mark J. Cowley, and Ryan L. Davis. "Next-generation sequencing and emerging technologies." Seminars in thrombosis and hemostasis. Vol.45. No.07. Thieme Medical Publishers, 2019.]) (such as DNA sequencing).
[0341] For example, one or more sequencing methods may be used, including but not limited to whole-genome sequencing (WGS), deep sequencing, high-throughput sequencing (HTS), de novo sequencing, second-generation sequencing, next-generation sequencing, third-generation sequencing, large-scale sequencing, shotgun sequencing, long-read sequencing, and short-read sequencing. For example, Hi-seq sequencing can be used. For example, Mi-seq sequencing can be used. For example, two or more sequencing methods may be used to analyze the DNA to be analyzed. In specific examples, a process involving Hi-seq and Mi-seq may be included in the analysis of the DNA to be analyzed. In one embodiment, the sequencing depth in a sequencing method used to analyze the target genomic DNA may be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000×. In one embodiment, the sequencing depth may be within the range of two values selected from the above values. In one embodiment, the sequencing depth may be less than, equal to, or greater than the above values. In certain embodiments, the sequencing depth in the sequencing used for analysis may be about 10 to 40 ×. The sequencing depth is not particularly limited, and any sequencing depth that allows for the confirmation of the presence or absence and / or location of the tag sequence in the genomic DNA being analyzed may be permitted.
[0342] In some embodiments, the analysis of the genomic DNA to be analyzed may include a tag-specific amplification process. Through tag-specific amplification, an amplified tag-specific library may be generated. In some embodiments, the analysis of the genomic DNA to be analyzed may include sequencing of the amplified tag-specific library.
[0343] Tagmentation information can be obtained through the analysis of the target genomic DNA. In some embodiments, the analysis of the target genomic DNA may include the steps of generating a tag-specific library from the target genomic DNA and sequencing the tag-specific library. In some embodiments, the analysis of the target genomic DNA may include the steps of generating an amplified tag-specific library from the target genomic DNA and sequencing the amplified tag-specific library. In some embodiments, the analysis of the target genomic DNA may include the steps of generating a tag-specific library from the target genomic DNA, amplifying the tag-specific library, and sequencing the amplified tag-specific library. For example, tag-specific primers and / or adapter-specific primers may be used in tag-specific amplification. For example, tag-specific amplification can be performed by PCR.
[0344] In some embodiments, the step of generating a tag-specific library from the genomic DNA to be analyzed may include one or more processes selected from a process of shearing the genomic DNA to be analyzed and a process of ligating the sheared genomic DNA using an adapter in order to generate a tag-specific library. For information on the amplification process of the tag-specific library, see [Tsai, Shengdar Q., et al. "GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases." Nature biotechnology 33.2(2015):187-197. and Liang, Shun-Qing, et al. "Genome-wide detection of CRISPR editing in vivo using GUIDE-tag." Nature communications 13.1(2022):1-14.].
[0345] In some embodiments, the analysis of the genomic DNA to be analyzed may further include one or more processes from cell disruption, incubation, RNA removal, and DNA purification. These processes can be carried out, for example, after contacting the genomic DNA with a prime editor protein and tpegRNA.
[0346] Retrieving tagging information Tagmentation information can be obtained through the analysis of the DNA targeted for analysis as described above. Tagmentation information is information obtained based on tag sequences and / or information related to tag sequences present in the genomic DNA targeted for analysis.
[0347] For example, tagmentation information may be information obtained based on information about tag sequences present in the DNA of one target genome. In another example, tagmentation information may be information obtained based on information about tag sequences present in the DNA of multiple target genomes. It will be understood that the analysis of target genome DNA includes all embodiments of the analysis of DNA of one or more target genomes. For example, tagmentation information may include, but is not limited to, one or more of the following: whether or not each tag sequence is inserted, the chromosome in which each tag sequence is inserted, the site (e.g., within a chromosome) in which each tag sequence is inserted, the insertion rate of the tag sequence, and the insertion rate for each site in which the tag sequence is inserted.
[0348] In some embodiments, tagging information may include one or more of the following: Whether or not the tag sequence is included in the target genomic DNA; The location of each tag sequence in the genomic DNA for one or more tag sequences; and The tag maintenance rate for one or more tag sequences.
[0349] For example, if the genomic DNA being analyzed contains a tag sequence, the presence of the tag sequence may be related to the presence of an on-target or candidate off-target region. As described above, one or more tag sequences may be present in the DNA of one genomic DNA, or one or more tag sequences may be present in the DNA of multiple genomic DNAs being analyzed. Consequently, whether or not a tag sequence is present in the genomic DNA being analyzed provides information about whether or not one or more tag sequences are present in the DNA of one or more genomic DNAs being analyzed. For example, in the case of DNA from multiple genomic DNAs being analyzed, if the first genomic DNA is not present with a tag sequence, but the second genomic DNA is, then it can be determined that the tag sequence is present in the genomic DNA being analyzed.
[0350] For example, for one or more tag sequences, the location of each tag sequence in the genomic DNA can be derived by analyzing the region where the tag sequence exists (this is sometimes called the "tagging site"). For instance, if, among the DNA of multiple target genomes, one target genome (the first target genome DNA) contains the first tag sequence and another target genome DNA (the second target genome DNA) contains the second tag sequence, the region of the first tag sequence may be called the first site, and the region of the second tag sequence may be called the second site. In another example, there may be multiple tag sequences in the DNA of one target genome, in which case one tag sequence may be called the first tag sequence and another tag sequence may be called the second tag sequence. In this case, the location of each tag sequence in the genomic DNA for one or more tag sequences may include the first site, the second site, or both the first and second sites. Here, the first and second sites are related to target sites (on-target sites and / or candidate off-target sites). The first and second sites, other than the on-target site, may both be candidate off-target sites. The first and second sites may refer to the same or different sites. In this case, information about the sites, such as the first and second sites, includes information about chromosome number and information about the site on a specific chromosome.
[0351] For example, the tagmentation rate of one or more tag sequences can be derived from the frequency of discovery at each tagging site. For instance, if, in the analysis of the target genomic DNA, a tag sequence is found 10 times at the first site and 5 times at the second site, the tagmentation rate at the first site is twice that of the second site. The tagmentation rate may, but is not limited to, be related to the likelihood that the corresponding off-target candidate is a true off-target.
[0352] In some embodiments, the process of obtaining tagmentation information by analyzing the target genomic DNA may further include additional processes for obtaining tagmentation information. For example, this process may further include processes for processing the information (or data) and / or processes for normalizing the obtained information (or data). For example, it may further include a process for comparing the obtained cleavage information with information about a predetermined on-target. The process of obtaining cleavage information may further include additional processes as described above, and is not particularly limited to these.
[0353] In some embodiments, tagmentation information may further include, but is not limited to, other information obtainable through analysis of the genomic DNA being analyzed (such as DNA sequencing).
[0354] Acquisition of off-target information Off-target information can be obtained based on tagmentation information. Those skilled in the art in the relevant field can obtain off-target information based on cleavage information without particular difficulty. Accordingly, the disclosure herein does not limit the process of the off-target prediction system. Those skilled in the art in the relevant field can obtain off-target information using tagmentation information obtained by analyzing the target genomic DNA, either through an appropriate process or without a separate process.
[0355] In some embodiments, the off-target prediction method of the present invention may include a process of verifying information about off-target candidates based on tagmentation information.
[0356] In some embodiments, information on off-target candidates may include information on the location of one or more off-target candidates in genomic DNA (such as information on candidate off-target regions). For example, information on off-target candidate sites may include information on each site (location in genomic DNA) of all off-target candidates. In other words, information on all candidate off-target regions can be obtained. Alternatively, information on one or more candidate off-target regions can be obtained instead of all candidate off-target regions. Among all off-target candidates, there may be true off-targets (such as actual off-targets resulting from the use of a prime editing system). Information on off-target candidate sites can be obtained based on the tagging information described above.
[0357] In one embodiment, information about off-target candidates may include the off-target score (e.g., off-target prediction score) of one or more off-target candidates. For example, information about off-target candidates may include the off-target score of each off-target candidate for all off-target candidates. For example, information about off-target candidates may include the off-target score of each off-target candidate for one or more off-target candidates. In other words, off-target scores can be obtained for all candidate off-target regions. Alternatively, off-target scores can be obtained for one or more candidate off-target regions instead of all candidate off-target regions. Information about the off-target scores of off-target candidates can be obtained based on the tagmentation information (e.g., information about the tagmentation rate) described above. In one embodiment, the rank of each off-target candidate may be calculated based on the obtained off-target score. For example, off-target candidates (e.g., candidate off-target regions) with high off-target scores may be ranked higher. For example, the off-target candidate with the highest off-target score may be ranked first. For example, a high off-target score for an off-target candidate may, but is not limited to, true off-target.
[0358] In one embodiment, information regarding off-target candidates may include information regarding the number of off-target candidates. For example, the total number of off-target candidates can be calculated. For example, in calculating the number of off-target candidates, overlapping sites may be counted as one. In another example, overlapping sites may be counted as multiple when calculating the number of off-target candidates. For example, if the number of identified candidate off-target regions X is 5, the total number can be counted as 1 or 5. Through such information regarding the number of off-target candidates, the total number of off-target candidates that may occur during the genome editing process by prime editing can be determined. In other words, the predicted total number of off-targets can be determined.
[0359] In one embodiment, information regarding off-target or off-target candidates may include, but is not limited to, one or more of the following: One or more off-target candidate sites in the genomic DNA of off-target candidates; The off-target score for each off-target candidate for one or more off-target candidates; and The number of predicted off-target candidates.
[0360] In some embodiments, the process of obtaining information about off-target candidates may further include additional processes for obtaining information about off-target candidates. For example, this process may further include processes for processing the information (or data) and / or processes for normalizing the obtained information (or data). For example, it may further include a process for comparing the obtained information about off-target candidates with a given on-target information. The process of obtaining information about off-target candidates may further include additional processes as described above, and is not limited to any other.
[0361] In some embodiments, information regarding off-target candidates may further include additional information that helps predict off-target occurrences that may arise from the use of the prime editing system, but is not limited to other information.
[0362] Comparison of off-target candidates with tpegRNA As described above, tags can be inserted into off-target candidate sites (i.e., candidate off-target regions). In conventional CRISPR / Cas systems, it is known that off-target can occur due to partial but sufficient matching between the guide sequence and the target sequence. Similarly, in prime editing systems, off-target is expected to result from partial but sufficient matching between the sequences of each element of tpegRNA and the target sequence, although the reasons for off-target are not limited herein. In some embodiments, off-target may result from one or more mismatches between the tpegRNA sequence and the off-target sequence. In this case, the mismatches include base mismatches (e.g., a difference of one or more nucleotides) and bulge mismatches (e.g., the addition or deletion of one or more nucleotides). In some embodiments, the off-target (or off-target candidate) sequence may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more mismatches compared to the corresponding tpegRNA sequence. In some embodiments, the off-target (or off-target candidate) sequence may have 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity compared to the corresponding tpegRNA sequence, or sequence identity within a range set by two values selected from the above values. For example, the tpegRNA spacer sequence and the off-target (or off-target candidate) spacer corresponding sequence may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches. In another example, the PAM sequence and the sequence corresponding to the off-target (or off-target candidate) PAM sequence may contain 1, 2, 3, 4, or 5 or more mismatches. For example, the sequence corresponding to the tpegRNA DNA synthesis template and the off-target (or off-target candidate) DNA synthesis template may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches.For example, sequences corresponding to homology regions of tpegRNA and homology regions of off-target (or off-target candidate) sequences may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches. For example, sequences corresponding to primer binding sites of tpegRNA and primer binding sites of off-target (or off-target candidate) sequences (such as sequences that function as primers) may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches. For example, one or more mismatches may exist in one or more sequences corresponding to spacers of off-target (or off-target candidate) sequences, PAM sequences of off-target (or off-target candidate) sequences, DNA synthesis templates of off-target (or off-target candidate) sequences, and primer binding sites of off-target (or off-target candidate) sequences.
[0363] Comparison of off-target candidates and on-target candidates As described above, tags can be inserted into off-target candidate sites (i.e., candidate off-target regions). Off-target candidate refers to an off-target predicted through a prediction system, which may or may not be a true off-target. In some embodiments, the off-target candidate region may refer to a specified site. In some embodiments, an on-target site or region, or an off-target candidate site or region, may be understood as a specific region, in which case this specific region may refer to a region consisting of approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 220, 240, 260, 280, 300, 320, 340, 360, 380, 400, 450, or 500 consecutive nucleotides, or a region consisting of a number of consecutive nucleotides greater than the above values. In some embodiments, a larger number of consecutive nucleotides allows for a more precise designation of the off-target or on-target region, because a larger number of nucleotides reduces the likelihood of identical sequences (duplicate sequences) existing in the genomic DNA.
[0364] An off-target or off-target candidate can be compared to the on-target sequence. In some embodiments, an off-target candidate or true off-target may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more mismatches (on-target mismatches) compared to the on-target sequence. In some embodiments, the off-target (off-target candidate) sequence may have 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99%, or 100% sequence identity compared to the corresponding on-target sequence, or within a range set by two values selected from the above values. The mismatches used in the comparison between the on-target and off-target are used to describe the differences in sequences between the off-target and the on-target. Furthermore, the term mismatch is used to include both nucleotide mismatches (e.g., a difference in nucleotides) and bulge mismatches (e.g., the addition or deletion of one or more nucleotides). For example, if the sequence corresponding to an off-target candidate spacer is GGCACTGaGGgTGGAGGTGG (SEQ ID NO: 51) and the sequence corresponding to an on-target spacer is GGCACTGCGGCTGGAGGTGG (SEQ ID NO: 52), the sequence corresponding to the off-target candidate spacer may be described as having two nucleotide mismatches (in lowercase) compared to the on-target sequence. In another example, if the sequence corresponding to an off-target candidate spacer is GGCACTGC--CTGGAGGTGG (SEQ ID NO: 53) and the sequence corresponding to an on-target spacer is GGCACTGCGGCTGGAGGTGG (SEQ ID NO: 54), the sequence corresponding to the off-target candidate spacer may be described as having two bulge mismatches (e.g., two buldion target mismatches) compared to the on-target sequence.In a further example, if the off-target candidate spacer sequence is GGCACTGCGGCTGGAGgTGG (SEQ ID NO: 55) and the on-target spacer sequence is GGCACT--GGCTGGAGGTGG (SEQ ID NO: 56), then the off-target candidate spacer sequence may be described as having one nucleotide mismatch and two bulge mismatches (a total of three mismatches) compared to the on-target sequence. Below, we will explain the off-target (or off-target candidate) sequence in comparison to the on-target sequence.
[0365] In some embodiments, the sequence corresponding to the off-target (or off-target candidate) spacer may include 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches (such as on-target mismatches). In some embodiments, the sequence corresponding to the off-target (or off-target candidate) PAM sequence may include 0, 1, 2, 3, 4, or 5 or more mismatches. In some embodiments, the sequence corresponding to the off-target (or off-target candidate) DNA synthesis template may include 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches. In some embodiments, the sequence corresponding to the off-target (or off-target candidate) homology region may include 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches. In some embodiments, the sequences corresponding to off-target (or off-target candidate) primer binding sites may contain 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches. In some embodiments, one or more mismatches may be present in one or more of the sequences corresponding to the off-target (or off-target candidate) spacer, the off-target (or off-target candidate) PAM sequence, the off-target (or off-target candidate) DNA synthesis template, and the off-target (or off-target candidate) primer binding sites.
[0366] In some embodiments, in one or more of the regions corresponding to the spacer, PAM, PBS, and DNA synthesis template, the off-target candidate (or off-target) region may contain 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 or more on-target mismatches, or may contain on-target mismatches within a range defined by two values selected from the above values. In some embodiments, in the regions corresponding to the spacer and DNA synthesis template, the off-target candidate (or off-target) region may contain 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 or more on-target mismatches, or may contain on-target mismatches within a range set by two values selected from the above values. In certain embodiments, the off-target candidate (or off-target) region may contain 0 to 20 on-target mismatches in the regions corresponding to the spacer and DNA synthesis template. In certain embodiments, the off-target candidate (or off-target) region may contain 1 to 15 on-target mismatches in the regions corresponding to the spacer and DNA synthesis template. In certain embodiments, the off-target candidate (or off-target) region may contain 1 to 10 on-target mismatches in the regions corresponding to the spacer and DNA synthesis template. In this case, an on-target mismatch refers to a mismatch determined by comparing the sequence of the on-target region with the sequence of the corresponding region. On-target mismatches may be counted on a single strand or on both strands. For example, an off-target candidate (or off-target) region may contain 0 to 10 on-target mismatches in the following regions of the spacer-unbound strand: (i) the region corresponding to the protospacer, and (ii) a region consisting of 5 to 20 nucleotides located downstream of the region corresponding to the protospacer.For example, an off-target candidate (or off-target) region may contain 0 to 10 on-target mismatches in a region ranging from -30 to +10 or -20 to +10 relative to the cleavage site (nick or DSB).
[0367] Relationship with the prime editing system being predicted The off-target prediction system of this application can be associated with a prime editing system that is subject to prediction. In this case, the prime editing system subject to prediction may refer to, but is not limited to, a prime editing system used in research or a prime editing system that has been determined to be used in treatment. In other words, the prime editing system subject to prediction refers to a prime editing system (or a genome editing process using a prime editing system) for which off-target prediction is necessary.
[0368] For example, if specific cells are used in the prime editing system being predicted, these specific cells can also be used in the method for predicting off-target effects according to this application. In another example, if specific cells are used in the prime editing system being predicted, cells other than the specific cells mentioned above can be used in the method for predicting off-target effects according to this application. For example, patient-derived cells can be used in the prime editing system being predicted, and the cells used in the off-target prediction system according to this application may be human cell lines.
[0369] In the method for predicting off-target effects of the present invention, for example, when using a tpegRNA having a specific sequence in the prime editing system to be predicted, a tpegRNA having the same or partially different sequence as the above sequence can be used. Similarly, when using a specific prime editor protein in the prime editing system to be predicted, a prime editor protein of the same or different type as the above protein can be used in the method for predicting off-target effects of the present invention. In another example, in the method for predicting off-target effects of the present invention, additional elements (e.g., dnMLH1, sgRNA, additional tpegRNA, etc.) may be used in addition to the elements of the prime editing system to be predicted, and are not particularly limited.
[0370] In this embodiment, a method for predicting off-target effects according to one embodiment of the present invention may further include a step of identifying the prime editing system to be predicted. The prime editing system to be predicted may be referred to as a predetermined prime editing system. A predetermined prime editing system may include a step of using one or more or a combination thereof of a predetermined cell (such as a cell to be subjected to genome editing by the prime editing system), a predetermined prime editor protein, and a predetermined pegRNA.
[0371] In one embodiment, the method for predicting off-target effects according to the present invention may further include a step of verifying or designing a predetermined prime editing system. By verifying a predetermined prime editing system, it becomes possible to design elements to be used appropriately in the off-target prediction system. In this case, the process of verifying the predetermined gene editing system can be performed before contacting the cell's genomic DNA with the prime editor protein and tpegRNA. An example of verifying a predetermined prime editing system (i.e., the one to be predicted) will be described below.
[0372] In one embodiment, the method for predicting off-target reactions according to the present invention may include a step of confirming a predetermined prime editing system. In this case, the step of confirming a predetermined prime editing system may include a step of confirming one or more of the following: a predetermined cell, a predetermined prime editor protein, and a predetermined pegRNA. The predetermined prime editing system, predetermined cell, predetermined prime editor protein, predetermined pegRNA, etc., may be used together with ordinal determinants, such as a first prime editing system, a first cell, a first prime editor protein, a first pegRNA, etc.
[0373] In certain embodiments, the step of verifying a predetermined prime editing system may include the step of verifying a predetermined cell. In certain embodiments, the same cell as the predetermined cell may be used in the off-target prediction system of the present invention. In certain embodiments, a different cell from the predetermined cell may be used in the off-target prediction system of the present invention. For example, the predetermined cell may be a human cell rather than a cell line, and a human cell line may be used in the off-target prediction system of the present invention. In some embodiments, the predetermined cell may be an animal cell or a plant cell. In some embodiments, the predetermined cell may be a human cell or a cell from a non-human animal (e.g., mouse, rat, monkey, chimpanzee, dog, cat, cow, pig, horse, sheep, etc.), but is not particularly limited. In some embodiments, the predetermined cell may be a patient-derived cell. In some embodiments, the predetermined cell may be a cell from a cell line (e.g., human, mouse, monkey, rat cell lines). Examples of cell lines derived from cells include, but are not limited to, 3T3 cells, A549 cells, HeLa cells, HEK 293 cells, K562 cells, Huh7 cells, Jurkat cells, OK cells, Ptk2 cells, or Vero cells.
[0374] In certain embodiments, the step of verifying a predetermined prime editing system may include the step of verifying a predetermined prime editor protein. In certain embodiments, the same prime editor protein as the predetermined prime editor protein may be used in the off-target prediction system of the present invention. In certain embodiments, a prime editor protein different from the predetermined prime editor protein may be used in the off-target prediction system of the present invention. For example, the predetermined prime editor protein may be a PE2 prime editor protein, but the prime editor protein used in the off-target prediction system of the present invention may be a PE2-nuclease prime editor protein or a PEmax-nuclease prime editor protein. Other types of prime editor proteins may be used to increase the tagmentation rate.
[0375] In certain embodiments, the step of confirming a predetermined prime editing system may include the step of confirming a predetermined pegRNA. In certain embodiments, the same tpegRNA as the predetermined pegRNA may be used in the off-target prediction system of the present invention (in this case, the same tpegRNA as the predetermined pegRNA indicates that all sequences except the tag template are the same). In certain embodiments, a tpegRNA different from the predetermined pegRNA may be used in the off-target prediction system of the present invention. The relationship between the predetermined pegRNA and tpegRNA used in the off-target prediction system of the present invention will be described below.
[0376] A given pegRNA can be called the first pegRNA, which includes a first spacer, a first DNA synthesis template, and a first primer binding site. The tpegRNA used in the off-target prediction system of this application is, for convenience, called the second tpegRNA. The second tpegRNA includes a second spacer, a second DNA synthesis template, a second tag template, and a second primer binding site. Furthermore, the second tpegRNA may further include a 3' engineering region. In this case, an etpegRNA developed based on an epegRNA other than the first pegRNA can be used in the off-target prediction method of this application.
[0377] In some embodiments, the second spacer may have the same arrangement as the first spacer, or it may have an arrangement that has about 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9% arrangement identity with the first spacer.
[0378] In some embodiments, the second primer binding site may have the same sequence as the first primer binding site, or the sequence may be approximately 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%. The sequences may have sequences with sequence identity of 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9%.
[0379] In some embodiments, the second DNA synthesis template may have the same sequence as the first DNA synthesis template, or it may have approximately 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 6% of the sequence of the first DNA synthesis template. The sequences may have sequences with sequence identity of 8%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9%.
[0380] In some embodiments, the second extension region may have the same sequence as the sequence of the first extension region excluding the tag template, or it may have a sequence of about 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, The sequences may have sequences with sequence identity of 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9%.
[0381] In some embodiments, the first pegRNA is pegRNA rather than epegRNA, but the tpegRNA used in the off-target prediction method of the present invention may further include a 3' engineering region (e.g., using etpegRNA).
[0382] In some embodiments, the first DNA synthesis template may include the first editing template, while the second DNA synthesis template may not include the editing template. In some embodiments, the first DNA synthesis template may include the first editing template, while the second DNA synthesis template may include the second editing template. In this case, the second editing template may have the same sequence as the first editing template, or it may have a sequence having approximately 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67% of the sequence of the first editing template. They may have sequence identity of 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9%.
[0383] In some embodiments, the second editing template may have a different sequence from the sequence of the first editing template. In some embodiments, the first DNA synthesis template may include a first homology region, and the second DNA synthesis template may include a second homology region. In some embodiments, the second homology region may hav...
Claims
1. A method for predicting off-target events that occur during the prime editing process using pegRNA, using tpegRNA corresponding to the pegRNA, The pegRNA comprises a spacer, a gRNA core, and an elongation region including an editing template and a primer binding site. The tpegRNA includes a spacer, a gRNA core, and an elongation region including a tag template and a primer binding site. The spacer of the pegRNA and the spacer of the tpegRNA have the same sequence, The primer binding site of the pegRNA and the primer binding site of the tpegRNA have the same sequence. The aforementioned method, (a) A step of preparing a population of cells, which includes one or more manipulated cells. Here, the manipulated cells contain manipulated genomic DNA including a tag sequence, The manipulated genomic DNA is as follows, involving the prime editor protein and the tpegRNA: (i) A process of contacting the genomic DNA with the prime editor protein and the tpegRNA, Here, the prime editor protein comprises a Cas protein and a reverse transcriptase. Here, the gRNA core of the tpegRNA interacts with the prime editor protein to form a complex of the prime editor protein and the tpegRNA. (ii) A process of inserting the tag sequence into the genomic DNA, Here, the insertion of the tag sequence is achieved through a reverse transcription process performed by the reverse transcriptase using the tag template of the tpegRNA as a reverse transcription template; Includes, generated through a process, (b) A step of obtaining tagment information by analyzing the results obtained through a process that includes sequencing the manipulated genomic DNA of one or more manipulated cells. Here, the tagging information includes information about one or more sites in which each tag sequence is inserted; and (c) A step of obtaining off-target information of the pegRNA based on the tagging information by comparing the tagging information with the on-target information of the pegRNA. Here, the off-target information of the pegRNA includes information regarding whether or not off-target candidates exist, and information regarding the regions of one or more off-target candidates. A method that includes [a certain feature].
2. The elongation region of the tpegRNA further includes an editing template, The method according to claim 1, wherein the editing template of the tpegRNA has the same sequence as the editing template of the pegRNA.
3. The gRNA core of the tpegRNA has the same sequence as the gRNA core of the pegRNA. The method according to claim 1.
4. The method according to claim 1, wherein the length of the tag template is 5 nt to 60 nt.
5. The method according to claim 4, wherein the sequence of the tag template of the tpegRNA is selected from sequence numbers 42 to 47.
6. The prime editor protein is selected from PE2, PE2max, and PE2-nuclease. The method according to claim 1, wherein in the tpegRNA, the spacer, the gRNA core, and the extension region are located in the order of 5'→3' direction.
7. The method according to claim 1, wherein the tpegRNA further comprises a 3' engineering region containing an RNA protection motif.
8. The method according to claim 1, wherein the manipulation of the DNA genome further involves one or more of dnMLH1, gRNA, additional Cas proteins, and additional prime editor proteins.
9. The method according to claim 1, wherein (b) comprises the step of specifically analyzing the manipulated genomic DNA tag.
10. The method for predicting off-target events is A step of obtaining on-target information of the pegRNA based on the tagging information, The method according to any one of claims 1 to 9, further comprising:
11. A method for predicting off-target events that occur during the prime editing process using pegRNA, using tpegRNA corresponding to the pegRNA, The pegRNA comprises a spacer, a gRNA core, and an elongation region including an editing template and a primer binding site. The tpegRNA includes a spacer, a gRNA core, and an elongation region including a tag template and a primer binding site. The spacer of the pegRNA and the spacer of the tpegRNA have the same sequence, The primer binding site of the pegRNA and the primer binding site of the tpegRNA have the same sequence. The aforementioned method, (a) The step of applying the tagmentation prime editing system to a population of cells, Here, the tagging prime editing system is Prime editor protein, or nucleic acid encoding the prime editor protein; and The tpegRNA, or the nucleic acid encoding the tpegRNA. Includes, As a result, the tagging prime editing system inserts a tag sequence into the genome of a portion of the population of cells. Here, the tag template of the tpegRNA is used as a template for inserting the tag sequence; (b) After (a), a step of obtaining tag information by sequencing the aforementioned group of cells, Here, the tagging information includes information about one or more sites where the tag sequence is inserted. Here, the tag sequence is inserted by the tagmentation prime editing system using the tag template of the tpegRNA as a template; and (c) A step of obtaining off-target information of the pegRNA based on the tagging information by comparing the tagging information with the on-target information of the pegRNA. Here, the off-target information of the pegRNA is Information regarding whether or not off-target candidates exist; and information regarding the regions of one or more off-target candidates. including, A method that includes [a certain feature].
12. A method for predicting off-target events that occur during the prime editing process using pegRNA, using tpegRNA corresponding to the pegRNA, The pegRNA comprises a spacer, a gRNA core, and an elongation region including an editing template and a primer binding site. The tpegRNA includes a spacer, a gRNA core, and an elongation region including a tag template and a primer binding site. The spacer of the pegRNA and the spacer of the tpegRNA have the same sequence, The primer binding site of the pegRNA and the primer binding site of the tpegRNA have the same sequence. The aforementioned method, (a) The step of treating a population of cells with the tpegRNA, Here, each cell in the population contains nucleic acid encoding the prime editor protein in order to express the prime editor protein. As a result, the tagmentation prime editing system inserts a tag sequence into the genome of a portion of the aforementioned population of cells. Here, the tag template of the tpegRNA is used as a template for inserting the tag sequence; (b) After (a), a step of obtaining tag information by sequencing the aforementioned group of cells, Here, the tagging information includes information about one or more sites where the tag sequence is inserted. Here, the tag sequence is inserted by the tagmentation prime editing system using the tag template of the tpegRNA as a template; and (c) A step of obtaining off-target information of the pegRNA based on the tagging information by comparing the tagging information with the on-target information of the pegRNA. Here, the off-target information of the pegRNA is Information regarding whether or not off-target candidates exist; and information regarding the regions of one or more off-target candidates. including, A method that includes [a certain feature].
Citation Information
Patent Citations
Off-target single nucleotide variants caused by single-base editing and high-specificity off-target-free single-base gene editing tool
EP3940078A1
Method to identify and validate genomic safe harbor sites for targeted genome engineering
US20200370067A1
Inhibition of unintended mutations in gene editing
WO2020156575A1
Methods and compositions for editing nucleotide sequences
WO2020191249A1
Methods and compositions for simultaneous editing of both strands of a target double-stranded nucleotide sequence
WO2021226558A1