A method for predicting possible off-targets in gene editing processes.
By physically destroying cells with a filter to expose genomic DNA to the Cas/gRNA complex, the method addresses off-target prediction limitations in CRISPR/Cas systems, offering accurate and efficient off-target identification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2026-04-14
AI Technical Summary
Current off-target prediction tools for genome editing processes, such as CRISPR/Cas systems, suffer from limitations like missing true off-target sites or excessive false positives, posing safety concerns due to potential severe side effects.
A method involving physical destruction of cells using a filter with pores smaller than cell size to expose genomic DNA to the Cas/gRNA complex, followed by analysis to identify cleavage sites and off-target candidates.
The method provides accurate and simple off-target prediction with reduced false-positive and missed rates, enhancing safety in genome editing processes.
Smart Images

Figure 0007845779000064 
Figure 0007845779000065 
Figure 0007845779000066
Abstract
Description
Technical Field
[0001] This application relates to a method for predicting possible off-targets in a gene editing process. The gene editing process may be, for example, a genomic DNA editing process using a CRISPR / Cas gene editing system.
Background Art
[0002] Since 2005, applications for investigational new drugs (INDs) have been submitted for various gene editors (e.g., zinc finger nuclease-based, TALEN-based, and CRISPR nuclease-based gene editors) (see [Mullard, Asher. "Gene-editing pipeline takes off." Nature Reviews Drug Discovery 19.6(2020):367-373.]). Unlike other drugs such as chemicals or antibodies that are generally associated with reversible side effects, the action of genome editing drugs is permanent. That is, actions that frequently occur at undesired positions in the genome editing process (i.e., off-target actions) pose important safety concerns, so it is particularly important to identify off-target sites throughout the entire genome for genome editing drugs. To confirm information about possible off-target actions in the genome editing process, many researchers have developed various methods for predicting off-target actions throughout the genome through various approaches.
[0003] However, currently developed off-target prediction tools (systems) have multiple limitations. For example, cell-based methods have problems such as missing true off-target sites. On the other hand, in vitro and in silico methods have problems such as showing an excessive number of false positive data points. [Disclosure] [Technical Problem]
[0004] In genome editing processes using gene editing tools (e.g., CRISPR / Cas gene editing systems), off-target problems can occur. Such off-target effects can result in severe side effects. Embodiments of the present invention provide a method for predicting off-target effects that may occur in genome editing processes. [Technical solution]
[0005] This invention provides a method for predicting off-target events that occur in gene editing processes using a gene editing system.
[0006] One embodiment of the present invention is a method for identifying off-target information that occurs in a genome editing process using a CRISPR / Cas genome editing system, (i) A step of preparing an initiation composition comprising Cas protein, guide RNA, and cells; (ii) A step of obtaining the composition to be analyzed by physically destroying the cells, wherein by physically destroying the cells, the genomic DNA comes into contact with the Cas / gRNA complex formed by the Cas protein and the guide RNA, thereby causing the genomic DNA to cleave at one or more cleavage sites; and (iii) The step of obtaining information about the cleavage site by analyzing the composition to be analyzed. The present invention provides a method that includes the following:
[0007] In a particular embodiment, the physical destruction of the cells may include passing the cells through a filter having pores, wherein the average diameter of the pores in the filter is smaller than the size of the cells.
[0008] In one particular embodiment, the force that causes the cells to pass through the filter may be pressure.
[0009] In a particular embodiment, the average diameter of the pores in the filter may be 5 to 15 μm.
[0010] In a particular embodiment, the physical destruction of the first cells may be achieved through the use of an extruder equipped with a filter having pores.
[0011] In one particular embodiment, the average diameter of the pores in the filter included in the extruder may be smaller than the size of the cells.
[0012] In a particular embodiment, the average diameter of the pores in the filter may be 5 to 15 μm.
[0013] In a particular embodiment, information regarding the rupture site is, The location of each of the one or more cleavage sites on the genomic DNA; The crack score of each of the one or more crack sites; and Number of rupture sites It may include one or more of the following:
[0014] In a particular embodiment, the method is (iv)(iii) Step of identifying off-target candidate information from the information about the fracture site obtained from (iv)(iii) You may want to prepare even more.
[0015] In a particular embodiment, the information regarding the off-target candidate is: The location on the genomic DNA of one or more off-target candidates; Off-target prediction scores for each of one or more off-target candidates; and Number of predicted off-target candidates It may include one or more of the following:
[0016] In a particular embodiment, the step of analyzing the composition to be analyzed may include the step of analyzing the cleaved genomic DNA contained in the composition to be analyzed by sequencing.
[0017] In a particular embodiment, the step of analyzing the composition to be analyzed may include the step of analyzing the cleaved genomic DNA contained in the composition to be analyzed by a PCR-based method.
[0018] In a particular embodiment, the cell membrane structure, including the cell membrane, may be destroyed by the physical destruction of the cell, thereby creating an environment in which the Cas / gRNA complex can come into contact with the genomic DNA derived from the cell.
[0019] In a particular embodiment, the cell may be physically destroyed, thereby disrupting the cell's membrane structure, including the nuclear membrane, which prepares an environment in which the Cas / gRNA complex can come into contact with the genomic DNA derived from the cell.
[0020] In a particular embodiment, the method is A step to identify a predetermined CRISPR / Cas genome editing system, wherein the step to identify the predetermined CRISPR / Cas genome editing system is performed before (i). You may want to prepare even more.
[0021] In one particular embodiment, the predetermined CRISPR / Cas genome editing system includes the use of a predetermined guide RNA having a predetermined guide sequence, wherein the predetermined guide sequence and the guide sequence of the guide RNA may be the same.
[0022] In a particular embodiment, the predetermined CRISPR / Cas genome editing system herein includes the use of predetermined cells, and the predetermined cells and the cells may be the same.
[0023] In a particular embodiment, the composition to be analyzed may include cleaved genomic DNA, in which the genomic DNA derived from physically destroyed cells is cleaved by the Cas / gRNA complex.
[0024] In a particular embodiment, the concentration of the Cas protein contained in the initial composition may be 4000 nM or more and 6000 nM or less.
[0025] In a particular embodiment, the concentration of the guide RNA contained in the initiation composition may be 4000 nM or more and 6000 nM or less.
[0026] In a particular embodiment, the concentration of the Cas / gRNA complex contained in the initial composition may be 4000 nM or more and 6000 nM or less.
[0027] In a particular embodiment, the concentration of the cells contained in the initial composition is 1 x 10 7 The value may be cells / mL.
[0028] In a particular embodiment, the step of obtaining the composition to be analyzed is: The process may further include incubating the composition obtained by destroying the cells.
[0029] In a particular embodiment, the step of obtaining the composition to be analyzed is: The process may further include the step of removing RNA from the composition obtained by destroying the aforementioned cells.
[0030] In a particular embodiment, the step of obtaining the composition to be analyzed may further include the step of purifying DNA from the composition obtained through the destruction of the cells.
[0031] One embodiment of the present invention is a method for identifying off-target information that occurs in a genome editing process using a CRISPR / Cas genome editing system, (i) A step of filling a first vessel of an extruder with an initiation composition comprising Cas protein, guide RNA, and cells; (ii) Using the extruder, (a) A step of applying pressure to the first container to move the components of the starting composition from the first container of the extruder to the second container of the extruder, Here, the components of the initial composition pass through a filter having holes located between the first and second containers of the extruder due to the applied pressure, and the mixture settles into the second container. Here, the cells, which are components larger in size than the diameter of the pores of the filter, are destroyed by the applied pressure and pass through the pores of the filter. Here, by physically destroying the cells, an environment is created in which the genomic DNA can come into contact with the Cas protein and guide RNA. As a result, the genomic DNA comes into contact with the Cas / gRNA complex, This brings the genomic DNA into contact with the Cas / gRNA complex, As a result, the genomic DNA is cleaved at one or more cleavage sites. A step of performing an extrusion process including the steps of obtaining the composition to be analyzed; and (iii) The step of analyzing the composition to be analyzed and obtaining information about the cleavage site. The present invention provides a method that includes the following:
[0032] In a particular embodiment, the pressure applied to the first vessel is generated through a process of pushing a piston designed to apply the pressure to the first vessel in a direction toward the first vessel and the filter.
[0033] One embodiment of the present invention is a method for identifying off-target information that occurs in a genome editing process using a CRISPR / Cas genome editing system, (i) A step of filling a first vessel of an extruder with an initiation composition comprising Cas protein, guide RNA, and cells; (ii) Using the extruder, (a) A step of applying pressure to the first container to move the components of the starting composition from the first container of the extruder to the second container of the extruder, Here, the components of the initial composition pass through a filter having holes located between the first and second containers of the extruder due to the applied pressure, and the mixture is contained in the second container. (b) A step of applying pressure to the second container to move the components of the mixture contained in the second container from the second container to the first container, Here, the components of the mixture contained in the second container pass through the filter having holes located between the first and second containers due to the applied pressure, and consequently move from the second container to the first container, thereby the mixture that has moved from the second container by passing through the filter due to the pressure settles into the first container, and (c) The step of repeating (a) and (b) a predetermined number of times, Here, the predetermined number of times is counted in increments of 0.5, where 0.5 represents the execution of a single process of (a) or (b), Here, the cells, which are components larger in size than the diameter of the pores of the filter, are destroyed by the applied pressure and pass through the pores of the filter. Here, by physically destroying the cells, an environment is created in which the genomic DNA can come into contact with the Cas protein and the guide RNA. As a result, the genomic DNA comes into contact with the Cas / gRNA complex, As a result, the genomic DNA is cleaved at one or more cleavage sites. A step of performing an extrusion process including the steps of obtaining the composition to be analyzed; and (iii) The step of analyzing the composition to be analyzed and obtaining information about the cleavage site. The present invention provides a method that includes the following:
[0034] In a particular embodiment, the pressure applied to the first vessel is generated through a process of pushing a piston designed to apply pressure to the first vessel in a direction toward the first vessel and the filter, and the pressure applied to the second vessel is generated through a process of pushing a piston designed to apply pressure to the second vessel in a direction toward the second vessel and the filter. [Advantageous effect]
[0035] This application provides a method for predicting off-target events that may occur during gene (e.g., genome) editing processes. This application provides a method for identifying candidate off-target events that may occur during genome editing processes. This application provides a method for predicting off-target events that can be performed more simply. The off-target prediction method of this application has the advantages of both in vitro-based and cell-based off-target prediction methods. The off-target prediction method of this application exhibits a lower false-positive rate. The off-target prediction method of this application exhibits a lower missed rate. In other words, when the off-target prediction method of this application is used, off-target events that may occur during genome editing processes can be predicted easily and accurately. [Brief explanation of the drawing]
[0036] [Figure 1] This document presents three categories of off-target prediction methods: cell-based, in vitro, and in silico.
[0037] [Figure 2] This is a schematic diagram illustrating a method for predicting off-target behavior, provided by one embodiment of the present application.
[0038] [Figure 3] This paper presents comparative results regarding off-target candidates predicted through different off-target prediction methods (Digenome-seq, Extru-seq, GUIDE-seq, and in silico). Comparative experiments on off-target prediction systems were conducted using sgRNAs targeting human PCSK9 and human albumin, respectively.
[0039] [Figure 4] This paper presents comparative results regarding off-target candidates predicted through different off-target prediction methods (Digenome-seq, Extru-seq, GUIDE-seq, and in silico). Comparative experiments on off-target prediction systems were conducted using sgRNAs targeting mouse PCSK9 and mouse albumin, respectively.
[0040] [Figure 5] The validation rates for the top off-target sites predicted by in silico, GUIDE-seq, Digenome-seq, and Extru-seq are shown. Results are presented for sgRNAs targeting human PCSK9, human albumin-targeting sgRNAs, mouse PCSK9-targeting sgRNAs, and mouse albumin-targeting sgRNAs.
[0041] [Figure 6] This paper presents comparative results regarding off-target candidates predicted through different off-target prediction methods (Digenome-seq, Extru-seq, GUIDE-seq, and DIG-seq). Comparative experiments on off-target prediction systems were conducted using sgRNAs targeting FANCF and sgRNAs targeting VEGFA, respectively.
[0042] [Figure 7] This paper presents comparative results regarding off-target candidates predicted through different off-target prediction methods (Digenome-seq, Extru-seq, GUIDE-seq, and DIG-seq). Comparative experiments on off-target prediction systems were conducted using sgRNAs targeting HBBs.
[0043] [Figure 8] This report shows the efficacy confirmation rates for top off-target sites predicted by DIG-seq, GUIDE-seq, Digenome-seq, and Extru-seq. Results are presented for sgRNAs targeting FANCF, VEGFA, and HBB.
[0044] [Figure 9] The comparison results for different off-target prediction methods are shown, analyzed through the common areas of the Venn diagrams (Figures 3 and 4, and Figures 6 and 7).
[0045] [Figure 10] This figure shows a comparison of efficacy confirmation results and off-target results predicted by GUIDE-seq and Extru-seq. Figure 10(a) shows results related to sgRNAs targeting human PCSK9. Figure 10(b) shows results related to human albumin-targeting sgRNAs.
[0046] [Figure 11] This figure shows a comparison of efficacy confirmation results and off-target results predicted by GUIDE-seq and Extru-seq. Figure 11(c) shows results related to sgRNAs targeting mouse PCSK9. Figure 11(d) shows results related to sgRNAs targeting mouse albumin.
[0047] [Figure 12] This figure shows a comparison of efficacy confirmation results and off-target results predicted by GUIDE-seq and Extru-seq. Figure 12(e) shows results related to sgRNAs targeting human FANCF. Figure 12(f) shows results related to sgRNAs targeting human VEGFA.
[0048] [Figure 13] This shows a comparison of efficacy confirmation results and off-target results predicted by GUIDE-seq and Extru-seq. Figure 13(g) shows results related to sgRNA targeting human HBB.
[0049] [Figure 14] This shows the missed rate for GUIDE-seq and Extru-seq off-target prediction methods, calculated based on the effectiveness confirmation results.
[0050] [Figure 15] This shows the distribution of the number of off-target mismatches missed by GUIDE-seq, confirmed based on efficacy verification results.
[0051] [Figure 16] The ROC curves for each off-target prediction method are shown. Figure 16(a) shows the results related to sgRNAs targeting human PCSK9. Figure 16(b) shows the results related to human albumin-targeting sgRNAs.
[0052] [Figure 17] The ROC curves for each off-target prediction method are shown. Figure 17(c) shows the results related to sgRNAs targeting mouse PCSK9. Figure 17(d) shows the results related to sgRNAs targeting mouse albumin.
[0053] [Figure 18] Figure 18(e) shows the results related to sgRNAs targeting human FANCF. Figure 18(f) shows the results related to sgRNAs targeting human VEGFA. Figure 18(g) shows the results related to sgRNAs targeting human HBB.
[0054] [Figure 19] Figures 16 to 18 show the AUC calculated using the ROC curve data. AUC was calculated for GUIDE-seq, Digenome-seq, Extru-seq, CROP, CFD, and DIG-seq, respectively.
[0055] [Figure 20] The results and experimental conditions of experiments performed to determine the optimal conditions for the average pore size of the filter, the Cas9 RNP concentration of the mixture, and the number of cells in extru-seq are shown. [Figure 21] The results and experimental conditions of experiments performed to determine the optimal conditions for the average pore size of the filter, the Cas9 RNP concentration of the mixture, and the number of cells in extru-seq are shown.
[0056] [Figure 22] Figure 22 shows the cleavage rates of on- and off-target sites recognized by sgRNA targeting the human PCSK9 site, as measured by quantitative PCR (qPCR). The results obtained via extru-seq are shown in Figure 22. [Figure 23] This figure shows the cleavage rates of on- and off-target sites recognized by sgRNA targeting the human PCSK9 site, as measured by quantitative PCR (qPCR). Figure 23 shows the results obtained via extru-seq.
[0057] [Figure 24] This shows WGS data from extru-seq analyzed using IGV to identify the cleavage pattern. [Figure 25] This shows WGS data from extru-seq analyzed using IGV to identify the cleavage pattern. [Figure 26] This shows WGS data from extru-seq analyzed using IGV to identify the cleavage pattern. [Figure 27] This shows WGS data from extru-seq analyzed using IGV to identify the cleavage pattern. [Figure 28] This shows WGS data from extru-seq analyzed using IGV to identify the cleavage pattern. [Figure 29] This shows WGS data from extru-seq analyzed using IGV to identify the cleavage pattern. [Figure 30] This shows WGS data from extru-seq analyzed using IGV to identify the cleavage pattern.
[0058] [Figure 31] This shows the cleavage rates of seven on-target sites for each target, obtained through qPCR and manual calculations based on IGV analysis of WGS data.
[0059] [Figure 32] This shows the results of dip sequencing performed on the non-cleaved group to confirm the degree of NHEJ formation after the extrusion process of extru-seq.
[0060] [Figure 33] The results regarding the cracking rate after SCR7 treatment, performed to confirm the degree of NHEJ formation after the extru-seq extrusion process, are shown.
[0061] [Figure 34]The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 34 shows the sequence read results for GUIDE-seq obtained from HEK293T using sgRNA targeting PCSK9. [Figure 35] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 35 shows the sequence read results for GUIDE-seq obtained from HEK293T using sgRNA targeting PCSK9. [Figure 36] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 36 shows the sequence read results for GUIDE-seq obtained from HEK293T using albumin-targeting sgRNA. [Figure 37] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 37 shows the sequence read results for GUIDE-seq obtained from HEK293T using albumin-targeting sgRNA. [Figure 38] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 38 shows the sequence read results for GUIDE-seq obtained from NIH-3T3 using sgRNA targeting PCSK9. [Figure 39] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 39 shows the sequence read results for GUIDE-seq obtained from NIH-3T3 using sgRNA targeting PCSK9. [Figure 40] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 40 shows the sequence read results for GUIDE-seq obtained from NIH-3T3 using albumin-targeting sgRNA. [Figure 41] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 41 shows the sequence read results for GUIDE-seq obtained from NIH-3T3 using albumin-targeting sgRNA.
[0062] [Figure 42] The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 42 shows the Manhattan plot results for Digenome-seq obtained from HEK293T using sgRNA targeting PCSK9. [Figure 43] The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 43 shows the Manhattan plot results for Digenome-seq obtained from HEK293T using sgRNA targeting PCSK9. [Figure 44] The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 44 shows the Manhattan plot results for Digenome-seq obtained from HEK293T using albumin-targeting sgRNA. [Figure 45] The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 45 shows the Manhattan plot results for Digenome-seq obtained from HEK293T using albumin-targeting sgRNA. [Figure 46]The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 46 shows the Manhattan plot results for Digenome-seq obtained from NIH-3T3 using sgRNA targeting PCSK9. [Figure 47] The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 47 shows the Manhattan plot results for Digenome-seq obtained from NIH-3T3 using sgRNA targeting PCSK9. [Figure 48] The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 48 shows the Manhattan plot results for Digenome-seq obtained from NIH-3T3 using albumin-targeting sgRNA. [Figure 49] The Manhattan plot results for off-target candidates predicted via Digenome-seq are shown. The Y-axis represents the DNA cleavage score. Figure 49 shows the Manhattan plot results for Digenome-seq obtained from NIH-3T3 using albumin-targeting sgRNA.
[0063] [Figure 50] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 50 shows the Manhattan plot results for extru-seq obtained from HEK293T using sgRNA targeting PCSK9. [Figure 51]The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 51 shows the Manhattan plot results for extru-seq obtained from HEK293T using sgRNA targeting PCSK9. [Figure 52] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 52 shows the Manhattan plot results for extru-seq obtained from HEK293T using albumin-targeting sgRNA. [Figure 53] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 53 shows the Manhattan plot results for extru-seq obtained from HEK293T using albumin-targeting sgRNA. [Figure 54] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 54 shows the Manhattan plot results for extru-seq obtained from NIH-3T3 using sgRNA targeting PCSK9. [Figure 55] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 55 shows the Manhattan plot results for extru-seq obtained from NIH-3T3 using sgRNA targeting PCSK9. [Figure 56] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 56 shows the Manhattan plot results for extru-seq obtained from NIH-3T3 using albumin-targeting sgRNA. [Figure 57]The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 57 shows the Manhattan plot results for extru-seq obtained from NIH-3T3 using albumin-targeting sgRNA.
[0064] [Figure 58] The results shown relate to a score based on the number of mismatches between on-target and off-target sequences, predicted using GUIDE-seq (a score calculated from sequence read count results).
[0065] [Figure 59] The results shown are related to a score (breakdown score in the Manhattan plot) based on the number of mismatches between on-target and off-target targets predicted using Digenome-seq.
[0066] [Figure 60] The results shown relate to a score (CROP score) based on the number of mismatches between on-target and off-target situations predicted using an in silico system.
[0067] [Figure 61] The results shown relate to a score (CFD score) based on the number of mismatches between on-target and off-target situations predicted using an in silico system.
[0068] [Figure 62] The results shown relate to a score (the split score in the Manhattan plot) based on the number of mismatches between on-target and off-target targets predicted using extru-seq.
[0069] [Figure 63] The results regarding the frequency of indel formation due to subretinal injection and systemic injection are shown. [Figure 64] The results regarding the frequency of indel formation due to subretinal injection and systemic injection are shown.
[0070] [Figure 65] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 65 shows the sequence read results for GUIDE-seq obtained from HeLa cells using sgRNA targeting FANCF. [Figure 66] The following shows the sequence read results for off-target candidates predicted via GUIDE-seq. Figure 66 shows the sequence read results for GUIDE-seq obtained from HeLa cells using sgRNA targeting VEGFA. Figure 67 shows the sequence read results for GUIDE-seq obtained from HeLa cells using sgRNA targeting HBB. [Figure 67] The sequence read results for off-target candidates predicted via GUIDE-seq are shown. Figure 67 shows the sequence read results for GUIDE-seq obtained from HeLa cells using sgRNA targeting HBBs.
[0071] [Figure 68] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 68 shows the Manhattan plot results for extru-seq obtained from HeLa cells using sgRNA targeting FANCF. [Figure 69] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 69 shows the Manhattan plot results for extru-seq obtained from HeLa cells using sgRNA targeting FANCF. [Figure 70]The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 70 shows the Manhattan plot results for extru-seq obtained from HeLa cells using sgRNA targeting VEGFA. [Figure 71] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 71 shows the Manhattan plot results for extru-seq obtained from HeLa cells using sgRNA targeting VEGFA. [Figure 72] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 72 shows the Manhattan plot results for extru-seq obtained from HeLa cells using HBB-targeting sgRNA. [Figure 73] The Manhattan plot results for off-target candidates predicted via extru-seq are shown. The Y-axis represents the DNA cleavage score. Figure 73 shows the Manhattan plot results for extru-seq obtained from HeLa cells using HBB-targeting sgRNA.
[0072] [Figure 74] This figure shows a Venn diagram comparing extru-seq results obtained from MSCs and extru-seq results obtained from HEK293T cells. Figure 74 shows results related to sgRNAs targeting human PCSK9. [Figure 75] A Venn diagram comparing extru-seq results obtained from MSCs and extru-seq results obtained from HEK293T cells is shown. Figure 75 shows results related to sgRNAs targeting human albumin.
[0073] [Figure 76]The p-values obtained by normalized rank-sum tests for each pair of off-target prediction methods for PCSK9 and albumin-targeting sgRNAs in MSCs and HEK293T cells are shown.
[0074] [Figure 77] The results for off-target sites, manually validated using IGV, are shown. [Figure 78] The results for off-target sites, manually validated using IGV, are shown. [Figure 79] The results for off-target sites, manually validated using IGV, are shown. [Figure 80] The results for off-target sites, manually validated using IGV, are shown. [Figure 81] The results for off-target sites, manually validated using IGV, are shown. [Figure 82] The results for off-target sites, manually validated using IGV, are shown. [Figure 83] The results for off-target sites, manually validated using IGV, are shown. [Figure 84] The results for off-target sites, manually validated using IGV, are shown. [Figure 85] The results for off-target sites, manually validated using IGV, are shown. [Figure 86] The results for off-target sites, manually validated using IGV, are shown. [Figure 87] The results for off-target sites, manually validated using IGV, are shown. [Figure 88] The results for off-target sites, manually validated using IGV, are shown. [Figure 89]The results for off-target sites, manually validated using IGV, are shown. [Figure 90] The results for off-target sites, manually validated using IGV, are shown. [Figure 91] The results for off-target sites, manually validated using IGV, are shown. [Figure 92] The results for off-target sites, manually validated using IGV, are shown. [Figure 93] The results for off-target sites, manually validated using IGV, are shown. [Figure 94] The results for off-target sites, manually validated using IGV, are shown. [Figure 95] The results for off-target sites, manually validated using IGV, are shown. [Figure 96] The results for off-target sites, manually validated using IGV, are shown. [Figure 97] The results for off-target sites, manually validated using IGV, are shown. [Figure 98] The results for off-target sites, manually validated using IGV, are shown. [Figure 99] The results for off-target sites, manually validated using IGV, are shown. [Figure 100] The results for off-target sites, manually validated using IGV, are shown. [Figure 101] The results for off-target sites, manually validated using IGV, are shown. [Figure 102] The results for off-target sites, manually validated using IGV, are shown. [Figure 103] The results for off-target sites, manually validated using IGV, are shown. [Figure 104] The results for off-target sites, manually validated using IGV, are shown. [Figure 105] The results for off-target sites, manually validated using IGV, are shown. [Figure 106] The results for off-target sites, manually validated using IGV, are shown. [Figure 107] The results for off-target sites, manually validated using IGV, are shown. [Figure 108] The results for off-target sites, manually validated using IGV, are shown. [Figure 109] The results for off-target sites, manually validated using IGV, are shown. [Figure 110] The results for off-target sites, manually validated using IGV, are shown. [Figure 111] The results for off-target sites, manually validated using IGV, are shown. [Figure 112] The results for off-target sites, manually validated using IGV, are shown. [Figure 113] The results for off-target sites, manually validated using IGV, are shown. [Figure 114] The results for off-target sites, manually validated using IGV, are shown. [Figure 115] The results for off-target sites, manually validated using IGV, are shown. [Figure 116] The results for off-target sites, manually validated using IGV, are shown.
[0075] [Figure 117] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 118] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 119] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 120] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 121] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 122] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 123] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 124] The results for false-positive off-target candidates manually excluded using IGV are shown from Digenome-seq and Extru-seq WGS data. [Figure 125] The results regarding false-positive off-target candidates manually excluded using IGV from Digenome-seq and Extru-seq WGS data are shown. [Aspects of the Invention]
[0076] [Definition of Terms] nucleic acid As used herein, the term “nucleic acid” means a subregion or the entire molecule of a molecule consisting of DNA (double-stranded or single-stranded), RNA (double-stranded or single-stranded), or a hybrid of DNA and RNA (double-stranded or single-stranded). “Nucleic acid” is used to mean, but is not limited to, a collection of nucleotides (a subregion or the entire molecule). The terms “nucleic acid” or “nucleic acid region” may be used to refer to a subregion of a molecule. The terms “nucleic acid” or “nucleic acid area” may be used to refer to the entire molecule. The term “nucleic acid” should be interpreted appropriately depending on the context, and the content of each context, including the explanation of the term “nucleic acid,” will be helpful in understanding the meaning of the term. Furthermore, the above terms include all meanings recognized by those skilled in the art and can be interpreted appropriately depending on the context.
[0077] "Linked" or "Linked" As used herein, the terms “linked” or “connected” refer to two or more elements present in a single conceptualizable structure that are directly or indirectly connected (e.g., via different elements such as linkers), and it is not intended that other additional elements cannot exist between these two or more elements. For example, a description such as “element B linked to element A” is intended to include, and should not be interpreted as limiting, both cases where one or more other elements interpose between elements A and B (i.e., element A is connected to element B via one or more third elements) and cases where no other elements interpose between elements A and B (i.e., elements A and B are directly connected).
[0078] Sequence identity As used herein, the term “sequence identity” is used in relation to the similarity between two or more nucleotide sequences. For example, the term “sequence identity” is used in conjunction with terms referring to a reference sequence and a ratio (e.g., a percentage). For example, the term “sequence identity” may be used to describe a sequence that is similar to or substantially identical to a reference nucleotide. When it is stated that “a sequence has 90% or more sequence identity with sequence A,” the reference sequence used is sequence A. For example, a percentage of sequence identity can be calculated by aligning the sequence to which the percentage of sequence identity has been measured with the reference sequence, and the percentage of sequence identity may be calculated by including all mismatches, deletions, and insertions in one or more nucleotides. The method for calculating and / or determining the percentage of sequence identity is not otherwise limited and may be calculated and / or determined through a reasonable method or algorithm that can be used by those skilled in the art.
[0079] Representation of amino acid sequences Unless otherwise specified, amino acid sequences in this specification are written in the direction from the N-terminus to the C-terminus, using either single-letter or three-letter notation for amino acids. For example, RNVP represents a peptide in which arginine, asparagine, valine, and proline are joined in order from the N-terminus to the C-terminus. In another example, Thr-Leu-Lys represents a peptide in which threonine, leucine, and lysine are joined in order from the N-terminus to the C-terminus. Amino acids that cannot be represented by single-letter notation are represented using multiple different letters and are explained in detail.
[0080] The notation for each amino acid is as follows: alanine (Ala, A); arginine (Arg, R); asparagine (Asn, N); aspartic acid (Asp, D); cysteine (Cys, C); glutamic acid (Glu, E); glutamine (Gln, Q); glycine (Gly, G); histidine (His, H); isoleucine (Ile, I); leucine (Leu, L); lysine (Lys, K); methionine (Met, M); phenylalanine (Phe, F); proline (Pro, P); serine (Ser, S); threonine (Thr, T); tryptophan (Trp, W); tyrosine (Tyr, Y); and valine (Val, V).
[0081] Representation of nucleic acid sequences The symbols A, T, C, G, and U used herein are to be interpreted as understood by those skilled in the art. Depending on the context and the art, they may be interpreted as bases, nucleosides, or nucleotides on DNA or RNA. For example, when referring to bases, each symbol may be interpreted as adenine (A), thymine (T), cytosine (C), guanine (G), or uracil (U); when referring to nucleosides, each symbol may be interpreted as adenosine (A), thymidine (T), cytidine (C), guanosine (G), or uridine (U); and when referring to nucleotides in a sequence, each symbol should be interpreted as indicating a nucleotide containing each nucleoside.
[0082] target array As used herein, “target sequence” refers to a specific sequence recognized by a guide RNA or gene editing tool (e.g., a Cas / gRNA complex) to cleave a target gene or target nucleic acid. The target sequence may be appropriately selected depending on the purpose. For example, “target sequence” may refer to a sequence contained in a target gene or target nucleic acid sequence, and a sequence complementary to a spacer sequence contained in the guide RNA. In another example, “target sequence” may refer to a sequence complementary to a sequence contained in a target gene or target nucleic acid sequence, and a sequence complementary to a spacer sequence contained in the guide RNA. Therefore, “target sequence” is used to refer to, and should not be interpreted as limiting, a sequence complementary to a spacer sequence contained in the guide RNA and / or a sequence substantially identical to the spacer sequence of the guide RNA. In some embodiments, the target sequence may be disclosed as a sequence containing a PAM sequence. In some embodiments, the target sequence may be disclosed as a sequence not containing a PAM sequence. The term “target sequence” should be interpreted appropriately depending on the context. Generally, spacer sequences are determined considering the sequence of the target gene or target nucleic acid and the PAM sequence recognized by the editing protein of the CRISPR / Cas system. The target sequence may refer only to the sequence of a specific strand that binds complementaryly to the guide RNA of the CRISPR / Cas complex, only to the sequence of a specific strand that does not bind complementaryly to the guide RNA, or to the entire target double helix containing a specific strand portion, which is interpreted appropriately depending on the context. The definitions of the term "target sequence" are provided to describe the strand on which the target sequence may reside, and it is not intended to distinguish between on-target and off-target sequences through the term "target sequence." That is, in some embodiments, an intended target sequence may be referred to as an on-target sequence, and an unintended target sequence may be referred to as an off-target sequence. With respect to on-target and off-target, the term "target sequence" may be interpreted appropriately depending on the context of the relevant paragraph.
[0083] Spacer linkage chain As used herein, the term "spacer-binding strand" refers to a strand containing a sequence complementary to a portion or the entire sequence of the spacer region of a guide nucleic acid (e.g., guide RNA) in a gene editing system (e.g., CRISPR / Cas gene editing system). DNA molecules such as genomes generally have a double-stranded structure. In a double-stranded structure, a strand that has a sequence complementary to a portion or the entire sequence of the spacer region of the guide nucleic acid, and thus forms a complementary bond with that portion or the entire sequence, may be called a spacer-binding strand.
[0084] Spacer unbound chain As used herein, the term "spacer-unbound strand" refers to a strand other than the "spacer-bound strand," which is a strand containing a sequence that forms a complementary bond with a portion or the entire spacer region of a guide nucleic acid (e.g., guide RNA) in a gene editing system (e.g., CRISPR / Cas gene editing system) involved with the guide nucleic acid. DNA molecules such as genomes generally have a double-stranded structure, and the term "spacer-unbound strand" may be used to refer to a strand other than the spacer-bound strand in a double-stranded structure.
[0085] functional equivalent The terms “functional equivalent” or “equivalent” refer to a second biomolecule that is functionally equivalent to a first biomolecule, but not necessarily structurally equivalent. For example, a “Cas9 equivalent” refers to a protein that has the same or substantially the same function as Cas9, but not necessarily the same amino acid sequence. Throughout this application, when a particular protein is referred to, it is intended that the particular protein referred to above encompasses all of its functional equivalents. For example, when described as “protein X,” the term “protein X” may be interpreted to encompass all functional equivalents of protein X. In this sense, a “functional equivalent” of protein X includes any homolog, paralog, ortholog, fragment, native, artificial, mutant, or synthetic version of protein X that ensures equivalent function. When described as a Cas protein, the term “Cas protein” may be interpreted to encompass all functional equivalents of the Cas protein.
[0086] Nuclear localization signal or sequence (NLS) The term “nuclear localization signal or sequence (NLS)” refers to an amino acid sequence that facilitates the delivery of a protein into the cell nucleus. For example, protein delivery is achieved by nuclear transport. NLSs are well known in the art and will be obvious to those skilled in the art. For example, exemplary sequences of NLSs may be disclosed in PCT application PCT / EP2000 / 011690 (publication number WO2021 / 038547), the contents of which relating to exemplary NLSs are incorporated herein by reference. In some embodiments, NLS is PKKKRKV (SEQ ID NO: 01), KRPAATKKAGQAKKKK (SEQ ID NO: 02), PAAKRVKLD (SEQ ID NO: 03), RQRRNELKRSP (SEQ ID NO: 04), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 05), RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 06), VSRKRPRP (SEQ ID NO: 07), PPKKARED (SEQ ID NO: 08), PO The present invention may include, but is not limited to, amino acid sequences such as PKKKPL (SEQ ID NO: 09), SALIKKKKKMAP (SEQ ID NO: 10), DRLRR (SEQ ID NO: 11), PKQKKRK (SEQ ID NO: 12), RKLKKKIKKL (SEQ ID NO: 13), REKKKFLKRR (SEQ ID NO: 14), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 15), RKCLQAGMNLEARKTKK (SEQ ID NO: 16), or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 17). NLS can be selectively fused to gene editing agents such as Cas proteins. NLS fused to a protein can be used to facilitate the movement of the fused protein to a desired location in the nucleus.
[0087] about As used herein, the term “about” means an approximation to a particular quantity, referring to a quantity, level, value, number, frequency, percentage, dimension, size, volume, weight, or length that has changed by approximately 30, 25, 20, 25, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1% from a reference quantity, level, value, number, frequency, percentage, dimension, size, volume, weight, or length.
[0088] Directionality of the disclosed sequence Nucleotide sequences disclosed herein (e.g., DNA sequences, RNA sequences, or DNA / RNA hybrid sequences) should be understood to be disclosed in the 5'→3' direction unless otherwise specified. Amino acid sequences disclosed herein should be understood to be disclosed in the N-terminus to C-terminus direction unless otherwise specified. For sequences disclosed in a direction other than those described above, the other direction is specified separately in the paragraph relating to the corresponding sequence.
[0089] Overview of gene editing systems This application relates to a method for predicting off-target effects that may occur in a gene editing process using a gene editing system. Off-target prediction is used to encompass the prediction of off-target sites. Before describing the method for predicting off-target effects provided by this application, gene editing systems related to off-target effects will be described. A gene editing system (e.g., a genome editing system) refers to a system used to achieve desired editing in a desired nucleic acid molecule (e.g., genomic DNA) through the use of gene editing tools such as editing proteins and guide nucleic acids. In many studies, gene editing systems are used to edit the genome of a cell, and the term “gene editing system” may be used interchangeably with genome editing system. However, the use of gene editing systems is not limited to genome editing. Furthermore, the term “gene editing system” may be used to refer to gene editing tools and may be interpreted as appropriate depending on the relevant context. Known gene editing systems include zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and the CRISPR / Cas gene editing system (see the following reference, which is incorporated herein by reference: Khan, Sikandar Hayat. "Genome-editing technologies: concept, pros, and cons of various genome-editing techniques and bioethical concerns for clinical application." Molecular Therapy-Nucleic Acids 16(2019):326-334.). Furthermore, there are base editing and prime editing techniques developed based on the CRISPR / Cas gene editing system.
[0090] One characteristic of the off-target prediction method provided by this application is that it involves cleaving the cell membrane structure through a physical method (e.g., using an extruder) to bring the components of the gene editing system (e.g., editing protein and / or guide nucleic acid) into contact with the genome. Therefore, the off-target prediction method of this application can be applied to all of the gene editing systems described above.
[0091] As an example of a gene editing system, the CRISPR / Cas gene editing system, which is being actively researched to achieve the goals of genome editing, is described in detail below.
[0092] CRISPR / Cas gene editing system Overview of the CRISPR / Cas gene editing system The term CRISPR / Cas gene editing system is used as a general term to refer to gene editing systems that involve editing proteins containing Cas proteins and guide nucleic acids (e.g., guide RNA) used to induce desired editing of a gene (e.g., genomic DNA) at a desired location. CRISPR / Cas gene system may be used as other terms that those skilled in the art will understand. For example, a CRISPR / Cas gene system may be referred to as CRISPR / Cas, a CRISPR / Cas system, a CRISRP system, or a Cas-based genome editing system, but this application is not limited thereto. Furthermore, the CRISPR / Cas gene editing system is used to encompass all developmental technologies developed based on the CRISPR / Cas gene, including base editing (see reference [Gaudelli, Nicole M., et al. "Programmable base editing of A· T to G· C in genomic DNA without DNA cleavage." Nature 551.7681(2017):464-471.]) and prime editing (see reference [Anzalone, Andrew V., et al. "Search-and-replace genome editing without double-strand breaks or donor DNA." Nature 576.7785(2019):149-157.]). The results of gene editing (e.g., genome editing) may include, but are not limited to, cleavage, indels, insertions, deletions, substitutions, base editing (e.g., which can be achieved by base editing), and writing (e.g., which can be achieved by prime editing). The following is a detailed description of the CRISPR / Cas gene editing system, including its origins.
[0093] CRISPR The "CRISPR" section is provided for the understanding of those skilled in the art, and the terms used in this section are not intended to limit the terms disclosed herein.
[0094] CRISPR is a family of DNA sequences (i.e., CRISPR clusters) in bacteria and archaea that represent snippets of previous infections brought about by viruses that have invaded prokaryotes. These DNA snippets are used by prokaryotic cells to detect and destroy DNA from subsequent attacks by similar viruses, and together with arrays of CRISPR-associated proteins (Cas proteins) and CRISPR-associated RNAs, they effectively constitute the prokaryotic immune defense system. CRISPR clusters are transcribed and processed into CRISPR RNA (crRNA). Subsequently, Cas9 / crRNA / tracrRNA cleaves linear or circular dsDNA targets complementary to the RNA in an endonuclease-like manner. In particular, target strands not complementary to the crRNA are first cleaved endonuclease-like and then trimmed 3'→5' exonuclease-like. DNA binding and cleavage typically require a protein and two RNAs. However, single guide RNAs (sgRNAs, or simply gRNAs) have been developed, and single-stranded RNAs have been engineered to incorporate both crRNA and tracrRNA characteristics into a single RNA species. For example, see, for reference, the full text of which is incorporated herein by reference [Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Science 337.6096(2012):816-821.]. Cas9 recognizes short motifs (PAMs or protospacer-adjacent motifs) in CRISPR repeat sequences to help distinguish between self and non-self.CRISPR biology, as well as Cas9 nuclease sequences and structures, are well known to those skilled in the art (see, for example, the literature whose entire contents are incorporated herein by reference: Ferretti, Joseph J., et al. "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Proceedings of the National Academy of Sciences 98.8(2001):4658-4663.; Deltcheva, Elitza, et al. "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Nature 471.7340(2011):602-607.; and Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Science 337.6096(2012):816-821.). Cas9 orthologs have been described in various species, including Streptococcus pyogenes (S. pyogenes) and Streptococcus thermophilus (S. thermophilus), but are not limited thereto. Additional preferred Cas9 nucleases and their sequences will become apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and their sequences include Cas9 sequences from organisms and loci disclosed in the literature [Chylinski, Krzysztof, Anais Le Rhun, and Emmanuelle Charpentier. The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems. RNA biology 10.5(2013):726-737.], the entirety of which is incorporated herein by reference.
[0095] CRISPR / Cas gene editing system The CRISPR / Cas gene editing system, developed from the aforementioned CRISPR, is a technology that edits genes (e.g., the cellular genome) at a desired location using Cas proteins derived from the cell's CRISPR system and guide nucleic acids that guide the Cas proteins to a target region. For example, Cas proteins form a Cas / gRNA complex with guide RNA (gRNA). The Cas / gRNA complex guides the guide RNA contained in the complex to the desired location. The Cas proteins contained in the Cas / gRNA complex induce double-strand breaks (DSBs) or nicks at the desired location. The CRISPR / Cas gene editing system can edit not only the cellular genome but also DNA molecules that are not located on the genome. Since the discovery of CRISPR, as mentioned above, single guide RNAs to which tracrRNA and crRNA are attached have been developed for the CRISPR / Cas genome editing system (see the reference cited herein for its full content [Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." science 337.6096(2012):816-821.]), and various types of Cas proteins (Cas equivalents) such as cas9, cas12a(cpf1), cas12b(c2c1), cas12e(casX), cas12k(c2c5), cas14, cas14a, cas13a(c2c2), cas13b(c2c6), cas9 nickase, and dead cas have been developed. Furthermore, based on the CRISPR / Cas gene editing system, base editing capable of achieving the purpose of base modification and prime editing capable of writing the desired edit have been developed. As described above, the CRISPR / Cas gene editing system can be used to encompass conventional CRISPR / Cas gene editing systems and gene editing technologies developed based thereon.To understand the CRISPR / Cas gene editing system, one can refer to the document WO2018 / 231018 (PCT No.), the full contents of which are incorporated herein by reference. For the benefit of those skilled in the art, the editing proteins (e.g., Cas proteins) that can be used in the CRISPR / Cas gene editing system are further described below.
[0096] CRISPR / Cas gene editing system 1 - Editing protein (editor protein) Overview of edited proteins and edited proteins in the CRISPR / Cas gene editing system The term "editor protein" may be used to refer to a protein that generates a double-stroke spread (DSB) or nick in a desired region to achieve gene editing, or a protein that assists in inducing editing. Generally, proteins that have nuclease activity to cleave nucleic acids may be referred to as editor proteins. In CRISPR / Cas gene editing systems, editor proteins may be used interchangeably with Cas proteins. A representative example of a Cas protein is Cas9. As used herein, the term "Cas protein" is used as a general term for gene editing proteins used in CRISPR / Cas gene editing systems that can generate a DSB or nick in a target region or an inactive Cas protein. Examples of Cas proteins include, but are not limited to, Cas9, Cas9 variants, Cas9 nickase (nCas9), dead Cas9, Cpf1 (Type V CRISPR-Cas system), C2c1 (Type V CRISPR-Cas system), C2c2 (Type VI CRISPR-Cas system), and C2c3 (Type V CRISPR-Cas system). An example of the addition of Cas proteins is disclosed in its entirety in the following document, which is incorporated herein by reference: [Abudayyeh, Omar O., et al. "C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector." Science 353.6299(2016):aaf5573.]In one embodiment, the Cas protein is derived from Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Staphylococcus aureus, Campylobacter jejuni, Nocardiopsis dasonvillei, Streptomyces pristinea espiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporandium roseum, Streptosporandium roseum, and aricyclo Bacillus acidocardarius, Bacillus pseudomycoides, Bacillus serenityredusens, Exigobacterium sibiricum, Lactobacillus delbruckii, Lactobacillus salivarius, Microsilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorance, Polaromonas, Crocosphaera watsoni, Cyanoseis, Microcystis erginosa, Synechococcus, Acethalobium alabaticum, Ammonifec *Cardioscopus degensii*, *Cardioscopus caldicosylpterus besii*, *Candidatus desulfordis*, *Clostridium botulinum*, *Clostridium difficile*, *Finegorgia magna*, *Natranaerobius thermophilus*, *Pelotomacrum thermopropionicum*, *Acidiciobacillus cardus*, *Acidiciobacillus ferrooxydance*, *Allochromatium binosum*, *Marinobacter*, *Nitrosococcus halophyllus*, *Nitrosococcus watsoni*, *Pseudoalthea* Cas9 or Cpf1 may be derived from various microorganisms such as Romonas haloplanchtis, Ctedonobacter racemifer, Metanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc, Arthrospira maxima, Arthrospira pratensis, Arthrospira, Lingus, Microcoleus ctonoplastes, Osillatoria, Petrotoga mobilis, Thermosipho africanus, and Acariochloris marina. The Cas9 proteins used in the CRISPR / Cas9 gene editing system are shown below.
[0097] Cas9 protein In the CRISPR / Cas9 gene editing system, proteins that possess nuclease activity to cleave nucleic acids are called Cas9 proteins. Cas9 proteins correspond to Class 2 (Type II) in the classification of the CRISPR / Cas system and include Cas9 proteins derived from Streptococcus pyogenes, Streptococcus thermophilus, the genus Streptococcus, Streptomyces pristinea espiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporandium roseum, and Streptosporandium roseum. Additional Cas9 proteins and their sequences are disclosed in full in the literature [Chylinski, Krzysztof, Anais Le Rhun, and Emmanuelle Charpentier. "The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems." RNA biology 10.5(2013):726-737.], which is incorporated herein by reference. For example, the DNA cleavage domain of Cas9 is known to contain two subdomains, such as the NHN nuclease subdomain and the RucC1 subdomain. The NHN subdomain cleaves the strand complementary to the gRNA, while the RucC1 subdomain cleaves the strand that is not complementary. Inactivation of these subdomains can silence the nuclease activity of Cas9. For example, both mutants D10A and H840A completely inactivate the nuclease activity of S. pyogenes Cas9 (see reference [Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." science 337.6096(2012):816-821.]). For example, mutant H840A provides Cas9 nicasse.
[0098] CRISPR / Cas gene editing system 2 - guide nucleic acid Guide to nucleic acids: Overview In the CRISPR / Cas gene editing system, the Cas protein associates with a guide nucleic acid (e.g., guide RNA) to form a Cas / guide nucleic acid complex (e.g., Cas / gRNA complex). The Cas / gRNA complex may be referred to as a ribonucleoprotein (RNP). The Cas / gRNA complex generates a double-strand break (DSB) or nick in a target region containing a sequence corresponding to (e.g., complementary to) the spacer sequence of the guide RNA (gRNA), and the DSB or nick is induced by the Cas protein. The site where the DSB or nick is generated may be near a PAM sequence on the genome. Protospacer-adjacent motifs (PAMs) on the genome and spacer sequences of gRNA are involved in Cas / gRNA targeting. The Cas protein (e.g., Cas9), guided to the target region by the PAMs and gRNA spacer sequences, generates a DSB in the target region.
[0099] In the CRISPR / Cas gene editing system, RNA that has the function of guiding the Cas protein to the target region in order to recognize a specific sequence contained in the target DNA molecule is called guide RNA.
[0100] When the composition of guide RNA is divided functionally, it can be mainly divided into 1) a scaffold sequence portion and 2) a guide domain containing the guide sequence. The scaffold sequence portion is the portion that interacts with the Cas protein (e.g., Cas9 protein) and the portion that can bind to the Cas protein to form a complex. Generally, the scaffold sequence portion includes tracrRNA and crRNA repeat sequence portions, and the scaffold sequence is determined depending on which Cas9 protein is used. The guide sequence is the portion that can bind complementaryly to a nucleotide sequence portion of a specific length in the target nucleic acid (e.g., a target DNA molecule or cell genome). The guide sequence can be artificially modified and is determined by the target nucleotide sequence of interest related to the desired gene editing.
[0101] In some embodiments, the guide RNA may be described as comprising crRNA and tracrRNA. The crRNA may include spacer and repeat sequences. The repeat sequence portion of the crRNA may interact with (e.g., bind complementaryly to) a portion of the tracrRNA. As described above, a single guide RNA (sgRNA) to which the crRNA and tracrRNA are ligated may be provided.
[0102] In one embodiment, the guide RNA may be provided as two strands. In one embodiment, the guide RNA may be provided as a single strand. TracrRNA and sgRNA to which crRNA is ligated have been developed (see the reference [Jinek, Martin, et al. "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." science 337.6096(2012):816-821.], the full content of which is incorporated herein by reference). In a particular embodiment, the guide RNA may be sgRNA.
[0103] Guide domain and guide sequence of guide RNA The guide nucleic acid (e.g., guide RNA) may include a guide domain containing a guide sequence. The guide sequence is used interchangeably with a spacer sequence. The guide sequence is determined by the target nucleotide sequence of interest, as it is an artificially designed portion. In some embodiments, the guide sequence may be designed to target a sequence adjacent to the PAM sequence located on the DNA molecule for editing. As described above, localization of the Cas / gRNA complex to the target site is induced. The structure of the guide nucleic acid may vary depending on the CRISPR type. For example, the guide RNA used in the CRISPR / Cas9 gene editing system may have a 5'-[guide domain]-[scaffold]-3' structure.
[0104] In one embodiment, the guide sequence may have a length of 5 to 40 nt. In one embodiment, the guide sequence contained in the guide domain of the guide RNA may have a length of 10 to 30 nt. In one embodiment, the guide sequence may have a length of 15 to 25 nt. In one embodiment, the guide sequence may have a length of 18 to 22 nt. In one embodiment, the guide sequence may have a length of 20 nt. In one embodiment, the target sequence, which is a sequence in the genome that forms a complementary bond with the guide sequence (including both target sequences present on the spacer-binding strand and target sequences present on the spacer-non-binding strand), may have a length of 5 to 40 nt or 5 to 40 bp. In one embodiment, the target sequence, which is a sequence in the genome that forms a complementary bond with the guide sequence, may have a length of 10 to 30 nt or 10 to 30 bp. In one embodiment, the target sequence may have a length of 15 to 25 nt or 15 to 25 bp. In one embodiment, the target sequence may have a length of 18 to 22 nt or 18 to 22 bp. In one embodiment, the target sequence may have a length of 20 nt or 20 bp.
[0105] CRISPR / Cas gene editing system 3 - Protospacer adjacent motif (PAM): Two conditions are required for the CRISPR / Cas9 gene editing system to cleave a target gene or target nucleic acid.
[0106] Firstly, a specific length of base sequence (nucleotide sequence) that the Cas9 protein can recognize must be present within the target gene or target nucleic acid. Here, the specific length of base sequence (nucleotide sequence) recognized by the Cas9 protein is called the protospacer adjacent motif (PAM) sequence. The PAM sequence is a unique sequence determined by the Cas9 protein. Secondly, there must be a sequence that can bind complementaryly to the spacer sequence contained in the guide RNA surrounding the specific length of the PAM sequence. Here, the PAM sequence may be used to include sequences present on the spacer-unbound strand and sequences present on the spacer-bound strand.
[0107] As described above, in the CRISPR / Cas gene editing system, the Cas / gRNA complex is guided to the target region by the PAM sequence and the guide sequence of the gRNA on the target DNA molecule (e.g., the genome of a cell). On the target DNA molecule, the PAM sequence may be located on the unbound strand of the guide sequence, which is not the strand to which the guide sequence of the guide RNA binds. The PAM sequence may be determined independently depending on the type of Cas protein used. In one embodiment, the PAM sequence may be any one selected from NGG (SEQ ID NO: 18); NNNNRYAC (SEQ ID NO: 19); NNAGAAW (SEQ ID NO: 20); NNNNGATT (SEQ ID NO: 21); NNGRR(T) (SEQ ID NO: 22); TTN (SEQ ID NO: 23); and NNNVRYAC (SEQ ID NO: 24) (shown in the 5'→3' direction). Each N may independently be A, T, C, or G. Each R may independently be A or G. Each Y may independently be C or T. Each W may independently be A or T. For example, when SpCas9 is used as the Cas protein, the PAM sequence may be NGG (SEQ ID NO: 18). For example, when Streptococcus thermophilus Cas9 (StCas9) is used as the Cas protein, the PAM sequence may be NNAGAAW (SEQ ID NO: 20). For example, when meningococcal Cas9 (NmCas9) is used, the PAM sequence may be NNNNGATT (SEQ ID NO: 21). For example, when Campylobacter jejuni Cas9 (CjCas9) is used, the PAM may be NNNVRYAC (SEQ ID NO: 24). In one embodiment, the PAM sequence may be attached to the 3' end of a target sequence located on the spacer-unbound strand (where the target sequence located on the spacer-unbound strand refers to a sequence that does not bind to the guide RNA). In one embodiment, the PAM sequence may be located at the 3' end of a target sequence located on the spacer-unbound strand. Target sequences located on the non-spacer-bound strand refer to sequences that do not bind to the guide sequence of the guide RNA. Target sequences located on the non-spacer-bound strand are complementary to target sequences located on the spacer-bound strand.
[0108] The location where the DSB or nick is generated may be near the PAM sequence on the genome. In one embodiment, the location where the DSB or nick is generated may be -0 to -20 or +0 to +20 based on the 5' or 3' end of the PAM sequence present on the non-spacer binding strand. In one embodiment, the location where the DSB or nick is generated may be -1 to -5 or +1 to +5 of the PAM sequence on the non-spacer binding strand. For example, in a CRISPR / Cas gene editing system using SpCas9, SpCas9 cleaves between the third and fourth nucleotides located upstream of the PAM sequence.
[0109] Genome editing process using the CRISPR / Cas gene editing system To aid those skilled in the art, the genome editing process using the CRISPR / Cas gene editing system is briefly explained using the following example.
[0110] For example, an environment may be provided in which a DNA molecule for editing can come into contact with a Cas / gRNA complex. The DNA molecule may be the DNA molecule to be edited. For the purpose of genome editing in a cell, a Cas protein or the nucleic acid encoding it and a guide RNA or the nucleic acid encoding it are introduced into the cell, thereby creating an environment in which the Cas protein and guide RNA can come into contact with the cell's genomic DNA. In this environment in which the Cas protein and guide RNA can come into contact with the cell's genomic DNA, the Cas protein and guide RNA can form a Cas / gRNA complex. Naturally, the Cas protein and gRNA can form a Cas / gRNA complex in a suitable environment even in the absence of the cell's genomic DNA. The guide sequence of the gRNA contained in the Cas / gRNA complex and the PAM sequence on the genome are involved in enabling the Cas / gRNA complex to be guided to a target region in which a pre-designed target sequence exists. The Cas / gRNA complex guided to the target region generates a DSB (for example, in the case of Cas9) in the target region. Subsequently, when the DNA in which the DSB was generated (cleaved) is repaired in the DNA repair process, gene editing is performed at the target region or target site. There are two main pathways for repairing DSBs generated on DNA: homologous recombination repair (HDR) and non-homologous end joining (NHEJ). Of these, HDR, a natural DNA repair system, can be used to modify genomes in various organisms, including humans. HDR-mediated repair can be primarily used to insert a desired sequence into a target region or site, or to induce specific point mutations, but this application is not limited to these uses. HDR-mediated repair may be performed by the DNA repair system HDR and an HDR template (e.g., a donor template that can be provided from outside the cell). NHEJ refers to the process of repairing DSBs on DNA, and in contrast to HDR, it joins cleaved ends without an HDR template. That is, the repair process does not require an HDR template.NHEJ may be DNA repair mechanisms that can be primarily selected to induce indels. Insertions / deletions may refer to mutations in which several nucleotides are deleted, any nucleotide is inserted, and / or a mixture of insertions and deletions in the nucleotide sequence of the nucleic acid prior to gene editing. The development of several indels in the target gene may inactivate the corresponding gene. The DNA repair mechanisms HDR and NHEJ are disclosed in detail in the literature [Sander, Jeffry D., and J. Keith Joung. "CRISPR-Cas systems for editing, regulating and targeting genomes." Nature biotechnology 32.4(2014):347-355.], which is incorporated herein by reference in full.
[0111] Up to this point, gene editing systems, including the CRISPR / Cas gene editing system, have been described in detail. This application relates to a method for predicting off-target effects that may occur in gene editing (e.g., genome editing) processes using gene editing systems. Below, the off-target effects that may occur in gene editing systems will be described in detail.
[0112] Off-target In the field of gene editing (e.g., genome editing), off-target refers to gene modifications that occur at unintended locations. Gene modifications induced by off-target effects can be nonspecific. Developed genome editing tools include CRISPR / Cas gene editing systems, activator-like effector nucleases (TALENs), meganucleases, and zinc finger nucleases. These genome editing tools or gene editing systems are designed to perform editing in a target region through a special mechanism that allows them to bind to predetermined sequences (e.g., sequences within the target region). For example, in the CRISPR / Cas gene editing system, guide RNA (gRNA) induces the movement of the Cas / gRNA complex to the intended target site. PAM sequences in the genome can also be involved in the movement to the target site. However, the Cas / gRNA complex can still bind to sequences at unintended locations that are not within the target region. In this way, when the Cas / gRNA complex binds to a sequence at an unintended location and a double-strand break (DSB) is generated at an unintended location, unintended gene modifications occur. Off-target effects can lead to unintended gene modifications such as unintended point mutations, deletions, insertions, inversions, and translocations. Binding of genome editing tools to unwanted regions is known to result from partial but sufficient matching of the sequence in the unwanted region to the target sequence. While not theoretically bound, the full text of the following reference [Lin, Yanni, et al. "CRISPR / Cas9 systems have off-target activity with insertions or deletions between target DNA and guide RNA sequences." Nucleic acids research 42.11(2014):7473-7485.], which is incorporated herein by reference, shows that the mechanisms of off-target binding can be grouped into base mismatch tolerance and bulge mismatch.For example, off-target sites may include, but are not limited to, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more mismatches with the guide RNA sequence.
[0113] Off-target issues suggest the potential for disruption in critical coding regions that could lead to serious problems such as cancer. Furthermore, in this research field, off-target issues can also lead to confusion in variables in biological research, suggesting the potential for irreproducible results (see the full text of which is incorporated herein by reference [Eid, Ayman, and Magdy M. Mahfouz. "Genome editing: the road of CRISPR / Cas9 from bench to clinic." Experimental & Molecular Medicine 48.10(2016):e265-e265.]).
[0114] The off-target problem persists not only in CRISPR / Cas gene editing systems but also in base editing and prime editing developed based on CRISPR / Cas gene editing systems. In this specification, off-target may be used as the opposite concept to on-target and may refer to gene modification at unintended locations.
[0115] As described above, off-target effects can result in serious side effects in various forms (e.g., side effects that are difficult to identify, and / or irreversible side effects). Therefore, identifying potential off-target effects during the use of gene editing systems (e.g., genome editing systems) is crucial in therapeutic development and research. Identifying true off-target effects that occur in designed gene editing systems (e.g., CRISPR / Cas9 gene editing systems and specific guide RNAs) requires considerable cost and time. For this reason, various methods for identifying off-target candidates, i.e., predicting off-target effects, have been studied and developed. However, methods for predicting potential off-target effects in gene editing processes (e.g., genome editing processes using gene editing systems) developed up to the filing date of this application still have various problems. Below, we disclose off-target prediction systems that have been studied and developed to date, and the problems associated with them.
[0116] Known off-target prediction systems and their limitations Known off-target prediction systems As described above, various methods have been developed to predict off-target effects that may occur in genome editing using gene editing systems (e.g., CRISPR / Cas gene editing systems). Conventional off-target prediction or off-target candidate identification methods can be classified into three categories according to the mechanism of action (MOA): cell-based off-target prediction systems, in vitro off-target prediction systems, and in silico off-target prediction systems. Examples of prediction systems included in the above categories are as follows: -Cell-based systems: GUIDE-seq (see reference [Tsai, SQ, Zheng, Z., Nguyen, NT, Liebers, M., Topkar, VV, Thapar, V., ... & Joung, JK (2015). GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nature biotechnology, 33(2), 187-197.]), GUIDE-tag, DISCOVER-seq, BLISS, BLESS, integrase-deficient lentiviral vector-mediated DNA cleavage capture, HTGTS, ONE-seq, CReVIS-Seq, ITR-seq, and TAG-seq. - In vitro systems: Digenome-seq (see reference [Kim, Daesik, et al. "Digenome-seq: genome-wide profiling of CRISPR-Cas9 off-target effects in human cells." Nature methods 12.3(2015):237-243.]), DIG-seq (see reference [Kim, Daesik, and Jin-Soo Kim. "DIG-seq: a genome-wide CRISPR off-target profiling method using chromatin DNA." Genome research 28.12(2018):1894-1900.]), SITE-Seq, CIRCLE-Seq, and CHANGE-seq. - In silico systems: Cas-OFFinder (see reference [Bae, Sangsu, Jeongbin Park, and Jin-Soo Kim. "Cas-OFFinder: a fast and versatile algorithm that searches for potential an off-target effect of Cas9 RNA-guided endonucleases." Bioinformatics 30.10(2014):1473-1475.]), CHOPCHOP, and CRISPOR, etc.
[0117] As mentioned above, various methods have been developed to predict off-target events that occur during genome editing processes using CRISPR / Cas genome editing systems (or tools), but the prediction methods or tools developed so far still have limitations. The limitations of each category are shown below.
[0118] Problems with cell-based classes For example, cell-based methods have been shown to have problems such as missing true off-target locations and reducing the efficiency of predictive methods in clinically relevant cell types (e.g., cells used clinically) (see references [Wienert, Beeke, et al. "Unbiased detection of CRISPR an off-target effect in vivo using DISCOVER-Seq." Science 364.6437(2019):286-289; and Shapiro, Jenny, et al. "Increasing CRISPR efficiency and measuring its specificity in HSPCs using a clinically relevant system." Molecular Therapy-Methods & Clinical Development 17(2020):1097-1107.]).
[0119] Furthermore, the inventors of this application have further identified potential problems in cell-based prediction methods (e.g., GUIDE-seq) through their own experiments. For example, the inventors have observed a higher miss rate in cell-based prediction methods and, through experiments, have confirmed that off-target prediction results can vary depending on the cell type (see the experimental examples in this application).
[0120] Issues concerning in vitro and in silico classes In vitro off-target profiling methods and in silico off-target profiling methods have the problem of providing an excessive number of false-positive data points and failing to reflect the intracellular environment, which may be cell-specific, such as chromatin structure and epigenetic modifications (see reference [Kim, Daesik, and Jin-Soo Kim. "DIG-seq: a genome-wide CRISPR off-target profiling method using chromatin DNA." Genome research 28.12(2018):1894-1900.]).
[0121] Furthermore, the inventors of this application have identified further problems that may arise in in vitro-based prediction methods (e.g., Digenome-seq, DIG-seq) and in silico-based prediction methods in their own experiments. For example, the inventors of this application have confirmed a higher false positive rate in in vitro-based prediction methods (see the experimental examples of this application).
[0122] Off-target prediction systems used in the IND process for gene therapy Because each method has its own advantages and disadvantages, various predictive methods were combined to determine or confirm off-target locations in the CRISPR / Cas genome editing system. In a recent study conducted at Intellia, GUIDE-seq, SITE-Seq, and Cas-OFFinder were used to identify potential off-target locations (NTLA-2001, see reference [Gillmore, Julian D., et al. "CRISPR-Cas9 in vivo gene editing for transthyretin amyloidosis." New England Journal of Medicine 385.6(2021):493-502.]). EDITAS Medicine used GUIDE-seq, Digenome-seq, and Cas-OFFinder off-target prediction tools for the candidate therapeutic agent EDIT-101 (see reference [Maeder ML, Stefanidakis M, Wilson CJ, Baral R, Barrera LA, Bounoutas GS, Bumcrot D, Chao H, Ciulla DM, DaSilva JA et al: Development of a gene-editing approach to restore vision loss in Leber congenital amaurosis type 10. Nat Med 2019, 25(2):229-233.]).
[0123] However, using a combination of various prediction methods requires considerable effort and cost, making it difficult to use in general populations. Furthermore, the use of various off-target prediction methods does not guarantee the detection of more off-target candidates. For example, with respect to NTLA-2001, the seven effective off-target locations identified by SITE-Seq as used herein include all of the effective off-target locations found by GUIDE-seq (three effective off-targets were found) and the effective off-target locations found by Cas-OFFinder (three effective off-targets were found). In this case, the output of one in vitro method is the same as the output of the combination of the three methods. With respect to NTLA-2001, SITE-Seq identified 475 off-target candidates, of which 468 were identified as false positives. In clinical trials, it is difficult to validate all 475 off-target candidates for cells in each patient or organ.
[0124] Therefore, for various reasons, including those described above, it is necessary to develop a concentrated off-target prediction method. The inventors of this application have developed an effective off-target prediction method that is more accurate and has a lower false-positive rate. The inventors of this application confirmed the superior performance of the novel off-target prediction method by comparing the performance of the novel off-target prediction method with that of conventional methods in detail using various and multiple test methods. In particular, the inventors of this application confirmed that the off-target prediction method of this application performs better than the in vitro off-target prediction method by comparing it with the in vitro off-target prediction method. Furthermore, the inventors of this application confirmed that the off-target prediction method of this application performs better than the cell-based off-target prediction method by comparing it with the cell-based off-target prediction method. Furthermore, the inventors of this application comprehensively confirmed through various and multiple tests that the off-target prediction method of this application performs better than other off-target prediction methods (see the experimental examples of this application).
[0125] The off-target prediction method provided herein has all the advantages of cell-based and in vitro prediction methods. Regarding the advantages of cell-based prediction methods, the off-target prediction method provided herein can provide an environment in which the Cas / gRNA complex can come into contact with genomic DNA in which chromatin structure and epigenetic modifications are maintained. Regarding the advantages of in vitro prediction methods, the off-target prediction method provided herein can prevent DNA repair mechanisms from accumulating cut rates, thereby preventing the missed occurrence of true off-targets. The off-target prediction system (method or tool) provided herein is described in detail below.
[0126] Overview of the Off-Target Prediction System Provided by This Application Overview of the Off-Target Prediction System Provided by This Application This application provides a method for predicting off-targets that may occur in a gene editing process. This application provides a method for predicting off-targets that may occur in a genome editing process. In one embodiment, the gene editing process may be performed using a CRISPR / Cas gene editing system. In one embodiment, the genome editing process may be performed using a CRISPR / Cas gene editing system. In one embodiment of this application, a method is provided for predicting off-targets that may occur in a genome editing process using a CRISPR / Cas gene editing system. Off-target encompasses the concept of off-target sites. For example, an off-target site or location may be described as an off-target. In this application, off-target prediction may mean the identification of off-target candidates. Off-target prediction may mean the identification of off-target candidates. The descriptions of “off-target,” “off-target prediction,” and “off-target candidates” as used herein should not be construed as limiting.
[0127] The novel off-target prediction system provided herein physically cleaves a cell, thereby bringing into contact the cell's genomic DNA and gene-editing proteins (e.g., Cas proteins such as the Cas9 protein) and gRNA, or Cas / gRNA complexes (e.g., Cas9 / gRNA complexes). By physically cleaving the cell, the cell membrane and / or nuclear membrane may be cleaved, creating an environment in which the Cas / gRNA complex can come into contact with the genome.
[0128] In one embodiment, the physical disruption of cells may be carried out using a filter with pores of an appropriate size to have less impact on the cellular genomic DNA. For example, an extruder equipped with a filter with appropriately sized pores may be used to provide an environment in which the cellular genomic DNA can come into contact with the Cas / gRNA complex. Pressure may be applied to the region containing the cells, and this pressure may cause the cells to pass through pores having a diameter smaller than the size of the cells, and the cells may be disrupted as they pass through the pores. If the pore size is appropriately adjusted, the genomic DNA or the structure of the genomic DNA in the cells (e.g., structures corresponding to epigenetic features such as chromatin structure) may not be disrupted or altered during the cell disruption process.
[0129] Depending on the characteristics of the off-target system of this invention, an environment can be created in which Cas proteins and gRNA (or Cas / gRNA complexes) can approach or come into contact with the genomic DNA of the cell, while at the same time, more complete genomic DNA can be maintained until it is cleaved by the Cas / gRNA complex.
[0130] The off-target prediction system (or method) provided herein may be referred to as Extru-seq.
[0131] The off-target prediction system provided by this application can be broadly divided into two processes: obtaining the composition to be analyzed (the target composition) and analyzing the composition to be analyzed.
[0132] Here, the acquisition of the composition to be analyzed may be carried out by a process that includes providing an environment in which genomic DNA can come into contact with the Cas / gRNA complex.
[0133] The analysis of the composition to be analyzed may be carried out by a process that includes analyzing the DNA contained in the composition (e.g., cleaved DNA or uncleaved DNA).
[0134] The following discloses a method for providing an environment in which genomic DNA can come into contact with a Cas / gRNA complex.
[0135] A method for providing an environment in which genomic DNA can come into contact with a Cas / gRNA complex. As described above, the method provided by this application for predicting off-targets that may occur in a gene editing process can provide an environment in which genomic DNA can come into contact with the Cas / gRNA complex.
[0136] A process of disrupting the cell (e.g., using physical force) may be performed to provide an environment in which genomic DNA present in the cell (e.g., the nucleus) can come into contact with the Cas / gRNA complex. Through cell disruption, cellular membrane structures such as the cell membrane and / or nuclear membrane may be disrupted, or a space may be created in the membrane structure in which the Cas protein and gRNA, or the Cas / gRNA complex, can approach the genomic DNA. For example, the nuclear membrane of the cell may be disrupted to expose the genomic DNA to the Cas protein and gRNA. In another example, the cell membrane may be disrupted, an environment may be provided in which the Cas protein can come into contact with the gRNA, and the Cas protein and gRNA (or Cas / gRNA complex) may approach the genomic DNA through the nuclear membrane of the cell. In some embodiments, the Cas protein may be fused or ligated with an NLS (i.e., an NLS-ligated Cas protein may be provided), and the NLS fused or ligated to the Cas protein may help the Cas protein (or Cas / gRNA complex) to pass through the nuclear membrane of the cell. In one embodiment, cell disruption may result in the disruption of the cell's membrane structure. In one embodiment, cell disruption may result in the disruption of the cell's membrane. In one embodiment, cell disruption may result in the disruption of the cell's nuclear membrane.
[0137] In one embodiment, the cell disruption described above can be achieved by passing cells through a porous structure having pores. The porous structure may be a filter or membrane having pores. For example, cell disruption may be performed by passing cells through a filter having pores. For example, cell disruption may be performed by passing cells through pores having a diameter smaller than the size of the cell. For example, cell disruption may be performed by passing cells through pores having a diameter smaller than the size of the cell nucleus. Here, the driving force that enables the cells to pass through the filter may be pressure. In particular, pressure is applied to the area where the cells are located, and the applied pressure forces the cells to pass through pores smaller than the size of the cell. Here, the cells can be disrupted as they pass through pores smaller than the size of the cell. In one embodiment, cell disruption may be performed by an extrusion process.
[0138] In one embodiment, cell disruption and contact between genomic DNA and Cas / gRNA complexes can be achieved by the use of an extruder. The extruder and its use will be described in detail later in this disclosure.
[0139] Although the above explanation has been illustrated through Cas proteins and gRNAs or Cas / gRNA complexes, they can be fully applied to editing proteins used in gene editing systems other than the CRISPR / Cas gene editing system.
[0140] When genomic DNA comes into contact with a Cas / gRNA complex, cleavage occurs at on-target and off-target sites on the genomic DNA. Here, cleavage can be achieved by double-strand breaks (DSBs) or nicks induced by the Cas / gRNA complex (particularly the Cas protein). Since the cell's DNA repair mechanism can be disrupted during the cell disruption process, the cleaved DNA cannot be repaired. By analyzing the cleaved or uncleaved DNA, sites where off-target events are likely to occur can be analyzed. That is, off-target (or off-target sites) can be predicted, or off-target candidates (or candidate off-target sites) can be identified.
[0141] Advantages of the method for predicting off-target events disclosed in this application The inventors of this application have tested in detail the off-target prediction method provided herein. By comparing the off-target prediction method of this application with other off-target prediction methods, it has been confirmed that the off-target prediction method of this application performs better than other off-target prediction methods (see the experimental examples of this application). The off-target prediction method of this application exhibits several advantages not possessed by other off-target prediction methods. The off-target prediction method of this application may possess the advantages of both cell-based and in vitro off-target prediction methods.
[0142] The off-target prediction method of this application may have a lower false-positive rate than in vitro off-target prediction methods. For example, in vitro off-target prediction methods have a high false-positive rate in off-target prediction results because they have difficulty reflecting epigenetic features such as chromatin structure and epigenetic modifications. Detecting sites that are not true off-target as off-target candidates can be represented by false-positive results. A high false-positive rate may be associated with a low efficacy confirmation rate. Furthermore, in conventional in vitro off-target prediction methods, cell-specific epigenetic features are not easily reflected in the in vitro off-target prediction results. However, in the off-target prediction method of this application, cells are physically destroyed without the use of chemical additives to maintain the structure of genomic DNA, so the cell-specific environment can be partially maintained, resulting in a lower false-positive rate. The off-target prediction method of this application may exhibit a high efficacy confirmation rate. Furthermore, epigenetic features can be reflected in the off-target prediction results.
[0143] The off-target prediction method of this application may exhibit a lower miss rate than cell-based off-target prediction methods. A lower miss rate may mean missing true off-targets. For example, false negative results, such as when a true off-target site is not detected as an off-target candidate, increase the miss rate. For example, in the process of cell-based prediction methods, the involvement of DNA repair mechanisms may be unavoidable, and repair cleavage sites repaired by DNA repair mechanisms may interfere with the identification of true off-targets or off-target candidates. However, in the off-target prediction method of this application, since cells can be destroyed, DNA repair mechanisms cannot be involved.
[0144] The off-target prediction method of this application can be applied without limiting the cell type. For example, cell-based prediction methods may be difficult to perform on some cells and may be difficult to apply to cells used in actual clinical practice. If off-target prediction is performed based on cells unrelated to the cells used in actual clinical practice, inaccurate results may be obtained. For example, since epigenetic features vary by cell type, the use of different cell types may lead to inaccurate results. However, the off-target prediction method has no or fewer limitations on cell type.
[0145] Furthermore, the off-target prediction method of this application can be implemented more simply and at lower cost than cell-based prediction methods or in vitro off-target prediction methods.
[0146] The off-target prediction method of this application involves the physical destruction of cells, which may result in the aforementioned advantages. The inventors of this application have tested and confirmed the effectiveness of the off-target prediction method of this application through numerous and many types of experiments. The advantages of the off-target prediction method of this application have been confirmed through the experimental examples of this application.
[0147] In one embodiment, the efficacy confirmation rate calculated based on the top 10 off-target candidates identified by the off-target prediction method of this application may be 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100%, or higher, but is not limited to these. In one embodiment, the efficacy confirmation rate calculated based on the top 10 off-target candidates identified by the off-target prediction method of this application may be within a range formed by two values selected from the aforementioned values, but is not limited to these. The efficacy confirmation rate may be influenced by the type of gene editing tool and the type of cell used in the off-target prediction system.
[0148] In one embodiment, the miss rate of the off-target prediction method of this application may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40%, or lower, but the present invention is not limited thereto. In one embodiment, the miss rate of the off-target prediction method of this application may be within a range formed by two values selected from the aforementioned values, but the present application is not limited thereto. The miss rate may be influenced by the type of gene editing tool and the type of cell used in the off-target prediction system.
[0149] In one embodiment, a receiver operating characteristic (ROC) curve for the off-target prediction method of the present invention may be plotted. In one embodiment, the area under the receiver operating characteristic curve (AUC) for the off-target prediction method of the present invention may be calculated. The ROC curve and AUC are powerful tools that can indicate the diagnostic capability of a binary classifier system. The ROC curve can be plotted by associating the true positive rate (TPR) and the false-positive rate (FPR), or by associating sensitivity and specificity. For example, the ROC curve can be plotted with TPR on the y-axis and FPR on the x-axis. For example, the ROC curve can be plotted with sensitivity on the y-axis and specificity on the x-axis. The closer the AUC is to 1 (i.e., the larger the area under the AUC), the higher the performance model of the system. In one embodiment, the AUC of the off-target prediction method of the present application may be calculated, and the AUC is approximately 0.4, 0.42, 0.44, 0.46, 0.48, 0.5, 0.52, 0.54, 0.56, 0.58, 0.6, 0.62, 0.64, 0.66, 0.68, 0.7, 0.72, 0.74, 0.75, 0.76, 0.77 The AUC may be 0.78, 0.79, 0.8, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, or greater than or equal to 1, but the Application is not limited thereto. In one embodiment, the AUC calculated for the off-target prediction method of the Application may be within the range of two values selected from the above values, but the Application is not limited thereto. The AUC may be influenced by the gene editing tool and cell type used in the off-target prediction system.
[0150] As described above, the off-target prediction method of this application comprises obtaining the composition to be analyzed and analyzing the composition to be analyzed. It will be apparent to those skilled in the art that, in addition to the two processes described above, the method may further comprise additional processes. The acquisition of the composition to be analyzed will be described in detail below.
[0151] Acquisition of the composition to be analyzed The off-target prediction method of the present application may include a process of obtaining an analyte composition, where the analyte composition may refer to a composition containing cleaved DNA and / or uncleaved DNA. The off-target prediction method of the present application may be realized through a method comprising obtaining an analyte composition and analyzing the analyte composition (e.g., analysis of cleaved DNA contained in the analyte composition). In order to obtain an analyte composition containing cleaved DNA, the genomic DNA must come into contact with a Cas protein and gRNA (or a Cas / gRNA complex). The Cas / gRNA complex cleaves on-target and / or off-target sites upon contact with the genomic DNA. To bring the Cas / gRNA complex into contact with the genomic DNA, the cell may be disrupted. That is, an environment in which the genomic DNA can come into contact with the Cas / gRNA complex may be provided by cell disruption (e.g., disruption of the cell's membrane structure).
[0152] One of the main processes for obtaining the composition to be analyzed is the physical destruction of cells. In one embodiment, cells may be destroyed by a physical method. In one embodiment, cells may be destroyed by physical force. In one embodiment, the composition to be analyzed may be obtained by destroying cells from an initial composition. In a particular embodiment, the initial composition may contain cells. In a particular embodiment, the initial composition may contain cells and gene editing tools (e.g., Cas protein and gRNA).
[0153] The physical disruption of cells, which is one of the main features of the off-target prediction method of the present application, is disclosed in detail below.
[0154] Physical disruption of cells Overview of physical disruption of cells As described above, the off-target prediction method of one embodiment of the present application may include physically disrupting cells. Here, it is important to note that chemical additives that may cause damage to genomic DNA or genomic DNA structures (e.g., chromatin structures, etc.) are not used for the main purpose of disrupting cells. In one embodiment, the physical disruption of cells may be performed by forcing the cells to pass through a porous structure with pores smaller than the cell size. In one embodiment, the porous structure may be a filter with pores. Hereinafter, the process of physically disrupting cells performed by forcing the cells to pass through a porous structure with pores smaller than the cell size will be described in detail.
[0155] Disruption of cells using a filter with pores and pressure In one embodiment, the physical disruption of cells may be performed by a method that includes passing the cells through a filter with pores smaller than the cell size. Here, the force for passing the cells through the filter may be pressure.
[0156] For example, pressure may be applied to a first container in which a first composition containing cells is located. Here, the applied pressure disrupts the cells while passing them through a filter having pores smaller than the cell size. That is, the driving force for passing the cells through the filter may be pressure. When pressure is applied to the first container or the first composition, the mixed solution and the contained components (e.g., cells) in the first container may be released from the first container through the pores of the filter. In this process, the cells can be disrupted by pores smaller than the cell size.
[0157] In one embodiment, the cell membrane may be disrupted by pores smaller than the cell size. In one embodiment, the cell membrane and the nuclear membrane may be disrupted by pores smaller than the cell size.
[0158] When describing multiple cells, some or all of the cells may be destroyed in the process of passing through pores smaller than the cell size. Of the multiple cells, some may not be destroyed by passing through pores larger than the cell size, or they may not be destroyed by passing through pores smaller than the cell size either.
[0159] In one embodiment, the first composition located in the first container may further include a tool used in a gene editing system (e.g., a gene editing tool). For example, the first composition located in the first container may further include a Cas protein and gRNA. When pressure is applied to the first container, the cells are destroyed, and components of the destroyed cells may move to the second container, and the Cas protein and gRNA may move to the second container through the pores. In the second container located on the opposite side of the first container based on a filter, the cellular Cas protein and gRNA (or Cas / gRNA complex) and genomic DNA may come into contact. Here, contact between Cas / gRNA and genomic DNA may occur in a newly formed vesicle (e.g., a liposome) and / or in an extravesicular environment that is not inside the vesicle, but is not otherwise limited.
[0160] In another embodiment, the first composition located in the first container may contain cells, and the second container located on the opposite side of the first container based on a filter may contain tools used in a gene editing system. For example, if pressure is applied to the first container, the cells are destroyed, and components of the destroyed cells move to the second container where the gene editing tools are located. Thus, in the second container, the gene editing tools (e.g., Cas protein and gRNA) may come into contact with DNA molecules derived from the cells.
[0161] In the second container, when contact is established between the gene editing tool and DNA molecules, an environment is created in which the DNA molecules (e.g., genomic DNA) can be cleaved by the gene editing tool.
[0162] Filter with holes As mentioned above, cell destruction can be achieved by passing cells through a porous structure containing pores.
[0163] In one embodiment, the porous structure may be a filter with pores. In one embodiment, the filter may be formed from, but is not limited to, polycarbonate, cellulose, mixed cellulose ester membrane, glass, polyethersulfone, nylon, polytetrafluoroethylene (PTFE), and PVDF, or a combination thereof, and may be a filter conventionally used in the bio and / or chemical fields. In a particular embodiment, the filter may be, but is not limited to, a polycarbonate membrane filter.
[0164] In one embodiment, the filter may have pores having a diameter smaller than the cell size. In one embodiment, the filter may have pores having a diameter smaller than the average size of the cells. In one embodiment, the filter may have pores having a diameter smaller than the size of the cell nucleus. In one embodiment, the filter may have pores having a size smaller than the average size of the cell nucleus. The filter may be suitably designed depending on the cell type. In one embodiment, the average diameter of the pores contained in the filter may be smaller than the cell size (e.g., cell diameter). In one embodiment, the average diameter of the pores contained in the filter may be smaller than the size of the cell nucleus (e.g., diameter of the cell nucleus).
[0165] In one embodiment, the filter has diameters of approximately 0.1 μm, 0.2 μm, 0.3 μm, 0.4 μm, 0.5 μm, 0.6 μm, 0.7 μm, 0.8 μm, 0.9 μm, 1 μm, 1.5 μm, 2 μm, 2.5 μm, 3 μm, 3.5 μm, 4 μm, 4.5 μm, 4.4 μm, 5 μm, 5.5 μm, 6 μm, 6.5 μm, 7 μm, 7.5 μm, 8 μm, 8.5 μm, 9 μm, 9.5μm, 10μm, 11μm, 12μm, 13μm, 14μm, 15μm, 16μm, 17μm, 18μm, 19μm, 20μm, 21μm, 22μm, 23μm, 24μm, 25μm, 26μm, 27μm, 28μm, 29μm, 30μm, 31μm, 32μm, 33μm, 34μm, 35μm, 36μm, 37μm, 38μm, 39μm, 40μm, 4 1μm, 42μm, 43μm, 44μm, 45μm, 46μm, 47μm, 48μm, 49μm, 50μm, 51μm, 52μm, 53μm, 54μm, 55μm, 56μm, 5 7μm, 58μm, 59μm, 60μm, 61μm, 62μm, 63μm, 64μm, 65μm, 66μm, 67μm, 68μm, 69μm, 70μm, 71μm, 72μm, 73 The filter may have holes having one of the following diameters: μm, 74μm, 75μm, 76μm, 77μm, 78μm, 79μm, 80μm, 81μm, 82μm, 83μm, 84μm, 85μm, 86μm, 87μm, 88μm, 89μm, 90μm, 91μm, 92μm, 93μm, 94μm, 95μm, 96μm, 97μm, 98μm, 99μm, and 100μm. In one embodiment, the filter may have holes having a diameter of one of the above values or smaller.
[0166] In one embodiment, the average diameter of the pores contained in the filter is approximately 0.1 μm, 0.2 μm, 0.3 μm, 0.4 μm, 0.5 μm, 0.6 μm, 0.7 μm, 0.8 μm, 0.9 μm, 1 μm, 1.5 μm, 2 μm, 2.5 μm, 3 μm, 3.5 μm, 4 μm, 4.5 μm, 4.4 μm, 5 μm, 5.5 μm, 6 μm, 6.5 μm, 7 μm, 7.5 μm, 8 μm, 8.5 μm. 9μm, 9.5μm, 10μm, 11μm, 12μm, 13μm, 14μm, 15μm, 16μm, 17μm, 18μm, 19μm, 20μm, 21μm, 22μm, 23μm, 24μm, 25μm, 26μm, 27μm, 28μm, 29μm, 30μm, 31μm, 32μm, 33μm, 34μm, 35μm, 36μm, 37μm, 38μm, 39μm, 40μm, 41μm, 4 2μm, 43μm, 44μm, 45μm, 46μm, 47μm, 48μm, 49μm, 50μm, 51μm, 52μm, 53μm, 54μm, 55μm, 56μm, 57μm, 58μm, 5 9μm, 60μm, 61μm, 62μm, 63μm, 64μm, 65μm, 66μm, 67μm, 68μm, 69μm, 70μm, 71μm, 72μm, 73μm, 74μm, 75μm, 76 The pore diameter may be any one selected from μm, 77μm, 78μm, 79μm, 80μm, 81μm, 82μm, 83μm, 84μm, 85μm, 86μm, 87μm, 88μm, 89μm, 90μm, 91μm, 92μm, 93μm, 94μm, 95μm, 96μm, 97μm, 98μm, 99μm, and 100μm, or any one of the aforementioned values or smaller. In a particular embodiment, the average diameter of the pores contained in the filter may be approximately 5μm, 6μm, 7μm, 8μm, 9μm, 10μm, 11μm, 12μm, 13μm, 14μm, or 15μm. In a particular embodiment, the average diameter of the pores contained in the filter may be 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, 11 μm, 12 μm, 13 μm, 14 μm, or 15 μm, or smaller.
[0167] In one embodiment, the average pore diameters in the filter are 0.1 μm, 0.2 μm, 0.3 μm, 0.4 μm, 0.5 μm, 0.6 μm, 0.7 μm, 0.8 μm, 0.9 μm, 1 μm, 1.5 μm, 2 μm, 2.5 μm, 3 μm, 3.5 μm, 4 μm, 4.54 μm, 5 μm, 5.5 μm, 6 μm, 6.5 μm, 7 μm, 7.5 μm, 8 μm, and 8.5 μm. , 9μm, 9.5μm, 10μm, 11μm, 12μm, 13μm, 14μm, 15μm, 16μm, 17μm, 18μm, 19μm, 20μm, 21μm, 22μm, 23μm, 24μm, 25μm, 26μm, 27μm, 28μm, 29μm, 30μm, 31μm, 32μm, 33μm, 34μm, 35μm, 36μm, 37μm, 38μm, 39μm, 40 μm, 41μm, 42μm, 43μm, 44μm, 45μm, 46μm, 47μm, 48μm, 49μm, 50μm, 51μm, 52μm, 53μm, 54μm, 55μm, 56μ m, 57μm, 58μm, 59μm, 60μm, 61μm, 62μm, 63μm, 64μm, 65μm, 66μm, 67μm, 68μm, 69μm, 70μm, 71μm, 72μm, The values may be within the range of two values selected from 73μm, 74μm, 75μm, 76μm, 77μm, 78μm, 79μm, 80μm, 81μm, 82μm, 83μm, 84μm, 85μm, 86μm, 87μm, 88μm, 89μm, 90μm, 91μm, 92μm, 93μm, 94μm, 95μm, 96μm, 97μm, 98μm, 99μm, and 100μm.
[0168] In some embodiments, one or more filters may be used to achieve physical cell destruction. For example, one filter may be used, for instance, a first filter having pores with a first average diameter. In another example, multiple filters may be used. For example, a first filter having pores with a first average diameter may be used first, and a second filter having pores with a second average diameter (i.e., a filter having a pore profile different from that of the first filter) may be used second. The types and number of filters that can be used to achieve physical cell destruction are not otherwise limited.
[0169] pressure As mentioned above, for a cell to pass through a pore smaller than its own size, a force must be applied to the region in which the cell is located.
[0170] The force that causes a cell to pass through a pore (for example, while being destroyed) may be pressure. That is, if pressure is applied to the region in which the cell is located (for example, the container containing the cell), the cell may pass through a pore smaller than the size of the cell while being destroyed. Here, the pressure may be applied in a variety of ways and is not otherwise limited.
[0171] In one embodiment, the application of pressure may be performed by a human. For example, pressure may be applied by pushing a piston designed to apply pressure to a container containing cells. In one embodiment, the application of pressure may be achieved by a machine or device. For example, pressure may be applied by pushing a piston designed to apply pressure to a container containing cells through a machine. In another example, the application of pressure may be achieved by centrifugation. In one embodiment, pressure may be centrifugal force or osmotic pressure. The magnitude or intensity of the applied force (e.g., pressure) is not otherwise limited. For example, a minimum or greater force or pressure may be applied to allow cells to pass through pores and / or filters.
[0172] In one embodiment, cell disruption may be achieved using an extruder. A method for disrupting cells using an extruder is described in detail below.
[0173] A method of destroying cells using an extruder Extruder Overview In this specification, an extruder may refer to a tool or machine having a container and a porous structure having holes, designed to pass a composition filled in the container through the porous structure having holes by force applied to the container. An example of an extruder used in the bio and chemical fields is the Avanti® Mini-Extruder. The Mini-Extruder has two containers, each contained in two syringes, and a porous filter (or membrane) located between the two containers. The inventors of the present application found that such an extruder structure is suitable for physically destroying cells and have used the extruder to destroy cells. The above description of the Avanti® Mini-Extruder is an example to aid the understanding of those skilled in the art, and the extruders disclosed herein are not limited to the Mini-Extruder. The extruders disclosed herein may be understood to encompass a tool or machine having at least one container and a porous structure (filter or membrane) and capable of achieving cell destruction.
[0174] In one embodiment, the term “extrusion” may be understood to encompass a series of processes that pass components located in a container through a porous structure (filter or membrane) under pressure. For example, an example of extrusion is the process in which pressure is applied to a first container in which a composition containing cells is located, causing the cells to pass through a filter while being destroyed by the pressure. In another example, an example of extrusion is the process of passing Cas proteins and / or gRNAs from the first container through a filter under pressure to a region other than the first container (e.g., a second container).
[0175] In one embodiment, the extruder may be a unidirectional extruder designed to pass through the filter once, but the present application is not limited thereto. In one embodiment, the extruder may be a bidirectional extruder designed to pass through the filter multiple times, but the present invention is not limited thereto. The bidirectional extruder may include at least two containers and a filter located between the two containers. For example, the above-mentioned Avanti (registered trademark) Mini-Extruder may be a bidirectional extruder. In one embodiment, when a bidirectional extruder is used, the cell disruption rate can be increased by passing the cells through the filter multiple times, but it is not limited in other ways.
[0176] Hereinafter, using the above-mentioned extruder, a method for generating an environment in which a gene editing tool and a DNA molecule can come into contact, including a cell disruption process, will be described in more detail.
[0177] Cell disruption through an extruder The off-target prediction method of the present application may include the use of an extruder. For example, an extruder including a first container, a second container, and a filter may be used. Here, the filter may be located between the first container and the second container. Hereinafter, an example of the use of an extruder including a first container, a second container, and a filter will be disclosed.
[0178] In one embodiment, a starting composition containing cells, Cas protein, and gRNA may be filled into the first container. Pressure may be applied to the first container where the starting composition is located. For example, applying pressure may be performed by pushing a piston connected to the first container, which is designed to apply pressure to the first container. That is, the pressure may be applied to the first container by pushing a piston connected to the first container in the direction of the first container and the filter. In one embodiment, pressure may be applied to move the components of the starting composition (including cells, Cas protein, and gRNA) to the second container.
[0179] When pressure is applied to the first container, components contained in the initial composition may move to the second container through a filter with pores. Here, cells larger than the size of the pores may be destroyed as they pass through the filter. As described above, cell destruction may be destruction of the cell membrane, or destruction of both the cell membrane and the nuclear membrane. Finally, a mixture containing the destroyed cells, Cas protein, and components obtained from gRNA may be contained in the second container. In this process, the Cas / gRNA complex comes into contact with DNA (e.g., genomic DNA), which is one of the components obtained from the destroyed cells. Furthermore, the mixture in the second container may or may not contain undestroyed cells.
[0180] Subsequently, optionally, pressure may be applied to the second container in which the mixture is located to transfer the components of the mixture (through the filter) to the first container. Thus, the mixture may be contained in the first container. Subsequently, optionally, pressure may be applied to the first container in which the mixture is located to transfer the components of the mixture (through the filter) to the second container. Similarly, the extruder may be used so that the components of the composition filled into the extruder or components derived from the composition pass through its filter several times. Passing through the filter several times may increase the rate of cell disruption and / or increase the rate of contact between the Cas / gRNA complex and genomic DNA.
[0181] In one embodiment, the extrusion may be performed n times. In one embodiment, the filter pass may be performed n times, where n is an integer. Here, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, n may be 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100, but the present invention is not limited thereto. In one embodiment, n may be any one or fewer of the above values. In one embodiment, n may be any one or more of the above values. In one embodiment, n may be within the range determined by any two selected from the above values.
[0182] The composition to be analyzed may be obtained through a process that includes the use of an extruder. In one embodiment, any one or more of incubation, RNA removal, and DNA purification may be further performed after the extrusion process.
[0183] In one embodiment, an incubation process may be performed after the extrusion process to accumulate the cleavage rate. That is, after the cell disruption process, a further process may be performed to incubate a composition containing components of the disrupted cells. For example, the incubation time may be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 26, 28, 30, 32, 36, 38, 40, 42, 44, 46, or 48 hours, or longer than the values mentioned above, but is not otherwise limited. For example, after the extrusion process is completed, the composition to be analyzed may finally be obtained through incubation (e.g., incubation at 37°C) and RNA removal, where the composition to be analyzed is the composition to be used in subsequent analytical processes. Here, for example, the DNA contained in the composition to be analyzed may be DNA suitable for analysis (e.g., sequencing). For example, the composition to be analyzed may contain cleaved DNA. In addition to cleaved DNA, the composition to be analyzed may contain uncleaved DNA. The process of forming cleaved DNA has been described in detail above. For example, the Cas / gRNA complex may come into contact with DNA (e.g., genomic DNA), and the DNA may be cleaved via DSBs or nicks induced by the Cas / gRNA complex.
[0184] Analysis of the composition to be analyzed The composition to be analyzed As described above, the composition to be analyzed may be obtained by a method comprising contacting a gene editing tool (e.g., a Cas / gRNA complex) with genomic DNA, according to one embodiment of the present application. In one embodiment, the composition to be analyzed may contain cleaved genomic DNA. In one embodiment, the composition to be analyzed may contain one or more cleaved DNA (e.g., double-stranded DNA or single-stranded DNA). In one embodiment, the cleaved genomic DNA may contain one or more cleavage at one or more cleavage sites. For example, as described above, one or more cleavage may occur via DSBs or nicks induced by the Cas / gRNA complex in contact with the genomic DNA. In one embodiment, the cleavage site may be an off-target or on-target site. That is, cleavage may occur via DSBs or nicks induced (or generated) at an off-target or on-target site by the Cas / gRNA complex in contact with the genomic DNA.
[0185] In one embodiment, the composition under analysis may reflect the advantages of an in vitro off-target prediction system. For example, the cleaved genomic DNA contained in the composition under analysis does not have to be repaired genomic DNA, because some or all of the DNA repair mechanisms have been inactivated.
[0186] In one embodiment, the composition being analyzed may reflect the advantages of a cell-based off-target prediction system. For example, cleaved genomic DNA contained in the composition being analyzed may reflect cell-specific epigenetic features.
[0187] After the composition to be analyzed has been obtained, it may be analyzed to obtain information about genomic DNA cleavage. Thus, information about off-target candidates that may be generated with the use of gene editing systems may be obtained. Information about off-target candidates may be applied to predict off-targets; that is, off-targets that may be generated when using gene editing systems (e.g., CRISPR / Cas gene editing systems) can be predicted.
[0188] The following discloses a method for obtaining information about off-target candidates by analyzing a composition to be analyzed.
[0189] Summary of the analysis of the composition being analyzed As described above, the method of the present application comprises analyzing a composition to be analyzed, which contains the cleaved genomic DNA obtained above. Information about the cleavage of genomic DNA (e.g., information about one or more cleavage sites and / or cleavage scores at one or more cleavage sites) can be obtained by analyzing the composition to be analyzed. Based on the information about the cleavage of genomic DNA, information about off-target candidates (e.g., information about one or more off-targets and / or scores for one or more off-targets) can be obtained.
[0190] Analysis of the composition to be analyzed and acquisition of information on genomic DNA cleavage. In one embodiment, information about genomic DNA cleavage can be obtained by analyzing the DNA contained in the composition to be analyzed (e.g., cleaved and / or uncleaved genomic DNA). In one embodiment, information about genomic DNA cleavage can be obtained by analyzing the cleaved DNA contained in the composition to be analyzed. In one embodiment, information about genomic DNA cleavage can be obtained by analyzing one or more cleavage sites. Here, the analytical method capable of identifying the cleavage site of cleaved DNA is not particularly limited. For example, any analytical method capable of identifying the cleavage site of cleaved DNA can be fully used in the off-target prediction method of this application.
[0191] In one embodiment, DNA analysis may be performed through DNA analysis methods well known to those skilled in the art. In one embodiment, DNA analysis may be performed by any one or more selected from PCR-based assays (see [Cameron, Peter, et al. "Mapping the genomic landscape of CRISPR-Cas9 cleavage." Nature methods 14.6(2017):600-606.]) and sequencing (see [Metzker, Michael L. "Sequencing technologies—the next generation." Nature reviews genetics 11.1(2010):31-46.; and Kumar, Kishore R., Mark J. Cowley, and Ryan L. Davis. "Next-generation sequencing and emerging technologies." Seminars in thrombosis and hemostasis. Vol. 45. No. 07. Thieme Medical Publishers, 2019.]) (e.g., DNA sequencing).
[0192] For example, sequencing may be, but is not limited to, any one or more sequencing methods known as whole-genome sequencing (WGS), deep sequencing, high-throughput sequencing (HTS), de novo sequencing, second-generation sequencing, next-generation sequencing, third-generation sequencing, large-scale sequencing, shotgun sequencing, long-read sequencing, and short-read sequencing.
[0193] In one embodiment, the sequencing depth of the sequencing method used in the analysis of the composition to be analyzed may be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000x. In one embodiment, the sequencing depth may be within a range of two values selected from the aforementioned values. In one embodiment, the sequencing depth may be lower or higher than the aforementioned values. In a particular embodiment, the sequencing depth used in the analysis may be approximately 10 to 40x. The sequencing depth is not otherwise limited, and any sequencing depth that allows for the identification of cleavage sites on the DNA is sufficient.
[0194] Information about DNA cleavage In one embodiment, information about DNA cleavage (e.g., information about the cleavage site of genomic DNA) can be obtained by analyzing the composition to be analyzed.
[0195] In one embodiment, the information about DNA cleavage may include information about one or more cleavage sites, where the cleavage sites can be generated by a gene editing tool.
[0196] In one embodiment, information about DNA cleavage may include information about the location of one or more cleavage sites on genomic DNA. For example, information about DNA cleavage may include location information for each cleavage site on genomic DNA out of all cleavage sites on the cleaved DNA contained in the composition being analyzed. For example, information about DNA cleavage may include location information for each cleavage site on genomic DNA out of one or more cleavage sites on the cleaved DNA contained in the composition being analyzed. That is, through analysis, location information for all cleavage sites or location information for some cleavage sites can be obtained. The obtained location information may be associated with off-target candidates and / or on-target sites. For example, by comparing the location information for an identified cleavage site with predetermined on-target sites, it can be determined whether a cleavage site is associated with an off-target candidate or an on-target.
[0197] In one embodiment, information about DNA cleavage may include cleavage scores for one or more cleavage sites. For example, information about DNA cleavage may include cleavage scores for each of the cleavage sites on the cleaved DNA contained in the composition being analyzed. For example, information about DNA cleavage may include cleavage scores for each of the cleavage sites on the cleaved DNA of the composition being analyzed. That is, through analysis, cleavage scores for all cleavage sites or scores for some cleavage sites can be obtained. In one embodiment, the cleavage score may be calculated from sequence reads. In one embodiment, the cleavage score may be calculated from the results of a Manhattan plot. The calculation mechanism for the cleavage score is not limited to this and can be suitably selected depending on which analytical method is used. In one embodiment, a cleavage rank may be calculated based on the cleavage score. For example, cleavage sites showing higher cleavage scores may be ranked higher. For example, the cleavage site showing the highest cleavage score may be ranked first. In one embodiment, the crack score may be related to the crack rate of the corresponding crack site. Information about the crack score obtained therefrom may be associated with scores for off-target and / or on-target candidates.
[0198] In one embodiment, information about DNA cleavage may include information about the number of cleavage sites generated. For example, the total number of cleavage sites may be calculated. For example, in one calculation of the number of cleavage sites, overlapping sites may be counted as 1. In another example, in another calculation of the number of cleavage sites, overlapping sites may be counted as multiple. For example, if 5 DNA molecules cleave at cleavage site x, the result may be counted as 1 or 5, as needed. From the information about the number of cleavage sites, the total number of off-target candidates that may be generated by the use of the gene editing system can be determined.
[0199] In one embodiment, information about DNA cleavage obtained by analyzing the composition under analysis may, but is not limited to, Location of one or more cleavage sites on genomic DNA; Destruction score for one or more destructive sites; and Number of cracks that occurred It may include any one or more of the following:
[0200] In one embodiment, the process of obtaining information about DNA cleavage by analyzing a composition to be analyzed may further include additional processes for obtaining information about DNA cleavage. For example, these may further include processes such as information (or data) processing and / or normalization of the obtained information (or data). For example, the process may further include comparing the obtained cleavage information with previously determined on-target information. The process of obtaining cleavage information may further include, but is not limited to, the additional processes described above.
[0201] In one embodiment, information about DNA cleavage may further include, but is not limited to, other information that can be obtained through the analysis of the composition being analyzed (e.g., DNA sequencing).
[0202] Obtaining information about off-target activities In one embodiment, information about off-target candidates may be obtained from the acquired cleavage information. A person skilled in the art related to the present application can easily obtain off-target information based on cleavage information, and therefore this disclosure does not limit the process of the off-target prediction system of the present application. A person skilled in the art related to the present application can obtain off-target information using cleavage information (e.g., information about DNA cleavage) obtained by analyzing the composition to be analyzed, either through a suitable process or without additional processes.
[0203] In one embodiment, the off-target prediction method of the present invention may include a process of identifying information about off-target candidates from information obtained about cracks.
[0204] In one embodiment, information about off-target candidates may include information about the location of one or more off-target candidates on genomic DNA (e.g., information about candidate off-target sites). For example, information about the location of off-target candidates may include information about each location (location on genomic DNA) of all off-target candidates. For example, information about the location of off-target candidates may include information about each location of one or more off-target candidates. That is, location information for all candidate off-target sites may be obtained, or location information for one or more candidate off-target sites may be obtained instead of information for all candidate off-target sites. Among the off-target candidates, there may be true off-targets (e.g., genuine off-targets generated from the use of a gene editing system). Information about the location of off-target candidates may be obtained based on the cleavage information described above (e.g., location information for one or more cleavage sites).
[0205] In one embodiment, information about off-target candidates may include off-target scores (e.g., off-target prediction scores) for one or more off-target candidates. For example, information about off-target candidates may include off-target scores for each off-target candidate for all off-target candidates. For example, information about off-target candidates may include off-target scores for each off-target candidate for one or more off-target candidates. That is, off-target scores may be obtained for all candidate off-target sites, or off-target scores may be obtained for one or more candidate off-target sites, but not for all candidate off-target sites. Information about off-target scores for off-target candidates may be obtained based on the aforementioned rupture information (e.g., scores for one or more rupture sites). In one embodiment, ranking of off-target candidates may be calculated based on the obtained off-target scores. For example, off-target candidates with higher off-target scores (e.g., candidate off-target sites) may be ranked higher. For example, the off-target candidate with the highest off-target score may be ranked first. For example, a high off-target score for an off-target candidate may be associated with a true off-target, but is not necessarily limited to one.
[0206] In one embodiment, information about off-target candidates may include information about the number of off-target candidates. For example, the total number of off-target candidates may be calculated. For example, in calculating the number of off-target candidates, overlapping sites may be counted as 1. In another example, in calculating the number of off-target candidates, overlapping sites may be counted as multiple. For example, if 5 candidate off-target sites x are found, they may be counted as 1 or 5. From the information about the number of off-target candidates, the total number of off-target candidates that can be generated by using the gene editing system can be determined. That is, the predicted total number of off-targets can be determined.
[0207] In one embodiment, information about off-target candidates is not limited to the following: The location on genomic DNA for each of one or more off-target candidates; Off-target scores for each of one or more off-target candidates; and Number of predicted off-target candidates It may include any one or more of the following:
[0208] In one embodiment, the process of obtaining information about off-target candidates may further include additional processes for obtaining information about off-target candidates. For example, these may further include processes such as information (or data) processing and / or normalization of the obtained information (or data). For example, these may further include a process of comparing the obtained information about off-target candidates with previously determined on-target information. As described above, the process of obtaining information about off-target candidates may further include, but is not limited to, additional processes.
[0209] In one embodiment, information about off-target candidates may further include, but not be otherwise limited to, additional information that helps predict off-target effects that may be generated by using the gene editing system.
[0210] Relationship with gene editing systems that are being predicted The off-target prediction system of this application may be associated with a gene editing system to which predictions are to be made. Here, the gene editing system to which predictions are to be made may refer to, but is not limited to, a gene editing system that has been determined to be used for research or therapeutic purposes. That is, the gene editing system to which predictions are to be made may refer to a gene editing system (or gene editing process) that needs to predict off-target effects.
[0211] For example, if a specific cell is used in the gene editing system to be predicted, that specific cell may also be used in the method for predicting off-target effects according to the present invention. For example, if a guide RNA having a specific guide sequence is used in the gene editing system to be predicted, that guide RNA having the same guide sequence may also be used in the method for predicting off-target effects according to the present invention.
[0212] In this embodiment, a method for predicting off-target effects according to one embodiment of the present invention may further include a process for confirming the gene editing system to be predicted. The gene editing system to be predicted may be referred to as a predetermined gene editing system. A predetermined gene editing (e.g., genome editing) system may include cells to be gene edited (genome edited) and one or more predetermined gene editing tools. A predetermined gene editing tool may include, for example, multiple types of guide RNA, guide sequences, and gene editing proteins (e.g., Cas proteins).
[0213] In one embodiment, the method for predicting off-target effects according to the present invention may further comprise confirming or designing a predetermined gene editing system. The predetermined gene editing system may be confirmed, thereby designing components suitably used in the off-target prediction system. Here, the process of confirming the predetermined gene editing system may be performed before the composition to be analyzed is obtained. Below, an example of confirming a predetermined gene editing system (to be predicted) is given based on a CRISPR / Cas gene editing system. The exemplary description based on a CRISPR / Cas gene editing system is not intended to limit the embodiments of the off-target prediction system of the present invention, and it is suggested that such exemplary description is sufficiently applicable to different gene editing systems in the same or identical context as described below.
[0214] In one embodiment, the method for predicting off-target effects according to the present invention may include a predetermined CRISPR / Cas gene editing system. Herein, confirming the predetermined CRISPR / Cas gene editing system may include confirming any one or more of the following pieces of information: predetermined cells (i.e., cells for editing to be used in editing the CRISPR / Cas-based gene to be predicted), predetermined type of Cas protein (i.e., type of Cas protein to be used in editing the CRISPR / Cas-based gene to be predicted), and predetermined guide RNA (guide RNA sequence or guide sequence).
[0215] In a particular embodiment, confirming a predetermined CRISPR / Cas gene editing system may include confirming predetermined cells. In a particular embodiment, the same cells as the predetermined cells may be used in the off-target prediction system of the present invention. As a result, cell-specific characteristics may be reflected in the results of the off-target prediction system. The cells to be genome edited are not otherwise limited. In one embodiment, the predetermined cells may be animal cells or plant cells. In one embodiment, the predetermined cells may be human cells or cells derived from non-human animals (e.g., mice, rats, dogs, cats, cattle, pigs, horses, and sheep), but are not otherwise limited. In a particular embodiment, the predetermined cells may be human cells.
[0216] In a particular embodiment, confirming a predetermined CRISPR / Cas gene editing system may include confirming a predetermined Cas protein. In a particular embodiment, the same Cas protein as the predetermined Cas protein may be used in the off-target prediction system of the present invention. As a result, properties that may be affected by the Cas protein may be reflected in the results of the off-target prediction system. In one embodiment, this may be a gene editing system using SpCas9.
[0217] In a particular embodiment, confirming a predetermined CRISPR / Cas gene editing system may include confirming a predetermined guide RNA. In a particular embodiment, the same guide RNA as the predetermined guide RNA may be used in the off-target prediction system of the present invention. As a result, characteristics that may be affected by the guide RNA may be reflected in the results of the off-target prediction system.
[0218] In a particular embodiment, confirming a predetermined CRISPR / Cas gene editing system may include confirming a predetermined guide sequence. In a particular embodiment, the off-target prediction system of the present invention may use a guide RNA having the same guide sequence as the predetermined guide sequence. As a result, characteristics influenced by the guide sequence may be reflected in the results of the off-target prediction system.
[0219] In a particular embodiment, the off-target prediction system of the present invention may use any one or more selected from cells identical to a predetermined cell, a Cas protein identical to a predetermined Cas protein, and a guide RNA having the same guide sequence as a predetermined guide sequence.
[0220] The above description does not limit the requirement that the same components used in the off-target prediction system must necessarily be used in the off-target prediction system as those used in the predetermined gene editing system. The off-target prediction system of this application may be suitably selected by those skilled in the art according to the purpose of the off-target prediction system used. For example, a different type of Cas protein (e.g., a Cas protein known to have similar properties) may be used in the off-target prediction system. In another example, a different type of cell (e.g., a cell known to have similar properties) may be used in the off-target prediction system. In yet another example, a different type of guide RNA (e.g., a guide RNA modified to be more effectively applied to the off-target prediction system) may be used in the off-target prediction system.
[0221] It can be used in conjunction with different off-target prediction systems. In one embodiment, the off-target prediction system of the present invention may be used in conjunction with different off-target prediction systems. For example, the off-target prediction system of the present invention may be used with any one or more selected from in silico-based off-target prediction systems, in vitro-based off-target prediction systems, and cell-based off-target prediction systems. For example, the off-target prediction system of the present invention may be used with any one or more selected from Cas-OFFinder, CHOPCHOP, CRISPOR, Digenome-seq, DIG-seq, SITE-seq, CIRCLE-seq, CHANGE-seq, GUIDE-seq, GUIDE-tag, DISCOVER-seq, BLISS, BLESS, integrase-deficient lentiviral vector-mediated DNA cleavage capture, HTGTS, ONE-seq, CReVIS-Seq, ITR-seq, and TAG-seq. To more effectively discover true off-target sites, an off-target prediction system different from the off-target prediction system of this application may be used, which may be, but not limited to, an off-target prediction system developed before the filing date of this application or an off-target prediction system developed after the filing date of this application.
[0222] Initial composition and components that may be included in the initial composition Overview of the starting composition As described above, according to one embodiment of the present application, the composition to be analyzed can be obtained by disrupting cells. Furthermore, by analyzing the composition to be analyzed, information about genomic DNA cleavage can be obtained.
[0223] In one embodiment, a composition to be analyzed can be obtained from a starting composition containing cells by destroying the cells. In one embodiment, the starting composition may further contain gene editing tools (e.g., Cas protein and guide RNA) in addition to the cells. Hereinafter, conditions regarding components that can be included in the starting composition of the off-target prediction method according to one embodiment of the present application are disclosed.
[0224] Cells and cell concentration In one embodiment, the starting composition may contain cells. In one embodiment, the concentration of the cells contained in the starting composition is approximately 1x10 5 cells / mL, 2x10 5 cells / mL, 3x10 5 cells / mL, 4x10 5 cells / mL, 5x10 5 cells / mL, 6x10 5 cells / mL, 7x10 5 cells / mL, 8x10 5 cells / mL, 9x10 5 cells / mL, 1x10 6 cells / mL, 2x10 6 cells / mL, 3x10 6 cells / mL, 4x10 6 cells / mL, 5x10 6 cells / mL, 6x10 6 cells / mL, 7x10 6 cells / mL, 8x10 6 cells / mL, 9x10 6 cells / mL, 1x10 7 cells / mL, 2x10 7 cells / mL, 3x10 7 cells / mL, 4x10 7 cells / mL, 5x10 7 cells / mL, 6x10 7 cells / mL, 7x10 7 cells / mL, 8x10 7 cells / mL, 9x10 7 cells / mL, 1x10 8 cells / mL, 2x10 8 cells / mL, 3x10 8 cells / mL, 4x108 cells / mL, 5x10 8 cells / mL, 6x10 8 cells / mL, 7x10 8 cells / mL, 8x10 8 cells / mL, 9x10 8 cells / mL, 1x10 9 cells / mL, 2x10 9 cells / mL, 3x10 9 cells / mL, 4x10 9 cells / mL, 5x10 9 cells / mL, 6x10 9 cells / mL, 7x10 9 cells / mL, 8x10 9 cells / mL, 9x10 9 cells / mL, 1x10 10 cells / mL, 2x10 10 cells / mL, 3x10 10 cells / mL, 4x10 10 cells / mL, 5x10 10 cells / mL, 6x10 10 cells / mL, 7x10 10 cells / mL, 8x10 10 cells / mL, or 9x10 10 The cell concentration may be, but is not limited to, cells / mL. In one embodiment, the cell concentration in the initial composition may be in the range between two values selected from the above values. In one embodiment, the cell concentration in the initial composition may be any one of the above values, or higher or lower. In a particular embodiment, the cell concentration in the initial composition is approximately 1 x 10⁻¹⁶ 6 cells / mL, 2x10 6 cells / mL, 3x10 6 cells / mL, 4x10 6 cells / mL, 5x10 6 cells / mL, 6x10 6 cells / mL, 7x10 6 cells / mL, 8x10 6 cells / mL, 9x10 6 cells / mL, 1x10 7 cells / mL, 2x10 7 cells / mL, 3x107 cells / mL, 4x10 7 cells / mL, 5x10 7 cells / mL, 6x10 7 cells / mL, 7x10 7 cells / mL, 8x10 7 cells / mL, 9x10 7 cells / mL, or 1x10 8 The value may be cells / mL.
[0225] The cells that may be used in the off-target prediction system of this application are not otherwise limited. In one embodiment, the cells may be animal cells or plant cells. In one embodiment, the cells may be human cells or cells derived from non-human animals (e.g., mice, rats, dogs, cats, cattle, pigs, horses, and sheep), but are not otherwise limited. In one particular embodiment, the cells may be human cells.
[0226] Gene editing tools and concentrations of the editing tools In one embodiment, the initiation composition may include a gene editing tool. In one embodiment, the initiation composition may include a Cas protein and a gRNA.
[0227] In one embodiment, the concentrations of Cas protein contained in the starting composition are approximately 10 nM, 20 nM, 30 nM, 40 nM, 50 nM, 60 nM, 70 nM, 80 nM, 90 nM, 100 nM, 200 nM, 300 nM, 400 nM, 500 nM, 600 nM, 700 nM, 800 nM, 900 nM, 1000 nM (1 μM), 2000 nM, and 3000 nM. The concentration may be, but is not limited to, nM, 4000nM, 5000nM, 6000nM, 7000nM, 8000nM, 9000nM, 10000nM (10μM), 20000nM, 30000nM, 40000nM, 50000nM, 60000nM, 70000nM, 80000nM, 90000nM, or 100000nM (100μM). In one embodiment, the concentration of Cas protein in the initial composition may be in the range between two values selected from the above values. In one embodiment, the concentration of Cas protein in the initial composition may be any one of the above values, or higher or lower. In a particular embodiment, the concentration of Cas protein contained in the starting composition may be approximately 1000 nM (1 μM), 2000 nM, 3000 nM, 4000 nM, 5000 nM, 6000 nM, 7000 nM, 8000 nM, 9000 nM, or 10000 nM (10 μM).
[0228] In one embodiment, the concentrations of guide RNA contained in the initiation composition are approximately 10 nM, 20 nM, 30 nM, 40 nM, 50 nM, 60 nM, 70 nM, 80 nM, 90 nM, 100 nM, 200 nM, 300 nM, 400 nM, 500 nM, 600 nM, 700 nM, 800 nM, 900 nM, 1000 nM (1 μM), 2000 nM, and 3000 nM. The concentration may be M, 4000nM, 5000nM, 6000nM, 7000nM, 8000nM, 9000nM, 10000nM (10μM), 20000nM, 30000nM, 40000nM, 50000nM, 60000nM, 70000nM, 80000nM, 90000nM, or 100000nM (100μM), but is not limited to these values. In one embodiment, the concentration of the guide RNA contained in the initiation composition may be in the range between two values selected from the above values. In one embodiment, the concentration of the guide RNA contained in the initiation composition may be any one of the above values, or higher or lower. In a particular embodiment, the concentration of guide RNA contained in the initiation composition may be approximately 1000 nM (1 μM), 2000 nM, 3000 nM, 4000 nM, 5000 nM, 6000 nM, 7000 nM, 8000 nM, 9000 nM, or 10000 nM (10 μM).
[0229] In one embodiment, the initiation composition may contain a ribonucleoprotein (RNP) (e.g., a Cas / gRNA complex). In this case, the off-target prediction method of the present invention may further comprise mixing a guide RNA and a Cas protein and pre-incubating the mixture, so that the Cas protein and gRNA are present in the initiation composition in the form of an RNP. That is, the method may further comprise a process of incubating a mixture containing the guide RNA and a Cas protein before providing the initiation composition. For example, an RNP (Cas / gRNA complex) may be obtained from an incubated mixture containing the guide RNA and a Cas protein, and the obtained RNP may be mixed with cells to obtain the initiation composition. In one embodiment, the concentrations of RNP (e.g., Cas / gRNA complex) contained in the starting composition are approximately 10 nM, 20 nM, 30 nM, 40 nM, 50 nM, 60 nM, 70 nM, 80 nM, 90 nM, 100 nM, 200 nM, 300 nM, 400 nM, 500 nM, 600 nM, 700 nM, 800 nM, 900 nM, 1000 nM (1 μM), 2000 nM The RNP concentration may be, but is not limited to, M, 3000nM, 4000nM, 5000nM, 6000nM, 7000nM, 8000nM, 9000nM, 10000nM (10μM), 20000nM, 30000nM, 40000nM, 50000nM, 60000nM, 70000nM, 80000nM, 90000nM, or 100000nM (100μM). In one embodiment, the concentration of RNP contained in the initial composition may be in the range between two values selected from the above values. In one embodiment, the concentration of RNP contained in the initial composition may be any one of the above values, or higher or lower. In a particular embodiment, the concentration of RNP contained in the starting composition may be approximately 1000 nM (1 μM), 2000 nM, 3000 nM, 4000 nM, 5000 nM, 6000 nM, 7000 nM, 8000 nM, 9000 nM, or 10000 nM (10 μM).
[0230] Exemplary Embodiment (1) of the Off-Target Prediction System of the Present Application Hereinafter, exemplary (non-limiting) embodiments of the off-target prediction system of the present application are disclosed. The following embodiments may feature the mechanism of the off-target prediction system of the present application. Some or all of the following embodiments may include one or all of the embodiments disclosed in embodiments featuring the use of an extruder as described below.
[0231] One embodiment of the present application provides a method for predicting off-targets that may occur in a gene editing (e.g., genome editing) process. One embodiment of the present application provides a method for identifying off-target candidates that may occur in a genome editing process. One embodiment of the present application provides a method for predicting off-targets in a CRISPR / Cas gene editing system. One embodiment of the present application provides a method for identifying off-target candidates that may be generated in a gene editing process using a CRISPR / Cas gene editing system. The descriptions of methods for predicting off-targets that may be generated in a genome editing process or for identifying information about off-targets can be used without limitation.
[0232] One embodiment of the present application is (i) The step of obtaining a composition to be analyzed that contains cleaved genomic DNA; and (ii) Steps to obtain information about the rupture site by analyzing the composition to be analyzed. The present invention provides a method for predicting off-target events that may occur during a gene editing process.
[0233] In one particular embodiment, the cleaved genomic DNA contained in the composition to be analyzed may be cleaved genomic DNA formed by cleaving the genomic DNA of cells that have been physically destroyed by a gene editing system.
[0234] In a particular embodiment, cleaved genomic DNA may possess cell-specific epigenetic characteristics.
[0235] In one particular embodiment, the cleaved genomic DNA does not have to be repaired genomic DNA.
[0236] One embodiment of the present application is (i) A step of preparing an initiation composition comprising a gene editing tool and a first cell; (ii) A step of obtaining the composition to be analyzed by physically destroying the first cells, wherein the physical destruction of the first cells leads to the creation of an environment in which the genomic DNA in the cells can come into contact with the gene editing tool, thereby the genomic DNA and the gene editing tool come into contact, thereby the genomic DNA is cleaved at one or more cleavage sites; and (iii) Steps to obtain information about the cleavage site by analyzing the composition to be analyzed. The present invention provides a method for predicting off-target events that may occur in gene editing (e.g., genome editing) processes.
[0237] In one particular embodiment, a method for predicting off-target behavior is: (iv)(iii) The step of confirming information about off-target candidates from the information about the fracture site obtained in (iv)(iii) You may want to prepare even more.
[0238] One embodiment of the present application is (i) A step of preparing an initiation composition comprising a first Cas protein, a first guide RNA, and a first cell, wherein the Cas protein and the guide RNA are capable of forming a Cas / gRNA complex; (ii) A step of obtaining the composition to be analyzed by physically destroying the first cell, wherein by physically destroying the first cell, an environment is created in which the genomic DNA can come into contact with the Cas / gRNA complex, thereby bringing the genomic DNA and the Cas / gRNA complex into contact, and the genomic DNA is cleaved at one or more cleavage sites; and (iii) Step of obtaining information about the cleavage site by analyzing the composition to be analyzed. This invention provides a method for predicting off-target events that may occur in a CRISPR / Cas gene editing system.
[0239] In a particular embodiment, the method for predicting off-targets may further include a step of confirming information about off-target candidates from the information about the fracture site obtained in (iv)(iii).
[0240] In a particular embodiment, information about cleavage sites may include one or more of the following: the location of one or more cleavage sites on genomic DNA, a cleavage score for one or more cleavage sites, and the number of cleavage sites.
[0241] In one particular embodiment, the locations of one or more cleavage sites on the genomic DNA may be the locations of each of the one or more cleavage sites on the genomic DNA.
[0242] In a particular embodiment, the crack score for one or more crack sites may be the crack score for each of the one or more crack sites.
[0243] In one particular embodiment, the number of rupture sites may be the total number of rupture sites.
[0244] In a particular embodiment, information about off-target candidates may include one or more of the following: the location of one or more off-target candidates on genomic DNA; an off-target prediction score for one or more off-target candidates; and one or more of the predicted number of off-target candidates.
[0245] In a particular embodiment, the locations of one or more off-target candidate sites on genomic DNA may be the locations on each of the one or more off-target candidate sites on genomic DNA.
[0246] In a particular embodiment, the off-target prediction score for one or more off-target candidates may be the off-target prediction score for each of the one or more off-target candidates.
[0247] In one particular embodiment, the number of off-target candidates may be the total number of predicted off-target candidates.
[0248] In one particular embodiment, in (ii), the membrane structure including the cell membrane of the first cell may be destroyed by physically destroying the first cell, thereby providing an environment in which the Cas / gRNA complex can come into contact with the genomic DNA.
[0249] In one particular embodiment, in (ii), the membrane structure including the nuclear membrane of the first cell may be destroyed by physically destroying the first cell, thereby creating an environment in which the Cas / gRNA complex can come into contact with the genomic DNA.
[0250] In a particular embodiment, the physical destruction of the first cells may include passing the first cells through a filter having pores.
[0251] In a particular embodiment, the force that causes the first cell to pass through a pore smaller than the size of the first cell may be pressure.
[0252] In one particular embodiment, the average pore size of the filter may be smaller than the size of the first cell.
[0253] In a particular embodiment, the filter may have pores having a diameter smaller than the size of the first cell.
[0254] In a particular embodiment, the average pore size of the filter may be 5 to 15 μm.
[0255] In one particular embodiment, the average pore size of the filter may be approximately 8 μm.
[0256] In a particular embodiment, the filter may have pores having a diameter of 5 to 15 μm.
[0257] In a particular embodiment, the physical destruction of the first cells may be carried out by the use of an extruder.
[0258] In a particular embodiment, the extruder may be equipped with a filter having pores, the filter may have pores having a diameter smaller than the first cell size.
[0259] In a particular embodiment, the extruder may be equipped with a filter having pores, the average pore size of the filter may be 5 to 15 μm.
[0260] In one particular embodiment, genomic DNA exposed by physical disruption of a first cell may retain the cell-specific epigenetic features of the first cell (e.g., chromatin structure features).
[0261] In a particular embodiment, the information about off-target candidates may be information that reflects the epigenetic characteristics specific to the first cell.
[0262] In one particular embodiment, the environment in which the genomic DNA and the Cas / gRNA complex can come into contact may be an environment in which the DNA repair mechanism is inactivated.
[0263] In one particular embodiment, cell destruction inactivates the cell's DNA repair mechanism, and therefore, the cleaved DNA cannot be repaired.
[0264] Certain embodiments may provide a method for predicting off-target effects, further comprising confirming a predetermined CRISPR / Cas gene editing system that is subject to off-target prediction. Here, the predetermined CRISPR / Cas gene editing system may include any one or more of the following: the use of predetermined cells, the use of predetermined Cas proteins, and the use of predetermined guide RNA. Here, the confirmation of the predetermined CRISPR / Cas gene editing system may be performed before (i).
[0265] In one particular embodiment, the guide sequence of the first guide RNA may have the same sequence as the guide sequence of a predetermined guide RNA.
[0266] In a particular embodiment, a predetermined CRISPR / Cas gene editing system may include the use of predetermined cells, and the first cells and the predetermined cells may be the same.
[0267] In one particular embodiment, in (iii), the analysis of the composition to be analyzed may include analyzing the DNA contained in the composition to be analyzed by sequencing.
[0268] In a particular embodiment, in (iii), the analysis of the composition to be analyzed may include analyzing the cleaved genomic DNA contained in the composition to be analyzed by sequencing.
[0269] In one particular embodiment, in (iii), the analysis of the composition to be analyzed may include analyzing the DNA contained in the composition to be analyzed by a PCR-based analytical method.
[0270] In a particular embodiment, in (iii), the analysis of the composition to be analyzed may include analyzing the cleaved genomic DNA contained in the composition to be analyzed by a PCR-based analytical method.
[0271] In a particular embodiment, the concentration of Cas protein in the initial composition may be approximately 5000 nM.
[0272] In a particular embodiment, the concentration of the first cells contained in the initial composition is approximately 1 x 10 7 The value may be cells / mL.
[0273] In a particular embodiment, the acquisition of the composition to be analyzed may further include incubating the composition obtained through cell disruption.
[0274] In a particular embodiment, obtaining the composition to be analyzed may further include incubating the composition comprising disrupted cellular components, Cas protein, and guide RNA.
[0275] In a particular embodiment, the acquisition of the composition to be analyzed may further include removing RNA from the composition obtained by cell disruption.
[0276] In a particular embodiment, obtaining the composition to be analyzed may further include removing the RNA component from the composition, which includes disrupted cellular components, Cas proteins, and guide RNA.
[0277] In a particular embodiment, obtaining the composition to be analyzed may further include purifying DNA from the composition obtained by cell disruption.
[0278] In a particular embodiment, obtaining the composition to be analyzed may further include purifying DNA from a composition comprising disrupted cellular components, Cas protein, and guide RNA.
[0279] In a particular embodiment, the off-target prediction method of the present application may be used by combining one or more other off-target prediction methods, where the other off-target prediction methods may be any one or more selected from Cas-OFFinder, CHOPCHOP, CRISPOR, Digenome-seq, DIG-seq, SITE-Seq, CIRCLE-Seq, CHANGE-seq, GUIDE-seq, GUIDE-tag, DISCOVER-seq, BLISS, BLESS, integrase-deficient lentiviral vector-mediated DNA cleavage capture, HTGTS, ONE-seq, CReVIS-Seq, ITR-seq, and TAG-seq.
[0280] The following describes exemplary embodiments of the off-target prediction system of the present invention, characterized by the use of an extruder.
[0281] In one embodiment, a method for predicting off-target events that may occur in a CRISPR / Cas gene editing system may be provided, and the method is (i) A step of filling a first vessel of an extruder with an initiation composition comprising a first edited protein, a first guide RNA, and a first cell; (ii) Using an extruder, (a) A step of applying pressure to the first container to move the components of the starting composition from the first container of the extruder to the second container of the extruder, Here, the components of the initial composition are moved by the applied pressure from the first container to the second container through a filter with holes located between the first and second containers of the extruder, and the mixture is contained in the second container; Here, the first cells, which are larger than the diameter of the filter pores, are destroyed by the applied pressure, and at the same time, they pass through the filter pores. Here, through the physical disruption of the first cell, an environment is created in which genomic DNA in the cell can come into contact with the Cas / gRNA complex. As a result, genomic DNA comes into contact with the Cas / gRNA complex, Cleavage of genomic DNA occurs at one or more cleavage sites. A step of obtaining the composition to be analyzed by performing an extrusion process that includes the process; and (iii) The step of analyzing the composition to be analyzed to obtain information about the cleavage site. It is equipped with.
[0282] In one particular embodiment, the off-target prediction method is: (iv)(iii) This step involves confirming information about off-target candidates from the information about the cleavage site obtained from (iv)(iii) and predicting off-target events that may occur in the CRISPR / Cas gene editing system. You may want to prepare even more.
[0283] In a particular embodiment, the extrusion process in (ii) is (a) A step of applying pressure to the first container to move the components of the starting composition from the first container of the extruder to the second container of the extruder, Here, the components of the initial composition are moved by the applied pressure from the first container to the second container through a filter with holes located between the first and second containers of the extruder, and the mixture is contained in the second container. (b) The step of applying pressure to the second container to move the components of the mixture contained in the second container from the second container to the first container. Here, the components of the mixture contained in the second container move from the second container to the first container through a filter with holes located between the first and second containers due to the applied pressure, and consequently, the mixture that moves from the second container through the filter to the first container due to the pressure settles in the first container, and (c) The step of repeating the processes of (a) and (b) a predetermined number of times, Here, the predetermined number of times is counted in increments of 0.5, where 0.5 indicates the execution of a single process of (a) or (b). Here, the first cells, which are larger than the pore size of the filter, are destroyed by the applied pressure and, at the same time, pass through the pores of the filter. Here, through the physical disruption of the first cell, an environment is created in which genomic DNA in the cell can come into contact with the Cas / gRNA complex. As a result, genomic DNA comes into contact with the Cas / gRNA complex, This results in cleavage of genomic DNA at one or more cleavage sites. It may be provided.
[0284] In a particular embodiment, the pressure applied to the first vessel may be generated through a process of pushing a piston designed to apply pressure to the first vessel in the direction toward the first vessel and filter.
[0285] In a particular embodiment, the pressure applied to the first vessel may be generated through a process of pushing a piston designed to apply pressure to the first vessel in the direction of the first vessel and filter, and the pressure applied to the second vessel may be generated through a process of pushing a piston designed to apply pressure to the second vessel in the direction of the second vessel and filter.
[0286] In a particular embodiment, information about the cleavage sites may include one or more of the following: the location of one or more cleavage sites on genomic DNA, a cleavage score for one or more cleavage sites, and the number of cleavage sites.
[0287] In a particular embodiment, information about off-target candidates may include one or more of the following: the location on genomic DNA for one or more off-target candidates, an off-target prediction score for one or more off-target candidates, and the number of predicted off-target candidates.
[0288] In one particular embodiment, a membrane structure including the cell membrane of a first cell may be destroyed by physically destroying the first cell, thereby creating an environment in which the Cas / gRNA complex can come into contact with the genomic DNA of the first cell.
[0289] In one particular embodiment, in (ii), physical disruption of the first cell can lead to disruption of membrane structures, including the nuclear membrane of the first cell, thereby creating an environment in which the Cas / gRNA complex can come into contact with the genomic DNA of the first cell.
[0290] In a particular embodiment, the filter may have pores with a diameter smaller than the first cell size.
[0291] In a particular embodiment, the average diameter of the filter pores may be 5 to 15 μm.
[0292] In one particular embodiment, the average diameter of the filter pores may be 8 μm.
[0293] In a particular embodiment, the predetermined number of times may be 4 to 7.
[0294] In one particular embodiment, the predetermined number of times may be 5.5.
[0295] In one particular embodiment, in (ii), genomic DNA exposed through the physical disruption of the first cell may retain the epigenetic characteristics unique to the first cell.
[0296] In a particular embodiment, the information about the cleavage site obtained in (iii) may be information that reflects the epigenetic features unique to the first cell.
[0297] In a particular embodiment, the information about off-target candidates obtained in (iv) may be information that reflects the epigenetic characteristics specific to the first cell.
[0298] In one particular embodiment, the cellular DNA repair mechanism may be disrupted as a result of cellular damage, and consequently, the cleaved DNA may not be repaired.
[0299] In one particular embodiment, the off-target prediction method is: The step of identifying a CRISPR / Cas gene editing system that is subject to off-target prediction, where the CRISPR / Cas gene editing system subject to prediction includes the use of a Cas protein subject to prediction and the use of a guide RNA subject to prediction. You may want to prepare even more.
[0300] In one particular embodiment, the off-target prediction method is: The step of verifying the CRISPR / Cas gene editing system that is subject to off-target prediction, where the CRISPR / Cas gene editing system subject to prediction includes the use of the Cas protein subject to prediction and the use of the guide RNA subject to prediction, and the verification of the CRISPR / Cas gene editing system subject to prediction is performed before (i). You may want to prepare even more.
[0301] In one particular embodiment, the guide sequence of the first guide RNA may have the same sequence as the guide sequence of the guide RNA to be predicted.
[0302] In a particular embodiment, the CRISPR / Cas gene editing system being predicted may include the use of cells being predicted, where the first cells and the cells being predicted may be the same.
[0303] In a particular embodiment, the analysis of the composition to be analyzed in (iii) may include analyzing the cleaved genomic DNA contained in the composition to be analyzed through sequencing.
[0304] In a particular embodiment, the analysis of the composition to be analyzed in (iii) may include analyzing the cleaved genomic DNA contained in the composition to be analyzed through sequencing.
[0305] In a particular embodiment, the analysis of the composition to be analyzed in (iii) may include analyzing the DNA contained in the composition to be analyzed by a PCR-based analytical method.
[0306] In a particular embodiment, the analysis of the composition to be analyzed in (iii) may include analyzing the cleaved genomic DNA contained in the composition to be analyzed by a PCR-based analytical method.
[0307] In one particular embodiment, the concentration of Cas protein in the initial composition may be 5000 nM.
[0308] In a particular embodiment, the concentration of the first cells contained in the starting composition is 1 x 10 7 The value may be cells / mL.
[0309] In a particular embodiment, a further incubation process may be performed on the composition comprising disrupted cellular components, Cas protein, and guide RNA in order to obtain the composition to be analyzed.
[0310] In one particular embodiment, a further incubation process of the composition obtained through cell disruption may be performed in order to obtain the composition to be analyzed.
[0311] In a particular embodiment, a process of removing the RNA component from the composition, which includes disrupted cellular components, Cas protein, and guide RNA, may be further performed in order to obtain the composition to be analyzed.
[0312] In a particular embodiment, a further process of removing RNA from the composition obtained by cell disruption may be performed in order to obtain the composition to be analyzed.
[0313] In a particular embodiment, a further process of purifying the DNA of a composition comprising disrupted cellular components, Cas protein, and guide RNA may be performed in order to obtain the composition to be analyzed.
[0314] In a particular embodiment, a further process of purifying DNA from the composition obtained by cell disruption may be performed in order to obtain the composition to be analyzed.
[0315] In a particular embodiment, the off-target prediction method of the present application may be used in combination with one or more different off-target prediction methods, where the different off-target prediction methods may be any one or more selected from Cas-OFFinder, CHOPCHOP, CRISPOR, Digenome-seq, DIG-seq, SITE-Seq, CIRCLE-Seq, CHANGE-seq, GUIDE-seq, GUIDE-tag, DISCOVER-seq, BLISS, BLESS, integrase-deficient lentiviral vector-mediated DNA cleavage capture, HTGTS, ONE-seq, CReVIS-Seq, ITR-seq, and TAG-seq.
[0316] Exemplary Embodiment (2) of the Off-Target Prediction System of the Present Application Hereinafter, exemplary embodiments (non-limiting embodiments) are disclosed through a different explanatory format than the "exemplary embodiment (1) of the off-target prediction system of the present application" described above. Exemplary Embodiments Featuring a Mechanism A01. (i) A step of preparing an initiation composition comprising a first Cas protein, a first guide RNA, and a first cell, wherein the Cas protein and the first guide RNA are capable of forming a Cas / gRNA complex; (ii) Obtaining the composition to be analyzed by physically disrupting the first cell, wherein the physical disruption of the first cell prepares an environment in which genomic DNA can come into contact with the Cas / gRNA complex, thereby causing the genomic DNA and the Cas / gRNA complex to come into contact, thereby causing the genomic DNA to cleave at one or more cleavage sites; (iii) The step of obtaining information about the cleavage site by analyzing the composition to be analyzed; and (iv)(iii)Identify information about off-target candidates from the information about the cleavage site obtained from (iv)(iii) and predict off-target events that will occur in the CRISPR / Cas gene editing system. A method for predicting potential off-target events in a genome editing process using a CRISPR / Cas gene editing system, comprising the following features.
[0317] A02. Information about the site of the tear is, Location on genomic DNA of one or more cleavage sites, Destruction score for one or more destructive sites, and Number of rupture sites Method A01, which includes one or more of the following.
[0318] A03. Information about off-target candidates is available. Location on genomic DNA for one or more off-target candidates, Off-target prediction scores for one or more off-target candidates, and Number of predicted off-target candidates One of the methods A01 to A02, which includes one or more of the above.
[0319] A04 (ii) In this process, the membrane structure including the cell membrane of the first cell is destroyed by physically destroying the first cell, thereby creating an environment in which the Cas / gRNA complex can come into contact with the genomic DNA of the first cell. Choose one of the methods A01 through A03.
[0320] A05. (ii) In this process, the membrane structure including the nuclear membrane of the first cell is destroyed by physically destroying the first cell, thereby creating an environment in which the Cas / gRNA complex can come into contact with the genomic DNA of the first cell. Choose one of the methods from A01 to A04.
[0321] A06. Physically destroying the first cell comprises passing the first cell through a filter having pores smaller in size than the first cell, wherein the first cell is destroyed as it passes through pores smaller in size than the first cell, one of the methods A01 to A05.
[0322] A07. Physically destroying the first cells comprises passing a composition containing the first cells or a composition containing cellular components derived from the destroyed cells through a filter having pores that are half the size of the first cells or smaller, wherein the first cells are destroyed as they pass through pores that are smaller than the size of the first cells, one of the methods A01 to A06.
[0323] A08. The filter has pores with a diameter smaller than the size of the first cell, according to one of the methods A06 and A07.
[0324] A09. The average pore size of the filter is smaller than the first cell size, using one of the methods A06 to A08.
[0325] A10. The average pore size of the filter is 5 to 15 μm, using one of the methods A06 to A09.
[0326] A11. The average pore size of the filter is approximately 8 μm, using one of the methods A06 to A10.
[0327] A12. The first step is to physically destroy the cells, which is achieved by using an extruder, using one of the methods A01 to A05.
[0328] A13. The first cells are physically destroyed by the use of an extruder, and the filter contained in the extruder has pores with a diameter smaller than the size of the first cells, in one of the methods A01 to A05 and A12.
[0329] A14. The first method, A01 to A05, A12, and A13, is used to physically destroy the cells, and the average pore size of the filter contained in the extruder is 5 to 15 μm.
[0330] A15. (ii) The genomic DNA exposed by physically disrupting the first cell maintains the epigenetic features unique to the first cell in one of the methods A01 to A14.
[0331] A16. The information about the cleavage site obtained in (iii) is information that reflects the first cell-specific epigenetic features, one of the methods A01 to A15.
[0332] A17. (iv) The information about off-target candidates is one of the methods A01 to A16, which reflects the first cell-specific epigenetic features.
[0333] A18. The cellular DNA repair mechanism is disrupted by cell destruction, and therefore, the cleaved DNA is not repaired, one of the methods A01 to A17.
[0334] A19. The step of identifying a CRISPR / Cas gene editing system that is subject to off-target prediction, where the CRISPR / Cas gene editing system subject to prediction includes the use of a Cas protein subject to prediction and the use of a guide RNA subject to prediction. One of the methods A01 through A18 further includes the above.
[0335] A20. The step of verifying the CRISPR / Cas gene editing system that is subject to off-target prediction, where the CRISPR / Cas gene editing system subject to prediction includes the use of the Cas protein subject to prediction and the use of the guide RNA subject to prediction, and the verification of the CRISPR / Cas gene editing system subject to prediction is performed before (i). One of the methods A01 through A18 further includes the above.
[0336] A21. The guide sequence of the first guide RNA is one of the methods A19 to A20, having the same sequence as the guide sequence of the guide RNA to be predicted.
[0337] A22. The CRISPR / Cas gene editing system to be predicted involves the use of cells to be predicted, where the first cell and the cells to be predicted are the same, as in any one of the methods A19 to A21.
[0338] A23. (iii) The step of analyzing the composition to be analyzed is one of the methods A01 to A22, wherein the step of analyzing the DNA contained in the composition to be analyzed is performed by sequencing.
[0339] A24. (iii) The step of analyzing the composition to be analyzed is one of the methods A01 to A23, wherein the step of analyzing the cleaved genomic DNA contained in the composition to be analyzed is performed by sequencing.
[0340] A25. (iii) The step of analyzing the composition to be analyzed is one of the methods A01 to A22, which includes a step of analyzing the DNA contained in the composition to be analyzed through a PCR-based analytical method.
[0341] A26. (iii) The step of analyzing the composition to be analyzed is one of the methods A01 to A22 and A25, which includes a step of analyzing the genomic DNA contained in the composition to be analyzed through a PCR-based analytical method.
[0342] A27. The concentration of Cas protein in the starting composition is 5000 nM, according to one of the methods A01 to A26.
[0343] A28. The concentration of the first cells contained in the starting composition is 1 x 10⁻⁶. 7 One of the methods A01 to A27, where the concentration is cells / mL.
[0344] A29. The acquisition of the composition to be analyzed is Incubating a composition containing disrupted cellular components, Cas protein, and guide RNA. One of the methods A01 through A28, further including the above.
[0345] A30. The acquisition of the composition to be analyzed is Removal of RNA components from a composition containing destroyed cellular components, Cas protein, and guide RNA. One of the methods A01 through A29, further including the above.
[0346] A31. The acquisition of the composition to be analyzed is Purification of DNA from a composition containing disrupted cellular components, Cas protein, and guide RNA. One of the methods A01 through A30, further including the above.
[0347] Exemplary Embodiments Featuring the Use of an Extruder B01. (i) A step of filling a first container with an initiation composition comprising a first Cas protein, a first guide RNA, and a first cell; (ii) Using an extruder, (a) Applying pressure to the first container to move the components of the starting composition from the first container of the extruder to the second container of the extruder, Here, the components of the initial composition are moved by the applied pressure through a filter with holes located between the first and second extruders, from the first extruder to the second extruder, and thus the mixture settles into the second extruder; Here, the first cells, which are components larger than the diameter of the filter pores, are destroyed by the applied pressure and pass through the filter pores. Here, by physically disrupting the first cell, an environment is created in which genomic DNA can come into contact with the Cas protein and guide RNA. This causes the genomic DNA to come into contact with the Cas / gRNA complex, This causes genomic DNA to be cleaved at one or more cleavage sites. A step of obtaining the composition to be analyzed by performing an extrusion process that includes the process; and (iii) The step of analyzing the composition to be analyzed to obtain information about the cleavage site; and (iv)(iii)Identify information about off-target candidates from the information about the cleavage site obtained in (iv)(iii) and predict off-target events that may occur in the CRISPR / Cas gene editing system. A method for predicting potential off-target events in a genome editing process using a CRISPR / Cas gene editing system, comprising the following features.
[0348] B02. (ii) The extrusion process in (ii) is (a) A step of applying pressure to the first container to move the components of the starting composition from the first container of the extruder to the second container of the extruder, Here, the components of the initial composition are moved from the first vessel to the second vessel by the applied pressure, passing through a filter with holes located between the first and second vessels of the extruder, and thus the mixture settles into the second vessel. (b) The step of applying pressure to the second container to move the components of the mixture contained in the second container from the second container to the first container. Here, the components of the mixture contained in the second container move from the second container to the first container by the applied pressure through a filter with holes located between the first and second containers, and consequently, the mixture that has moved from the second container to the first container by passing through the filter due to pressure is contained in the first container, and (c) The step of repeating the processes of (a) and (b) a predetermined number of times, Here, the predetermined number of times is counted in increments of 0.5, where 0.5 represents the execution of a single process of (a) or (b), Here, the first cells, which are larger than the diameter of the filter pores, pass through the filter pores while being destroyed by the applied pressure. Here, through the physical disruption of the first cell, an environment is created in which genomic DNA in the cell can come into contact with the Cas / gRNA complex. This causes the genomic DNA to come into contact with the Cas / gRNA complex, This causes the genomic DNA to cleave at one or more cleavage sites. Method B01, which includes the process of...
[0349] B03. The pressure applied to the first vessel may be generated by any one of the methods B01 to B02, through a process of pushing a piston designed to apply pressure to the first vessel in the direction of the first vessel and filter.
[0350] B04. The method of B02, wherein the pressure applied to the first vessel may be generated through a process of pushing a piston designed to apply pressure to the first vessel in the direction of the first vessel and filter, and the pressure applied to the second vessel may be generated through a process of pushing a piston designed to apply pressure to the second vessel in the direction of the second vessel and filter.
[0351] B05. Information about the site of the tear is, Location on genomic DNA of one or more cleavage sites, Destruction score for one or more destructive sites, and Number of rupture sites One of the methods B01 to B04, which includes one or more of the above.
[0352] B06. Information about off-target candidates is available. Location on genomic DNA for one or more off-target candidates, Off-target prediction scores for one or more off-target candidates, and Number of predicted off-target candidates One of the methods B01 to B05, which includes one or more of the above.
[0353] B07. (ii) in any one of the methods B01 to B06, wherein the membrane structure including the cell membrane is destroyed through physical destruction of the first cell, thereby preparing an environment in which the Cas / gRNA complex can come into contact with the genomic DNA of the first cell.
[0354] B08. (ii) in any one of the methods B01 to B07, wherein the membrane structure including the nuclear membrane of the first cell is destroyed through physical destruction of the first cell, thereby preparing an environment in which the Cas / gRNA complex can come into contact with the genomic DNA of the first cell.
[0355] B09. The filter has pores with a diameter smaller than the first cell size, one of the methods B01 to B08.
[0356] B10. The average pore size of the filter is 5 to 15 μm, using one of the methods B01 to B09.
[0357] B11. The average pore size of the filter is 8 μm, using one of the methods B01 to B10.
[0358] B12. One of the methods B02 to B10, where the predetermined number of steps is between 4 and 7.
[0359] B13. The predetermined number of steps is 5.5, and one of the methods from B02 to B12.
[0360] B14. (ii) In one of the following ways, the genomic DNA exposed by the physical destruction of the first cell maintains the epigenetic features unique to the first cell: B01 to B13.
[0361] B15. The information about the cleavage site obtained in (iii) is information that reflects the first cell-specific epigenetic features, one of the methods B01 to B14.
[0362] B16. The information about off-target candidates obtained in (iv) is information that reflects the first cell-specific epigenetic features, one of the methods B01 to B15.
[0363] B17. One of the methods B01 to B16, in which the cell's DNA repair mechanism is disrupted due to cell damage, and as a result, the cleaved DNA is not repaired.
[0364] B18. The method is, The step of identifying a CRISPR / Cas gene editing system that is subject to off-target prediction, where the CRISPR / Cas gene editing system subject to prediction includes the use of a Cas protein subject to prediction and the use of a guide RNA subject to prediction. One of the methods B01 through B17 further provides the following:
[0365] B19. The method involves a step of verifying a CRISPR / Cas gene editing system that is subject to off-target prediction, where the CRISPR / Cas gene editing system subject to prediction includes the use of a Cas protein subject to prediction and the use of a guide RNA subject to prediction, and the verification of the CRISPR / Cas gene editing system subject to prediction is performed before (i). One of the methods B01 through B17 further provides the following:
[0366] B20. The guide sequence of the first guide RNA is one of the B18 and B19 methods, having the same sequence as the guide sequence of the guide RNA to be predicted.
[0367] B21. The CRISPR / Cas gene editing system to be predicted involves the use of cells to be predicted, where the first cell and the cells to be predicted are the same, one of the methods from B18 to B20.
[0368] B22. (iii) Analysis of the composition to be analyzed includes analyzing the cleaved genomic DNA contained in the composition to be analyzed by sequencing, One of the methods from B01 to B21.
[0369] B23. (iii) Analyzing the composition to be analyzed includes analyzing the cleaved genomic DNA contained in the composition to be analyzed through sequencing, One of the methods from B01 to B22.
[0370] B24. (iii) Analyzing the composition to be analyzed includes analyzing the DNA contained in the composition to be analyzed by PCR-based analytical methods, One of the methods from B01 to B21.
[0371] B25. (iii)(iii)The analysis of the composition to be analyzed in (iii) is to analyze the cleaved genomic DNA contained in the composition to be analyzed by a PCR-based analytical method. One of the methods from B01 to B21.
[0372] B26. The concentration of Cas protein in the starting composition is 5000 nM, according to one of the methods B01 to B25.
[0373] B27. The concentration of the first cells contained in the starting composition is 1 x 10⁻⁶. 7 One of the methods B01 to B26, where the concentration is cells / mL.
[0374] B28. The method involves obtaining the composition to be analyzed, Step 1: Incubating a composition containing disrupted cellular components, Cas protein, and guide RNA. One of the methods B01 through B27 further enhances the process.
[0375] B29. The method involves obtaining the composition to be analyzed, Steps to remove RNA components from a composition containing destroyed cellular components, Cas protein, and guide RNA. One of the methods B01 through B28 further enhances the process.
[0376] B30. The method involves obtaining the composition to be analyzed, A step of purifying the DNA of a composition containing disrupted cellular components, Cas protein, and guide RNA. One of the methods from B01 to B29 further enhances the process.
[0377] Expected use of the off-target prediction system of this invention The following examples of anticipated uses of the off-target prediction system of this application (conceptual scenarios in which a person skilled in the art would use the off-target prediction system of this application) will be described without limitation. The off-target prediction system of this application (e.g., Extru-seq) is an off-target prediction system characterized by the physical disruption of cells and is a more efficient and more accurate off-target prediction system that has the advantages of conventional in vitro-based off-target prediction systems and in vivo-based off-target prediction systems. Accordingly, all methods for identifying off-target candidates or predicting off-targets performed to achieve the objective of identifying off-targets that may occur in a gene editing process using the above-described characteristics of the off-target prediction system are included as one embodiment of the use or application of the off-target prediction method of this application, but the following examples are not intended to limit the scope of this application.
[0378] For example, the off-target prediction method (or system) of the present invention may be used by technicians or researchers who use CRISPR / Cas gene editing systems to edit cell genomes.
[0379] For example, a researcher selects a gene editing system to be used in cell genome editing. For example, a researcher selects the CRISPR / Cas gene editing system as the gene editing system to be used in cell genome editing. Furthermore, the researcher may select the cells that are the primary target of genome editing. In the process of selecting the gene editing system to be used for cell genome editing, an in silico-based off-target prediction method may be used to design an appropriate guide sequence. Here, the researcher will develop a treatment that involves the use of the gene editing system. In developing the treatment, information about off-targets of the selected gene editing system (in particular, the guide RNA) must be confirmed. Based on the selected gene editing system, the details of the off-target prediction method of this application were designed to suit the purpose. By performing the off-target prediction method of this application, information about off-target candidates that may occur with the use of the selected gene editing system was confirmed. Subsequently, using the confirmed information about off-target candidates, information about off-targets that would be problematic with the use of the selected gene editing system was confirmed. In particular, true off-targets were finally confirmed by verifying the candidate off-target sites identified by the off-target prediction method of this application. In this process, known off-target prediction methods (in silico, in vitro, and cell-based off-target prediction methods) may be used in combination to discover true off-target sites.
[0380] In another example, the off-target prediction system of the present invention may be used in the selection process of a gene editing system (particularly the guide sequence of the guide RNA). Researchers generate a guide RNA library containing various types of guide RNAs. The off-target prediction method is performed on a gene editing system containing one or more guide RNAs included in the guide RNA library. Subsequently, a gene editing system to be used in the development or study of a treatment is selected based on the results of the off-target prediction method of the present invention. In this process, known off-target prediction methods (in silico, in vitro, and cell-based off-target prediction methods) may be used in combination to discover true off-target sites.
[0381] As described above, the off-target prediction system of this application can be used in a variety of situations, and the manner in which the off-target prediction system is used is not limited to the examples described above. [Examples]
[0382] The present invention provided herein will be described in more detail below with reference to experimental examples or embodiments. These experimental examples are provided solely to illustrate the content disclosed herein, and it will be apparent to those skilled in the art that the scope of the present invention disclosed herein is not construed as being limited to the following experimental examples.
[0383] Experimental example Experimental method Experimental method 1. Design of promiscuous sgRNA Candidate target sequences containing protospacer-adjacent motifs (NGG PAMs) located in the mouse genome (mm10) PCSK9 and albumin genes were extracted using Cas-Designer (see reference [Park, Jeongbin, Sangsu Bae, and Jin-Soo Kim. Cas-Designer: a web-based tool for choice of CRISPR-Cas9 target sites. Bioinformatics 31.24(2015):4014-4016.]). The extracted sequences were aligned to the human genome (hg19). From the extracted sequences, sequences with one or more targets having zero mismatches when aligned to the human genome were selected. The selected candidates were analyzed using Cas-OFFinder (see reference [Bae, Sangsu, Jeongbin Park, and Jin-Soo Kim. Cas-OFFinder: a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 30.10(2014):1473-1475.]). Candidates with various sets of associated sequences containing varying numbers of mismatches (0 to 5 mismatches per site) widely distributed throughout the human and mouse genomes were selected as targets. Information on the indiscriminate sgRNA targets and guide sequences used in subsequent experiments is as follows.
[0384] Target sequences of single-stranded guide RNA (sgRNA) targeting mouse PCSK9 (excluding NGG PAM): AGGTGGGAAACTGAGGCTT (Sequence ID: 25)
[0385] Target sequences of sgRNAs targeting mouse albumin (excluding NGG PAM): ACATGCATATGTATGTGTG(Sequence code: 26)
[0386] As explained below in the experimental results, these sgRNAs perfectly matched target sequences present in the human genome (however, the PCSK9 and albumin loci are not targets in the human genome). The target and guide sequences targeted loci other than PCSK9 or albumin in the human genome, but for convenience, they are referred to as human PCSK9-targeting sgRNAs (sgRNAs that target human PCSK9) and human albumin-targeting sgRNAs (sgRNAs that target human albumin).
[0387] In other words, the target sequence of the indiscriminate sgRNA, referred to as the human PCSK9 targeting sgRNA, is the same as the target sequence of the mouse PCSK9 targeting sgRNA (sgRNA that targets mouse PCSK9). The target sequence of the human PCSK9 targeting sgRNA (excluding NGG PAM) is as follows: AGGTGGGGAAACTGAGGCTT (Sequence ID: 25).
[0388] The target sequence of indiscriminate sgRNA, also known as human albumin-targeting sgRNA, is the same as the target sequence of mouse albumin-targeting sgRNA (sgRNA that targets mouse albumin). The target sequence of human albumin-targeting sgRNA (excluding NGG PAM) is as follows: ACATGCATATGTATGTGTG (Sequence ID: 26).
[0389] Experimental method 2. Plasmid construction for sgRNA and Cas9 expression The Streptococcus pyogenes Cas9 sequence (see reference [Cho, Seung Woo, et al. "Targeted genome engineering in human cells with the Cas9 RNA-guided endonuclease." Nature biotechnology 31.3(2013):230-232.]) and designed indiscriminate sgRNA sequences (albumin-targeting sgRNA and PCSK9-targeting sgRNA sequences) were cloned into the backbone of the AAV plasmid used in previous studies (see reference [Kim, Eunji, et al. "In vivo genome editing with a small Cas9 orthologue derived from Campylobacter jejuni." Nature communications 8.1(2017):1-12.]) to generate Cas9 (pAAV-Cas9) and sgRNA (pAAV-albumin and pAAV-PCSK9) expression vectors. Cas9 expression was performed under the control of the CMV promoter, and sgRNA expression was performed under the control of the U6 promoter. Guide sequences targeting the FANCF, VEGFA, and HBB genes were cloned in the pRG2 vector (Addgene #104174).
[0390] Experimental method 3. GUIDE-seq Human HEK293T cells (ATCC, Cat#CRL-3216) and mouse NIH-3T3 cells (ATCC, Cat#CRL-1658) were maintained in Dulbecco's modified Eagle medium (DMEM) supplemented with 10% fetal bovine serum (FBS) and 1% penicillin-streptomycin under conditions of 5% CO2 and 37°C. HEK293T and NIH3T3 cells were subcultured every 72 hours to maintain 80% confluence. GUIDE-seq results were obtained at 2x10⁻⁶. 5HEK293T cells were transfected with lipofectamine 2000 using sgRNA expression plasmids (500 ng, pAAV-albumin or pAAV-PCSK9), Cas9 expression plasmids (500 ng, p3s-Cas9HC; Addgene plasmid #43945), and 5 pmol dsODN. 2x10 5 NIH-3T3 cells were transfected with sgRNA expression plasmid (250 ng, pAAV-albumin or pAAV-PCSK), Cas9 expression plasmid (500 ng, p3s-Cas9HC; Addgene plasmid #43945), and 100 pmol dsODN using the Amaxa P3 electroporation kit (V4XP-3032; program EN-158). Transfected cells were transferred to 24-well plates containing pre-cultured DMEM (1 mL / well) at 37°C. After 72 hours, genomic DNA was isolated using the QIAamp DNA mini-kit (Qiagen).
[0391] Human HeLa cells (ATCC, Cat#CCL-2) were maintained in DMEM supplemented with 10% FBS and 1% penicillin-streptomycin under conditions of 5% CO2 and 37°C. HeLa cells were subcultured every 72 hours to maintain 80% confluence. GUIDE-seq results were obtained at 2x10⁻⁶. 5 HeLa cells were transfected with sgRNA expression plasmids (500 ng, pRG2-FANCF, pRG2-VEGFA, or pRG2-HBB), Cas9 expression plasmid (500 ng, p3s-Cas9HC; Addgene plasmid #43945), and 25 pmol dsODN using Amaxa 4D-nucleofector (V4XC-1024; program CN-114). Transfected cells were transferred to 24-well plates (1 mL / well) containing DMEM pre-cultured at 37°C. After 72 hours, genomic DNA was isolated using the QIAamp DNA mini-kit (Qiagen).
[0392] 1000 nm purified DNA was fragmented using a Covaris system (duty factor: 10%, PIP: 50, cycles per burst: 200, time: 50 seconds, temperature: 20°C) and purified using Ampure XP beads (A63881). A sequencing library was generated from the DNA using the Illumina NEBNext® Ultra® II DNA Library Prep Kit (E7546L) according to the manufacturer's protocol. Subsequently, regions of the library containing dsODN sequences were amplified using dsODN-specific primers and sequenced using MiSeq (Illumina, TruSeq HT kit). Other procedures were the same as those described in previous studies (see reference [Tsai, Shengdar Q., et al. "GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases." Nature biotechnology 33.2(2015):187-197.]). For data analysis, GUIDE-seq (1.0.2; https: / / pypi.org / project / guide-seq / ), which is compatible with Python® 3, was used.
[0393] Experimental method 4. Plasmid construction and in vitro transcription reaction related to sgRNA transcription To improve the yield and accuracy of sgRNA transcription, the inventors of this invention modified the previously described method (see [Kim, Daesik, Beum-Chang Kang, and Jin-Soo Kim. Identifying genome-wide off-target sites of CRISPR RNA-guided nucleases and deaminases with Digenome-seq. Nature Protocols 16.2(2021):1170-1192.]). Briefly, an sgRNA template was generated by annealing two complementary oligonucleotides followed by PCR amplification. BamHI, BsaI, and KpnI restriction sites were attached to the ends of the sgRNA template through a second PCR. The tailed sgRNA template was inserted into a pUC19 plasmid digested with BamHI and KpnI. The sgRNA-encoding plasmid was linearized with Bsal, resulting in the generation of a suitable sgRNA terminal sequence. The linearized plasmids were incubated at 37°C for 8 hours with 7.5 U / μl T7 RNA polymerase (NEB, M0251L) in reaction butter (NEB, B9012S) containing 14 mM MgCl2 (NEB, B0510A), 10 mM DTT (Sigma, 43816), 0.02 U / μl yeast inorganic pyrophosphatase (NEB, M2403L), 1 U / μl mouse RNase inhibitor (NEB, M0314L), 4 mM ATP (NEB, N0451AA), 4 mM GTP (NEB, N0452AA), 4 mM UTP (NEB, N0453AA), and 4 mM CTP (NEB, N0454AA). Yeast inorganic phosphatase was included to enhance sgRNA synthesis. After the reaction, the mixture was mixed with DNase I and incubated to remove the DNA template; then, the transcribed sgRNA was purified using a PCR purification kit (Favorgen, #FAGCK001-1).
[0394] Experimental method 5. Digenome-seq Genomic DNA was purified from HEK293T cells (ATCC, Cat#CRL-3216) and NIH-3T3 cells (ATCC, Cat#CRL-1658) using the DNeasy Blood & Tissue Kit (Qiagen). Two types of genomic DNA (10 μg each) were incubated with Cas9 protein (10 μg) and albumin or PCSK9 targeting sgRNA (10 μg each) in 1 mL of a reaction solution containing NEB3 buffer [100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / mL bovine serum albumin (BSA), pH 7.9] for 8 hours at 37°C. Digested genomic DNA was treated with RNase A (50 μg / mL, Qiagen) for 10 minutes to degrade the sgRNA, and then purified using the DNeasy Blood & Tissue Kit (Qiagen).
[0395] Genomic DNA (1 μg) was fragmented into 300 bp segments using the Covaris system (Life Technologies) and blunt-ended using End Repair Mix (Thermo Fischer). The fragmented DNA was ligated to an adapter to generate a library, which was then subjected to WGS using a HiSeq X Ten Sequencer (Illumina) in Macrogen. WGS was performed at sequencing depths of 30 to 40x. DNA cleavage sites were identified using the Digenome 1.0 program (see reference [Park, Jeongbin, et al. "Digenome-seq web tool for profiling CRISPR specificity." Nature methods 14.6(2017):548-549.]).
[0396] Experimental method 6. In silico prediction of off-target sites Using Cas-OFFinder, we obtained genome-wide candidate off-target sites in hg19 with fewer than 7 mismatches with selected sgRNAs. Based on a previous paper ([Liu, Qiaoyue, et al. "Deep learning improves the ability of sgRNA off-target propensity prediction." BMC bioinformatics 21.1(2020):1-15.]), we used the CROP prediction model and optimization parameters (https: / / github.com / vaprilyanto / crop) to calculate the CROP score (a heuristic score indicating whether candidate off-target sites are edited). The CFD score (percent activity value provided in a penalty matrix based on each possible type of mismatch at each position in the guide RNA sequence) was calculated using the "crisprScore"R package (see reference [Doench, John G., et al. "Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9." Nature biotechnology 34.2(2016):184-191.]). For the two calculations, GX19 (GACATGCATATGTATGTGTG (SEQ ID NO: 27)) was used for albumin and GAGGTGGGAAACTGAGGCTT (SEQ ID NO: 28) was used for the PCSK9 sgRNA sequence and X20 target sequence.
[0397] Experimental method 7. Extru-seq To prepare extru-seq, transcribed sgRNA was refolded in 1X NEBuffer 3.1 reaction buffer (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / mL BSA, pH 7.9). The sgRNA was heated at 98°C for 2 minutes, and then the temperature was lowered at a rate of 0.1°C / second until it reached 20°C. To reduce reaction inhibition by high concentrations of glycerol, Cas9 buffer (10 mM Tris-HCl, 0.15 M NaCl, 50% glycerol, pH 7.4) was replaced with elution buffer (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, pH 8.0). Buffer exchange was performed using a 10K Amicon® Ultra-15 centrifuge filter (Millipore).
[0398] HEK293T (ATCC, Cat#CRL-3216), NIH-3T3 (ATCC, Cat#CRL-1658), and HeLa (ATCC, Cat#CCL-2) cells of each type were collected using 0.25% trypsin-EDTA, and human bone marrow mesenchymal stem cells (BM-MSCs; Lonza, Cat#PT-2501) were collected using 0.05% trypsin-EDTA. The collected cells were resuspended in Dulbecco's phosphate-buffered saline (PBS). Buffer-exchanged Cas9 (800 mg) and refolded sgRNA (530 μg) were pre-cultured at room temperature for 10 minutes to form RNP complexes (for multiplex extru-seq, buffer-exchanged Cas9 (800 mg) and five different refolded sgRNAs (106 μg each) were used). 1 x 10 7The cells were mixed with a 5000 nM RNP complex in 1 mL of 1X NEBuffer 3.1 reaction buffer (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 100 μg / mL BSA, pH 7.9). SCR7 pyrazine (Sigma, SML1546) (1 μM) was added to perform extru-seq in the presence of SCR7. After gentle pipetting, the suspended cells were extruded 11 times through an 8 μm pore size polycarbonate membrane filter (Whatman) using a mini extruder (Avanti Polar Lipids). The extruded samples were incubated at 37°C for 16 hours. RNase A (2 mg / mL) was added to remove sgRNA and RNA, and then genomic DNA was purified from the extruded samples using the FavorPrep Blood Genomic DNA Extraction mini kit (Favorgen, #FAGCK001-2). Whole-genome sequencing (WGS) was performed at sequencing depths of 30 to 40x. WGS was performed in Macrogen using Nova-seq instruments according to the manufacturer's standard protocol. DNA cleavage sites were identified using the Digenome-seq standalone program (http: / / www.rgenome.net / digenome-js / standalone). The analysis filtering options were a minimum depth of 10, a minimum score of 0.05, and a minimum ratio of 0.01; other options were left at their default settings. Furthermore, as developers of a new tool, the inventors of this application checked all sites identified by extru-seq using Integrative Genomic Viewer (IGV). Several loci appeared to be false-positive candidates (i.e., uncleaved sites by IGV (see Figures 117 to 125)). False positives (uncleaved sites) identified using IGV were treated as negative and excluded from the analysis. Figures 117 to 125 show false-positive off-target sites manually excluded from Digenome-seq and Extru-seq WGS data. They were visualized using IGV.These false positives were observed in Digenome-seq. The relevant BAM file is available under accession number PRJNA796642 in the NCBI Bioproject (https: / / www.ncbi.nlm.nih.gov / bioproject / ).
[0399] Referring to Figures 117 to 119, the sequence ACtTGtgTgTGTgTGTGgGGGG (sequence number: 49) is disclosed. In these figures, mismatches are indicated in lowercase letters, and bulges (if present) are indicated by dashes. In these figures, the PAM sequence is underlined.
[0400] Referring to Figures 120 to 122, the sequence AtATATATATaTATaTaTGGAG (sequence number: 50) (bulge-related markings are omitted) is disclosed. In these figures, mismatches are indicated in lowercase letters, and bulges (if present) are indicated by dashes. In these figures, the PAM sequence is underlined.
[0401] Referring to Figures 123 to 125, the sequence TAgATATATATGaATGgGTaGAG (sequence number: 51) (bulge-related markings are omitted) is disclosed. In these figures, mismatches are indicated in lowercase letters, and bulges (if present) are indicated by dashes. In these figures, the PAM sequence is underlined.
[0402] Experimental method 8. Unlike GUIDE-seq and Cas-OFFinder, which assign non-target results from Digenome-seq and Extru-seq to Cas-OFFinder results, the standalone Digenome-seq program does not have an sgRNA:off-target alignment function that provides information about the number of mismatches and bulge type (DNA or RNA) between guide and off-target sites. The web version of the Digenome-seq analysis tool (http: / / www.rgenome.net / digenome-js / #!) has an optional alignment function with an alignment score that does not provide any information about the number of mismatches or bulge type. Instead, the inventors of this application used Cas-OFFinder to identify off-target sites with up to 7 mismatches and 2 bulges related to the target sequence. The locations of off-target candidate sites identified by Digenome-seq and Extru-seq were compared with those identified by Cas-OFFinder. Information on mismatches and bulge types obtained from Cas-OFFinder could be assigned to loci identified by Digenome-seq and Extru-seq.
[0403] Experimental method 9. Efficacy confirmation of candidate off-target sites using human cell lines Human HEK293T and HeLa cells were maintained at 37°C in the presence of 5% CO2 in DMEM supplemented with 10% FBS (ATCC, CRL-3216) and 1% penicillin-streptomycin. To determine the indel frequency at candidate off-target sites, 2 x 10⁶ cells were used. 5 HEK293T cells and 8x10 4Each HeLa cell was transfected with sgRNA expression plasmids (500 ng, pAAV-albumin, pAAV-PCSK9, pRG2-HBB, pRG2-FANCF, or pRG2-VEGFA) and Cas9 expression plasmids (500 ng, pAAV-Cas9 or p3s-Cas9HC; Addgene plasmid #43945) using lipofectamine 2000 (vendor, quantity). After incubating the cells at 37°C for 3 days, genomic DNA was prepared using the FavorPrep Blood Genomic DNA Extraction Mini Kit (Favorgen, #FAGCK001-2). Target sites and potential off-target sites were then analyzed using deep sequencing. Deep sequencing libraries were generated by PCR. Each sample was labeled using TruSeq HT dual-index primers. Paired-end sequencing was performed on the pooled libraries using MiSeq (Illumina). In particular, multiple targets were combined by PCR performed using primers with different indices, and then subjected to deep sequencing analysis.
[0404] Deep sequencing data is available under accession number PRJNA796642 in the NCBI Bioproject (https: / / www.ncbi.nlm.nih.gov / bioproject / ). To determine whether a target was confirmed effective or false, the inventors of this application used the following criteria used in EDITAS Medicine (see reference [Maeder, Morgan L., et al. "Development of a gene-editing approach to restore vision loss in Leber congenital amaurosis type 10." Nature medicine 25.2(2019):229-233.]): First, the indel concentration in the sample must be higher than 0.1% for it to be confirmed effective. Second, the treated / control ratio must be higher than 2. The efficacy confirmation results for off-target candidates via deep sequencing are detailed in Table 1. Mismatches between the target and the confirmed effective off-target candidates are shown in lowercase. Regarding human PCSK9 in Table 1, sequence numbers 74 to 132 were assigned to off-target sequences 1 to 59 disclosed in human PCSK9 in the order of their disclosure. Regarding human albumin in Table 1, sequence numbers 133 to 174 were assigned to off-target sequences 1 to 42 disclosed in human albumin in the order of their disclosure. Regarding mouse PCSK9 in Table 1, sequence numbers 175 to 211 were assigned to off-target sequences 1 to 37 disclosed in mouse PCSK9 in the order of their disclosure. Regarding mouse albumin in Table 1, sequence numbers 212 to 249 were assigned to off-target sequences 1 to 38 disclosed in mouse albumin in the order of their disclosure. Regarding human HBB in Table 1, sequence numbers 250 to 293 were assigned to off-target sequences 1 to 44 disclosed in human HBB in the order of their disclosure.Regarding human VEGFA in Table 1, sequence numbers 294 to 343 were assigned to off-target sequences 1 to 50 disclosed in human VEGFA, in the order of their disclosure. Regarding human FANCF in Table 1, sequence numbers 344 to 383 were assigned to off-target sequences 1 to 40 disclosed in human FANCF, in the order of their disclosure. Table 1 is shown below.
[0405] Table 1. Target deep sequencing results for off-target effectiveness verification. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10] [Table 1-11] [Table 1-12] [Table 1-13] [Table 1-14] [Table 1-15] [Table 1-16] [Table 1-17] [Table 1-18] [Table 1-19] [Table 1-20] [Table 1-21]
[0406] Experimental method 10. AAV generation AAV8 carrying the desired clonal sequences (pAAV-PCSK9, pAAV-albumin, and pAAV-Cas9) was generated on a large scale by VigeneBioscience (10 13 Genome copy (GC) / mL). The generated AAV was distributed and stored at -70°C until use.
[0407] Experimental Method 11. Animal Research All animal experiments were approved by the Animal Care Committee (IACUC) at Yonsei University College of Medicine (IACUC number 2019-0215). C57BL / 6 mice were maintained in a 12:12 light-dark cycle.
[0408] Experimental Method 12. Two types of AAV8, each carrying pAAV-Cas9 and one of two pAAV-sgRNAs (pAAV-PCSK9 or pAAV-albumin), were delivered into C57BL / 6 mice via systemic (intravenous) and subretinal injection. Both injections were performed in a 1:1 GC (pAAV-Cas9:pAAV-sgRNA) ratio. The dose for intravenous injection was 2.5 x 10⁻⁶. 11 GC / animal, the dosage for subretinal injection is 1.5 x 10 10 It was GC / eye.
[0409] For systemic injection, 200 μL of AVV8 diluted in PBS was injected into 7- to 9-week-old male mice via tail vein injection. The dose was 2.5 x 10⁻⁶. 11 It was a GC AAV8.
[0410] For subretinal injection, male mice aged 7 to 9 weeks were selected. Under general anesthesia, one pupil per mouse was dilated with eye drops containing tropicamide and phenylfrine. During the experiment, the mice's body temperature was maintained at 37°C using a heating pad. A small incision was made 1 mm away from the corneal margin using a 1 / 2 30G needle. A Hamilton syringe with a 33G blunt needle, filled with 2 μL of a solution containing the AAV8 mixture, was inserted through the incision to the point of resistance (subretinal region). To avoid unnecessary tissue damage, the volume was injected carefully and gently, and after waiting 20 to 30 seconds for uniform diffusion, the syringe was slowly removed. Subsequently, an antibiotic ointment was applied to the surface of the eye. Four mice were used for different injection methods and different sgRNAs.
[0411] Experimental method 13. DNA preparation from collected organs and tissues Organs and tissues were removed 3 months and 2 weeks after injection. At the end of the experiment, the animals were euthanized by cardiac puncture under isoflurane anesthesia. Organs including the eyes, liver, spleen, lungs, kidneys, muscles, brain, and testicles were excised, flash-frozen in liquid nitrogen, and stored at -70°C until further analysis.
[0412] For subretinal injection, the neuroretina and retinal pigment epithelium (RPE) were isolated and prepared. The cornea, iris, lens, and vitreous humor were removed from the excised eye. The remaining ocular tissue was incubated in hyaluronidase solution for 45 minutes (37°C, 5% CO2). It was then incubated in cold PBS for 30 minutes to inactivate hyaluronidase activity. Next, the ocular tissue was transferred to fresh PBS, and the neuroretina was gently separated from the retina / RPE / choroid / scleral complex. The remaining retina / RPE / choroid / scleral complex was incubated in trypsin solution in 5% CO2 at 37°C for 45 minutes, gently shaken until the RPE sheet was completely detached. All isolated RPE sheets and RPE cells were collected. Genomic DNA was extracted using the DNeasy Blood & Tissue Kit (Qiagen, catalog number 69506) according to the manufacturer's instructions.
[0413] Experimental Method 14. Targeted Deep Sequencing Genomic DNA of mouse retinal pigment epithelium (RPE) cells was amplified using the REPLI-g single-cell kit (Qiagen) according to the manufacturer's protocol.
[0414] Target and potential off-target sites were analyzed through targeted deep sequencing. Deep sequencing libraries were generated by PCR. TruSeq HT dual-index primers were used to label each sample. Pair-end sequencing was performed on the pooled libraries using MiSeq (Illumina). In particular, multiple targets were combined by PCR performed with different indexed primers, followed by targeted deep sequencing analysis.
[0415] Experimental Method 15. Statistical Analysis The score / sequence read count was min-max normalized. In each group, the maximum value was normalized to 1 and the minimum value to 0. To test whether the score medians of two different groups were the same or not, a Wilcoxon rank-sum test was performed on samples that fell into each intersection of the Venn diagram. The results of the two-tailed, unpaired Mann-Whitney test, calculated by Prism (version 9.4.1), are shown.
[0416] Nucleic acid sequence of sgRNA and amino acid sequence of SpCas9 used in the experiment The following shows the sequences of sgRNAs used in this specification, their related sequences, and SpCas9 sequences. As mentioned above, human PCSK9 targeting sgRNAs actually target different gene loci other than PCSK9 on the human genome, but for convenience they are referred to as human PCSK9 targeting sgRNAs. Human albumin targeting sgRNAs actually target different gene loci other than albumin on the human genome, but for convenience they are referred to as human albumin targeting sgRNAs.
[0417] Total sequence of mouse PCSK9 targeting sgRNA GAGGUGGGAAACUGAGGCUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU(Sequence ID: 29)
[0418] Guide sequence of mouse PCSK9 targeting sgRNA GAGGUGGGAAACUGAGGCUU (Sequence ID: 30)
[0419] Target sequences of mouse PCSK9 targeting sgRNA (target sequences on the spatially unbound strand, excluding PAM) AGGTGGGAAACTGAGGCTT (Sequence ID: 25)
[0420] Complete sequence of mouse albumin-targeting sgRNA GACAUGCAUAUGUAUGUGUGUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU(Sequence ID: 31)
[0421] Guide sequences for mouse albumin targeting sgRNAs GACAUGCAUAUGUAUGUGUG (Sequence ID: 32)
[0422] Target sequences of mouse albumin-targeting sgRNAs (target sequences on the spatially unbound strand, excluding PAM) ACATGCATATGTATGTGTG(Sequence code: 26)
[0423] Complete sequence of human PCSK9 targeting sgRNA GAGGUGGGAAACUGAGGCUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU(Sequence ID: 29)
[0424] Guide sequence of human PCSK9 targeting sgRNA GAGGUGGGAAACUGAGGCUU (Sequence ID: 30)
[0425] Target sequences of human PCSK9 targeting sgRNA (target sequences on the spatially unbound strand, excluding PAM) AGGTGGGAAACTGAGGCTT (Sequence ID: 25)
[0426] Complete sequence of human albumin targeting sgRNA GACAUGCAUAUGUAUGUGUGUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU(Sequence ID: 31)
[0427] Guide sequences for human albumin targeting sgRNAs GACAUGCAUAUGUAUGUGUG (Sequence ID: 32)
[0428] Target sequences of human albumin-targeting sgRNA (target sequences on the spatially unbound strand, excluding PAM) ACATGCATATGTATGTGTG(Sequence code: 26)
[0429] Full sequence of FANCF targeting sgRNA GGAAUCCCUUCUGCAGCACCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU(Sequence ID: 33)
[0430] Guide sequence of FANCF targeting sgRNA GGAAUCCCUUCUGCAGCACC (Sequence ID: 34)
[0431] Target sequences of FANCF targeting sgRNA (target sequences on the spatially unbound strand, excluding PAM) GAATCCCTTCTGCAGCACC (Sequence ID: 35)
[0432] Complete sequence of VEGFA-targeting sgRNA GGGUGGGGGGAGUUUGCUCCGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU(Sequence ID: 36)
[0433] Guide sequence of VEGFA-targeting sgRNA GGGUGGGGGGAGUUUGCUCC (Sequence code: 37)
[0434] Target sequences of VEGFA-targeting sgRNAs (target sequences on the spatially unbound strand, excluding PAM) GGTGGGGGGAGTTTGCTCC (Sequence number: 38)
[0435] Full sequence of HBB targeting sgRNA GUUGCCCCACAGGGCAGUAAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU(Sequence ID: 39)
[0436] Guide sequence of HBB targeting sgRNA GUUGCCCCACAGGGCAGUAA (Sequence number: 40)
[0437] Target sequences of HBB-targeting sgRNAs (target sequences on the spatially unbound strand, excluding PAM) TTGCCCCACAGGGCAGTAA (Sequence ID: 41)
[0438] Amino acid sequence of SpCas9 MDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD (Sequence ID: 42)
[0439] DNA sequence encoding SpCas9 atggacaagaagtacagcatcggcctggacatcggtaccaacagcgtgggctgggccgtgatcaccgacgagtacaaggtgcccagcaagaagttcaaggtgctgggcaacaccgaccgccacagcatcaagaagaacctgatcggcgccctgctgttcgacagcggcgagaccgccgaggccacccgcctgaagcgcaccgcccgccgccgctacacccgccgcaagaaccgcatctgctacctgcaggagatcttcagcaacgagatggccaaggtggacgacagcttcttccaccgcctggaggagagcttcctggtggaggaggacaagaagcacgagcgccaccccatcttcggcaacatcgtggacgaggtggcctaccacgagaagtaccccaccatctaccacctgcgcaagaagctggtggacagcaccgacaaggccgacctgcgcctgatctacctggccctggcccacatgatcaagttccgcggccacttcctgatcgagggcgacctgaaccccgacaacagcgacgtggacaagctgttcatccagctggtgcagacctacaaccagctgttcgaggagaaccccatcaacgccagcggcgtggacgccaaggccatcctgagcgcccgcctgagcaagagccgccgcctggagaacctgatcgcccagctgcccggcgagaagaagaacggcctgttcggcaacctgatcgccctgagcctgggcctgacccccaacttcaagagcaacttcgacctggccgaggacgccaagctgcagctgagcaaggacacctacgacgacgacctggacaacctgctggcccagatcggcgaccagtacgccgacctgttcctggccgccaagaacctgagcgacgccatcctgctgagcgacatcctgcgcgtgaacaccgagatcaccaaggcccccctgagcgccagcatgatcaagcgctacgacgagcaccaccaggacctgaccctgctgaaggccctggtgcgccagcagctgcccgagaagtacaaggagatcttcttcgaccagagcaagaacggctacgccggctacatcgacggcggcgccagccaggaggagttctacaagttcatcaagcccatcctggagaagatggacggcaccgaggagctgctggtgaagctgaaccgcgaggacctgctgcgcaagcagcgcaccttcgacaacggcagcatcccccaccagatccacctgggcgagctgcacgccatcctgcgccgccaggaggacttctaccccttcctgaaggacaaccgcgagaagatcgagaagatcctgaccttccgcatcccctactacgtgggccccctggcccgcggcaacagccgcttcgcctggatgacccgcaagagcgaggagaccatcaccccctggaacttcgaggaggtggtggacaagggcgccagcgcccagagcttcatcgagcgcatgaccaacttcgacaagaacctgcccaacgagaaggtgctgcccaagcacagcctgctgtacgagtacttcaccgtgtacaacgagctgaccaaggtgaagtacgtgaccgagggcatgcgcaagcccgccttcctgagcggcgagcagaagaaggccatcgtggacctgctgttcaagaccaaccgcaaggtgaccgtgaagcagctgaaggaggactacttcaagaagatcgagtgcttcgacagcgtggagatcagcggcgtggaggaccgcttcaacgccagcctgggcacctaccacgacctgctgaagatcatcaaggacaaggacttcctggacaacgaggagaacgaggacatcctggaggacatcgtgctgaccctgaccctgttcgaggaccgcgagatgatcgaggagcgcctgaagacctacgcccacctgttcgacgacaaggtgatgaagcagctgaagcgccgccgctacaccggctggggccgcctgagccgcaagcttatcaacggcatccgcgacaagcagagcggcaagaccatcctggacttcctgaagagcgacggcttcgccaaccgcaacttcatgcagctgatccacgacgacagcctgaccttcaaggaggacatccagaaggcccaggtgagcggccagggcgacagcctgcacgagcacatcgccaacctggccggcagccccgccatcaagaagggcatcctgcagaccgtgaaggtggtggacgagctggtgaaggtgatgggccgccacaagcccgagaacatcgtgatcgagatggcccgcgagaaccagaccacccagaagggccagaagaacagccgcgagcgcatgaagcgcatcgaggagggcatcaaggagctgggcagccagatcctgaaggagcaccccgtggagaacacccagctgcagaacgagaagctgtacctgtactacctgcagaacggccgcgacatgtacgtggaccaggagctggacatcaaccgcctgagcgactacgacgtggaccacatcgtgccccagagcttcctgaaggacgacagcatcgacaacaaggtgctgacccgcagcgacaagaaccgcggcaagagcgacaacgtgcccagcgaggaggtggtgaagaagatgaagaactactggcgccagctgctgaacgccaagctgatcacccagcgcaagttcgacaacctgaccaaggccgagcgcggcggcctgagcgagctggacaaggccggcttcatcaagcgccagctggtggagacccgccagatcaccaagcacgtggcccagatcctggacagccgcatgaacaccaagtacgacgagaacgacaagctgatccgcgaggtgaaggtgatcaccctgaagagcaagctggtgagcgacttccgcaaggacttccagttctacaaggtgcgcgagatcaacaactaccaccacgcccacgacgcctacctgaacgccgtggtgggcaccgccctgatcaagaagtaccccaagctggagagcgagttcgtgtacggcgactacaaggtgtacgacgtgcgcaagatgatcgccaagagcgagcaggagatcggcaaggccaccgccaagtacttcttctacagcaacatcatgaacttcttcaagaccgagatcaccctggccaacggcgagatccgcaagcgccccctgatcgagaccaacggcgagaccggcgagatcgtgtgggacaagggccgcgacttcgccaccgtgcgcaaggtgctgagcatgccccaggtgaacatcgtgaagaagaccgaggtgcagaccggcggcttcagcaaggagagcatcctgcccaagcgcaacagcgacaagctgatcgcccgcaagaaggactgggaccccaagaagtacggcggcttcgacagccccaccgtggcctacagcgtgctggtggtggccaaggtggagaagggcaagagcaagaagctgaagagcgtgaaggagctgctgggcatcaccatcatggagcgcagcagcttcgagaagaaccccatcgacttcctggaggccaagggctacaaggaggtgaagaaggacctgatcatcaagctgcccaagtacagcctgttcgagctggagaacggccgcaagcgcatgctggccagcgccggcgagctgcagaagggcaacgagctggccctgcccagcaagtacgtgaacttcctgtacctggccagccactacgagaagctgaagggcagccccgaggacaacgagcagaagcagctgttcgtggagcagcacaagcactacctggacgagatcatcgagcagatcagcgagttcagcaagcgcgtgatcctggccgacgccaacctggacaaggtgctgagcgcctacaacaagcaccgcgacaagcccatccgcgagcaggccgagaacatcatccacctgttcaccctgaccaacctgggcgcccccgccgccttcaagtacttcgacaccaccatcgaccgcaagcgctacaccagcaccaaggaggtgctggacgccaccctgatccaccagagcatcaccggtctgtacgagacccgcatcgacctgagccagctgggcggcgac (Sequence ID: 43)
[0440] result Result 1 Selection of off-target prediction methods to be used for comparison Genome-wide off-target prediction methods can be categorized into three groups, according to their approaches: cell-based, in vitro, and in silico methods. Examples of approaches from these three groups are shown in Figure 1.
[0441] In IND studies on genome editing therapy, different combinations of off-target prediction methods have been used. Table 2 shows information on the off-target prediction methods used in IND studies on genome editing therapy.
[0442] Table 2. Off-target prediction methods used in gene editing drugs and IND studies (TALEN: transcription activator-like effector; NA: unavailable; LCA10: Leber congenital amaurosis 10; MPS: Mucopolysaccharidosis. ) [Table 2]
[0443] To compare the performance of off-target prediction methods, the inventors of this application selected one method for each of the categories described above. For cell-based off-target prediction methods, GUIDE-seq was selected. For in silico off-target prediction methods, CAS-OFFinder was selected. GUIDE-seq and CAS-OFFinder were most frequently used to predict off-target effects of Cas9 therapeutics, including EDIT101 and NTLA-2001. For in vitro off-target prediction methods, Digenome-seq was selected. This is because Digenome-seq was used in the EDIT101 study and is one of the most popular protocols in previous studies for comparison.
[0444] Result 2 Overview of Extru-seq and Extru-seq Condition Optimization The inventors of this application aimed to design a novel method that combines the positive features of cell-based and in vitro methods. To this end, an off-target prediction method, Extru-seq, was developed, characterized by the use of physical force to lyse cells and the mixing of genomic DNA with Cas9 and sgRNA. A schematic diagram of this novel off-target prediction method, Extru-seq, is shown in Figure 2.
[0445] For example, live HEK294T or NIH-3T3 cells are mixed with a pre-incubated Cas9-sgRNA RNP complex. Using an extruder (see reference [Goh, Wei Jiang, et al. "Bioinspired cell-derived nanovesicles versus exosomes as drug delivery systems: a cost-effective alternative." Scientific reports 7.1(2017):1-10.]), the mixture is passed through a filter (e.g., filter paper) with pore sizes smaller than the cell diameter. Because the mixture passes through a filter with pore sizes smaller than the cell diameter, the cells (e.g., cell membranes) are disrupted, allowing the Cas9 RNP to access the genomic DNA. Figures 20 and 21 show the results of experiments performed to determine the optimal conditions for the average pore size of the filter, the Cas9 RNP concentration in the mixture, and the number of cells in extru-seq.
[0446] In particular, Figure 20 shows the quality of genomic DNA cultured overnight with Cas9 RNP, analyzed via gel electrophoresis. Various numbers of NIH-3T3 cells and various pore sizes were tested. Here, "Con" represents control genomic DNA of sufficient quality for WGS analysis. "L" represents ladder DNA. Figure 21 shows information for each of samples 1 through 9. For example, the conditions for sample 8 (1X10) 7The electrophoretic results for the genomic DNA of the sample extruded under (cells / mL; pore size of 8 μm) are shown in line 8 of FIG. 20.
[0447] FIGS. 22 and 23 show the cleavage rates for the on- and off-target sites recognized by sgRNAs targeting the human PCSK9 locus, measured by quantitative PCR (qPCR). FIG. 22 shows the results of the cleavage rates for the on-target and off-target 2 sites for the samples using the human PCSK9 targeting sgRNA. FIG. 23 shows the results of the cleavage rates for the off-target 4 and off-target 7 sites for the samples using the human PCSK9 targeting sgRNA.
[0448] Through this experiment, the inventors of the present application determined the optimized conditions for Extru-seq for subsequent experiments as follows: a pore size of 8 μM, a Cas9 RNP concentration of 5000 nM, and 10 7 cells. Under the optimized conditions, it was confirmed that the quality of the genomic DNA cultured overnight at 37° C. with Cas9 RNP was high enough to construct a whole genome sequencing (WGS) library.
[0449] Result 3. Measurement of the NHEJ level after the extrusion step The inventors of the present application hypothesized that there is no DNA repair mechanism for religating the genomic DNA cleaved by Cas9 in the Extru-seq process. In fact, when the DNA was cleaved through the Extru-seq process and the cleavage rate of the target site was measured using quantitative PCR, an average rate of 70% was observed. The results regarding the cleavage rate of the target site are shown in FIGS. 24 to 31.
[0450] Figures 24 to 30 show WGS data from extru-seq analyzed using IGV to identify cleavage patterns. Figure 24 shows results obtained using human PCSK9-targeting sgRNA. Figure 25 shows results obtained using human albumin-targeting sgRNA. Figure 26 shows results obtained using mouse PCSK9-targeting sgRNA. Figure 27 shows results obtained using mouse albumin-targeting sgRNA. Figure 28 shows results obtained using human FANCF-targeting sgRNA. Figure 29 shows results obtained using human VEGFA-targeting sgRNA. Figure 30 shows results obtained using human HBB-targeting sgRNA.
[0451] Figure 31 shows the cleavage rates of seven on-target sites for each target, obtained through manual calculation based on IGV analysis of qPCR and WGS data. In Figure 31, the y-axis represents the cleavage rate.
[0452] The results shown in Figures 24 to 31 demonstrate the absence of DNA repair mechanisms such as NHEJ, indicating that extru-seq can reflect positive features observed in vitro.
[0453] Furthermore, the inventors of the present invention analyzed cleaved and uncleaved populations at the on-target site to investigate which NHEJs occurred or to what extent after the extrusion process. Firstly, considering that indel mutations accumulated in the uncleaved group when the NHEJ process was complete during the incubation period after the extrusion process, the uncleaved population in Extru-seq samples was analyzed through deep sequencing. The deep sequencing results for Extru-seq samples treated with the Cas9 RNP complex were compared with the results for control samples not treated with the Cas9 RNP complex. As a result of the comparison, there was no significant difference between the two samples. The deep sequencing results for the uncleaved population are shown in Figure 32. In particular, Figure 32 shows the indel frequencies measured by targeted deep sequencing for the uncleaved population of Extru-seq samples related to Figures 24 to 30. Indel frequencies were measured for untreated and Cas9-treated samples. In Figure 32, samples not treated with Cas9 are shown as Cas9(-), and samples treated with Cas9 are shown as Cas9(+). An independent Student t-test was used for the t-test. Error bars indicate the standard deviation (n=3).
[0454] These results demonstrate that the NHEJ levels after the extrusion process are not significant.
[0455] Secondly, using a protocol from multiplex Digenome-seq (see [Kim, Daesik, et al. "Genome-wide target specificities of CRISPR-Cas9 nucleases revealed by multiplex Digenome-seq." Genome Research 26.3(2016):406-415.]), the inventors of the present invention performed multiplex Extru-seq to measure the change in cleavage rates at five different on-target sites in the presence or absence of SCR7 (chemical DNA ligase IV or NHEJ inhibitors; see [Chu, Van Trung, et al. "Increasing the efficiency of homology-directed repair for CRISPR-Cas9-induced precise gene editing in mammalian cells." Nature Biotechnology 33.5(2015):543-548.]). When NHEJ occurs, the cleavage rate increases due to the presence of SCR7, and this effect may also accumulate during the incubation phase. However, the difference in mean cleavage rates at the on-target site of 5 with or without the presence of SCR7 was not significant. Results related to the presence or absence of SCR7 are shown in Figure 33. In particular, Figure 33 shows the cleavage rate (%) measured using qPCR at the target site of 5. This result was obtained by multiplex extru-seq in the presence (+SCR7 in Figure 33) or absence (-SCR7 in Figure 33) of 1 μM SCR7. The horizontal line in the graph represents the mean for the experiment (n=5). For the t-test, an independent Student's t-test was used.
[0456] These results further demonstrate that NHEJ does not have a significant effect on on-target fracture rates.
[0457] The inventors of this application hypothesized that cellular components other than genomic DNA would not be damaged, such that the cleavage pattern would be similar to that of cell-based off-target prediction methods. This hypothesis was tested by comparing Extru-seq results with those of cell-based and in vitro-based methods (described below). Extru-seq was confirmed to be an off-target prediction method that can reflect the positive features of both cell-based off-target prediction methods (complete cellular components other than genomic DNA) and in vitro off-target prediction methods (absence of DNA repair mechanisms).
[0458] Result 4. Design and use of indiscriminate guide arrays A second objective of this study was to conduct a standardized test that could effectively measure the performance metrics of each method. Previous studies used predicted guide sequences to recognize only a small number of off-target sites in the genome to compare other methods. As a result, only a small number of guide sequence-validated off-target loci were discovered, and therefore it was difficult to effectively compare different prediction methods using a statistically significant number of loci. In more recent publications (see references [Wienert, Beeke, et al. "Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq."Science 364.6437(2019):286-289.; and Akcakaya, Pinar, et al. "In vivo CRISPR editing with no detectable genome-wide off-target mutations."Nature 561.7723(2018):416-419.]), predicted indiscriminate guide sequences were used to recognize a large number of off-target loci. The use of indiscriminate guide sequences has provided a powerful testbed for genome-wide off-target prediction methods. However, these indiscriminate guide sequences were not used in this study. One of them was involved in a mouse guide sequence targeting PCSK9, which was not complementary to the human cell sequence (as shown in previous studies). Another, targeting VEGFA, lacked an off-target locus predicted as a single mismatch (see Table 3).
[0459] To overcome these limitations, this study sought two types of indiscriminate guide sequences targeting PCSK9 and albumin in the mouse genome. These guide sequences also perfectly matched target sequences present in the human genome (however, the PCSK9 and albumin loci are not targets in the human genome). Even if the guide sequences target loci other than PCSK9 or albumin in the human genome, they are referred to as human PCSK9 and human albumin for convenience. The number of off-target sequences related to these indiscriminate guide sequences was calculated using Cas-OFFinder in each of the human and mouse genomes. The selected guide sequences were confirmed to be associated with a large number of off-target sequences in both genomes. Information on the target sequences of guide sequences used in previous studies and the indiscriminate guide sequences used in this study, as well as the results of the investigation into off-target sites, are shown below in Table 3. In particular, Table 3 shows the results of the investigation into genome-wide off-target loci, including 0 to 6 mismatches. Off-target sites were predicted via Cas-OFFinder in the genomes hg19 (Table 3(a)) and mm10 (Table 3(b)). Referring to Table 3, the sequences GACCCCCTCCACCCCGCCTC (SEQ ID NO: 72) (VEGFA target sequence), AGCAGCAGCGGCGGCAACAG (SEQ ID NO: 73) (PCSK9 target sequence, previous study), ACATGCATATGTATGTGTG (SEQ ID NO: 26) (Albumin-target sequence), and AGGTGGGAAACTGAGGCTT (SEQ ID NO: 25) (PCSK9 target sequence) are shown.
[0460] Table 3. Research findings on whole-genome off-target sites according to target sequences or guide sequences [Table 3] [Table 4]
[0461] Result 5. Prediction of genome-wide off-target sites using GUIDE-seq, Digenome-seq, in silico methods, and Extru-seq. In silico predictions based on GUIDE-seq, Digenome-seq, Extru-seq, and Cas-OFFinder were performed using indiscriminate sgRNA sequences targeting PCSK9 and albumin, respectively. Off-target prediction systems for each sgRNA sequence were applied to human cell lines (HEK293T) and mouse cell lines (NIH-3T3). The results are shown in Figures 34 to 62, and Figures 3 and 4.
[0462] In particular, Figures 34 and 35 show GUIDE-seq results obtained from HEK293T cells using PCSK9-targeting sgRNA. In Figures 34 and 35, the target sequence (including the PAM sequence) AGGTGGGAAACTGAGGCTTNGG (Sequence ID: 44) is shown.
[0463] Figures 36 and 37 show GUIDE-seq results obtained from HEK293T cells using albumin-targeting sgRNA. In Figures 36 and 37, the target sequence (including the PAM sequence) ACATGCATATGTATGTGTGNGG (Sequence ID: 45) is shown.
[0464] Figures 38 and 39 show GUIDE-seq results obtained from NIH-3T3 cells using PCSK9-targeting sgRNA. In Figures 38 and 39, the target sequence (including the PAM sequence) AGGTGGGAAACTGAGGCTTNGG (Sequence ID: 44) is shown.
[0465] Figures 40 and 41 show GUIDE-seq results obtained from NIH-3T3 cells using albumin-targeting sgRNA. In Figures 40 and 41, the target sequence (including the PAM sequence) ACATGCATATGTATGTGTGNGG (Sequence ID: 45) is shown.
[0466] Relatively low-ranking off-target loci were omitted from the corresponding diagrams generated by the GUIDE-seq analysis program. These omitted loci were included in subsequent analyses.
[0467] Figures 42 and 43 show Manhattan plot results of Digenome-seq obtained from HEK293T cells using PCSK9-targeted sgRNA, where the y-axis represents the DNA cleavage score.
[0468] Figures 44 and 45 show Manhattan plot results of Digenome-seq obtained from HEK293T cells using albumin-targeting sgRNA, where the y-axis represents the DNA cleavage score.
[0469] Figures 46 and 47 show Manhattan plot results of Digenome-seq obtained from NIH-3T3 cells using PCSK9-targeted sgRNA. Here, the y-axis represents the DNA cleavage score.
[0470] Figures 48 and 49 show Manhattan plot results of Digenome-seq obtained from NIH-3T3 cells using albumin-targeting sgRNA, where the y-axis represents the DNA cleavage score.
[0471] Figures 50 and 51 show Manhattan plot results of extru-seq obtained from HEK293T cells using PCSK9-targeted sgRNA, where the y-axis represents the DNA cleavage score.
[0472] Figures 52 and 53 show Manhattan plot results of extru-seq obtained from HEK293T cells using albumin-targeting sgRNA, where the y-axis represents the DNA cleavage score.
[0473] Figures 54 and 55 show Manhattan plot results of extru-seq obtained from NIH-3T3 cells using PCSK9-targeted sgRNA. Here, the y-axis represents the DNA cleavage score.
[0474] Figures 56 and 57 show Manhattan plot results of extru-seq obtained from NIH-3T3 cells using albumin-targeting sgRNA, where the y-axis represents the DNA cleavage score.
[0475] The inventors of this application predicted off-target sites (candidate off-target sites) for sgRNAs (sgRNAs targeting human PCSK9, sgRNAs targeting human albumin, sgRNAs targeting mouse PCSK9, and sgRNAs targeting mouse albumin) using GUIDE-seq, Digenome-seq, Extru-seq, and in silico methods, and compared the results. The comparison results are shown in Figures 3 and 4 using Venn diagrams. In particular, Figure 3 shows the comparison results for sgRNAs targeting human PCSK9 and sgRNAs targeting human albumin. Figure 4 shows the comparison results for sgRNAs targeting mouse PCSK9 and sgRNAs targeting mouse albumin. For the results in Figures 3 and 4, human cell lines (HEK293T) and mouse cell lines (NIH-3T3) were used.
[0476] It was possible to rank candidate off-target loci using GUIDE-seq sequence read counts and DNA cleavage scores from Digenome-seq and Extru-seq. Regarding in silico prediction based on Cas-OFFinder, no scores exist that can be used for this ranking. Therefore, the inventors of this application calculated predictive scores for candidate off-target sites for ranking using two different scripts of machine learning studies (see references [Liu, Qiaoyue, et al. "Deep learning improves the ability of sgRNA off-target propensity prediction." BMC bioinformatics 21.1(2020):1-15.; and Doench, John G., et al. "Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9." Nature biotechnology 34.2(2016):184-191.]). The prediction scores are as follows: the CRISPR off-target predictor (CROP) score (a heuristic score indicating whether a candidate off-target site will be edited) and the cutting frequency determination (CFD) score (a percentage activity value provided to the penalty matrix based on each possible type of mismatch at each position in the guide RNA sequence). The distribution of sequence read counts, DNA cleavage, and in silico prediction scores for each candidate off-target locus is shown in tables against the number of mismatches with the guide sequence. Tables showing the distribution of in silico prediction scores according to the number of sequence read counts, DNA cleavage scores, and mismatches with the guide sequence are shown in Figures 58 to 62.
[0477] In particular, Figure 58 shows the results related to on-target and off-target scores (scores calculated from sequence read count results) based on the number of mismatches predicted using GUIDE-seq. The x-axis represents the number of mismatches, and the score for each number of mismatches is shown. That is, the number of mismatches for on-target and off-target sites is shown on the x-axis. The scores converted from sequence read counts are shown on the y-axis.
[0478] Figure 59 shows the results related to on-target and off-target scores (break scores in the Manhattan plot) based on the number of mismatches, as predicted using Digenome-seq. The number of mismatches is shown on the x-axis, and the score for each number of mismatches is shown. That is, the number of mismatches for on-target and off-target sites is shown on the x-axis, and the break scores are shown on the y-axis.
[0479] Figure 60 shows the results related to on-target and off-target scores (CROP scores) based on the number of mismatches, as predicted using the in silico system. The x-axis represents the number of mismatches, and the score for each number of mismatches is shown. That is, the number of mismatches for on-target and off-target sites is shown on the x-axis. The CROP score is shown on the y-axis.
[0480] Figure 61 shows the results related to on-target and off-target scores (CFD scores) based on the number of mismatches, as predicted using an in silico system. The x-axis represents the number of mismatches, and the score for each number of mismatches is shown. That is, the number of mismatches for on-target and off-target sites is shown on the x-axis. The CFD score is shown on the y-axis.
[0481] Figure 62 shows the results related to on-target and off-target scores (break scores in the Manhattan plot) based on the number of mismatches, as predicted using Extru-seq. The x-axis represents the number of mismatches, and the score for each number of mismatches is shown. That is, the number of mismatches for on-target and off-target sites is shown on the x-axis. The break scores are shown on the y-axis.
[0482] It was expected that the corresponding prediction score would decrease as the number of mismatches increased. GUIDE-seq and in silico prediction followed this trend, but outliers with high DNA cleavage scores were observed even when 4, 5, or 6 mismatches were present in the Digenome-seq results. Unlike the Digenome-seq results, when DNA cleavage scores for off-target site candidates were calculated using the Extru-seq approach, high DNA cleavage scores were not observed for off-target sgRNA off-target candidates with 4 or more mismatches. This result indicates that Extru-seq identified fewer false positives than Digenome-seq. The inventors of this application confirmed this idea by validating the effectiveness of off-target candidates with high scores.
[0483] Result 6. Effectiveness confirmation rate of GUIDE-seq and Extru-seq GUIDE-seq and Extru-seq demonstrated high efficacy confirmation rates. Efficacy confirmation of predicted off-target loci was performed in human cell lines and mouse models. In human cell line experiments, plasmids encoding Cas9 protein and sgRNA were transfected into HEK293T cells. In mouse experiments, sequences encoding Cas9 protein and sgRNA were packaged into adeno-associated virus (AAV) serotype 8 (i.e., AAV8). These AAVs were then delivered to C57BL / 6 mice via systemic or subretinal injection. Since only subretinal injection showed high on-target indel formation, model retinal pigment epithelial cells were used in efficacy confirmation experiments.
[0484] The results regarding indel formation frequency for subretinal and systemic injections are shown in Figures 63 and 64. In particular, the indel rate was calculated through the analysis of genomic DNA obtained from the organs of C57BL / 6 mice injected with two AAV8 vectors expressing Cas9 and PCSK9 or albumin-targeting sgRNAs, respectively. In Figures 63 and 64, results marked as iv represent the results for systemic injection, and results marked as subretinal represent the results for subretinal injection. In Figures 63 and 64, error bars indicate sem(n=3). NR represents neuroretina, and RPE represents retinal pigment epithelial cells.
[0485] The top 10 candidates for each prediction method were tested using targeted deep sequencing. The results showed that Extru-seq had a 92.5% efficacy confirmation rate, and GUIDE-seq had a 97.5% efficacy confirmation rate. However, Digenome-seq had a 45% efficacy confirmation rate, and the in silico method showed a 62.5% efficacy confirmation rate for CROP and a 67.5% efficacy confirmation rate for CFD. Extru-seq and GUIDE-seq showed significantly higher efficacy confirmation rates than Digenome-seq and the in silico method. The efficacy confirmation rate results for each prediction method are shown in Figure 5 and Table 4.
[0486] In particular, Figure 5 shows the efficacy confirmation rates of top off-target sites predicted by in silico, GUIDE-seq, Digenome-seq, and Extru-seq. This shows experimental results in human and mouse cells for indiscriminate sgRNAs targeting PCSK9 and albumin (*P<0.05, ns, not significant in two-sided unpaired Mann-Whitney test).
[0487] Table 4 below shows the efficacy confirmation results obtained through targeted deep sequencing of the top 10 off-target sites (using human PCSK9-targeting sgRNA, human albumin-targeting sgRNA, mouse PCSK9-targeting sgRNA, and mouse albumin-targeting sgRNA) as predicted by each method. For efficacy confirmation, the indel frequency at the off-target site must be greater than 0.1%, and the formula '(indel frequency at the off-target locus) / (indel frequency in the control without Cas9 treatment)>2' must be satisfied (see reference [Frangoul, Haydar, et al. "CRISPR-Cas9 gene editing for sickle cell disease and β-thalassemia." New England Journal of Medicine 384.3(2021):252-260.]). * indicates that the target was manually confirmed. Here, manually identified (off-target) sites refer to sites that were determined to be negative using Digenome software but positive using IGV software. As a result of IGV software verification, one additional site was identified in the human PCSK9 sample. Five additional sites were identified in the human albumin sample. Three additional sites were identified in the mouse PCSK9 sample. Three additional sites were identified in the mouse albumin sample.
[0488] Table 4. Efficacy confirmation results for the top 10 candidate off-target sites predicted by each off-target prediction method. (a) [Table 5] (b) [Table 6] (c) [Table 7] (d) [Table 8]
[0489] Information about manually identified targets is shown in Figures 77 to 116. In particular, WGS data visualized using IGV for off-target sites whose efficacy was manually validated from extru-seq is shown in Figures 77 to 116. In Figures 77 to 116, the sequences of the off-target sites are shown, and sequences that mismatch with the guide sequence are shown in lowercase. PAM sequences are underlined.
[0490] Figures 77 and 78 disclose the off-target sequence AGGTGGGAAACTGAGGCccAGG (Sequence No. 52).
[0491] Figures 79 and 80 disclose the off-target sequence tgATGCATATGTATGTGTGGaGG (Sequence ID: 53).
[0492] Figures 81 and 82 disclose the off-target sequence AaATGCATATGTATGaGTGTGG (sequence number: 54).
[0493] Figures 83 and 84 disclose the off-target sequence CATGCATATGcATGTGgGAGG (Sequence ID: 55).
[0494] Figures 85 and 86 disclose the off-target sequence AgATGCATAgGTATGTGTGTGG (Sequence ID: 56).
[0495] Figures 87 and 88 disclose the off-target sequence ACtTGCATATcTATGTGTGTGG (Sequence No. 57).
[0496] Figures 89 and 90 disclose the off-target sequence ccGTGGGAAACTGAGGCTTGGG (Sequence ID: 58).
[0497] Figures 91 and 92 disclose the off-target sequence AGGTGGGAAACTGAGGCTgAGG (Sequence ID: 59).
[0498] Figures 93 and 94 disclose the off-target sequence AGGaGGGAAACTGAGGCTcAGG (Sequence ID: 60).
[0499] Figures 95 and 96 disclose the off-target sequence AaATaCATATGTATGTGTGTGG (Sequence ID: 61).
[0500] Figures 97 and 98 disclose the off-target sequence ACATGtATATGTATaTGTGTGG (Sequence ID: 62).
[0501] Figures 99 and 100 disclose the off-target sequence ACATatATATGTATGTGTGTGG (sequence number: 63).
[0502] Figures 101 and 102 disclose the off-target sequence GGGTGGGtGGAGTTTGCTaCTGG (Sequence ID: 64).
[0503] Figures 103 and 104 disclose the off-target sequence aGGTGGtGGGAGcTTGtTCCTGG (Sequence ID: 65).
[0504] Figures 105 and 106 disclose the off-target sequence GGtgGGGGtGgGTTTGCTCCTGG (Sequence ID: 66).
[0505] Figures 107 and 108 disclose the off-target sequence GGGcaaGGGGAGgTTGCTCCTGG (Sequence No. 67).
[0506] Figures 109 and 110 disclose the off-target sequence GGAtTgCCaTCcGCAGCACCTGG (Sequence ID: 68).
[0507] Figures 111 and 112 disclose the off-target sequence GGAgTCCCTcCTGCAGCACCTGA (Sequence ID: 69).
[0508] Figures 113 and 114 disclose the off-target sequence aGAggCCCcTCTGCAGCACCAGG (Sequence ID: 70).
[0509] Figures 115 and 116 disclose the off-target sequence accATCCCTcCTGCAGCACCAGG (sequence number: 71).
[0510] Result 7. Further comparison of Extru-seq, GUIDE-seq, Digenome-seq, and DIG-seq Digenome-seq uses purified genomic DNA from which components such as chromatin proteins have been lost. To overcome this problem, previous studies have developed an improved version of Digenome-seq, named DIG-seq. DIG-seq, which uses cell-free chromatin DNA rather than histone-free DNA, predicted fewer false positives than Digenome-seq. The neutral detergent used to lyse cells in the DIG-seq approach may affect the chromatin state of the cell DNA, which may also affect the Cas9 cleavage mechanism. Therefore, the inventors of this application predicted that Extru-seq, which uses physical force for cell lysis, would better reflect the characteristics of cell-based methods compared to DIG-seq.
[0511] To compare extru-seq with other in vitro methods, the inventors of this application performed GUIDE-seq and extru-seq using guide sequences targeting FANCF, VEGFA, and HBB in HeLa cells. Each of the FANCF, VEGFA, and HBB-targeting guide sequences was used to compare DIG-seq and Digenome-seq in previous studies (see reference [Kim, Daesik, and Jin-Soo Kim. DIG-seq: a genome-wide CRISPR off-target profiling method using chromatin DNA. Genome research 28.12(2018):1894-1900.]).
[0512] The GUIDE-seq results in HeLa cells are shown in Figures 65 to 67 (using sgRNAs targeting FANCF, VEGFA, and HBB). In Figure 65, the target sequence (including PAM) GAATCCCTTCTGCAGCACCNGG (Sequence ID: 46) is shown. In Figure 66, the target sequence (including PAM) GGTGGGGGGAGTTTGCTCCNGG (Sequence ID: 47) is shown. In Figure 67, the target sequence (including PAM) TTGCCCCACAGGGCAGTAANGG (Sequence ID: 48) is shown. In particular, Figure 65 shows the results (sequence read results) of predicting off-target sequences through GUIDE-seq using FANCF-targeting sgRNA. Figure 66 shows the results (sequence read results) of predicting off-target sequences through GUIDE-seq using VEGFA-targeting sgRNA. Figure 67 shows the results (sequence read results) of predicting off-target sequences through GUIDE-seq using HBB-targeting sgRNA.
[0513] Extru-seq results in HeLa cells are shown in Figures 68 to 73 (using sgRNAs targeting FANCF, VEGFA, and HBB). In particular, Figures 68 and 69 show the results of predicting off-target sequencing through extru-seq using FANCF-targeting sgRNA (Manhattan plot results). Figures 70 and 71 show the results of predicting off-target sequencing through extru-seq using VEGFA-targeting sgRNA (Manhattan plot results). Figures 72 and 73 show the results of predicting off-target sequencing through extru-seq using HBB-targeting sgRNA (Manhattan plot results). The y-axis represents the DNA cleavage score.
[0514] Figures 6 and 7 show a comparison of results regarding off-target candidates according to each method (Extru-seq, GUIDE-seq, DIG-seq, Digenome-seq) using Venn diagrams (with HeLa cells, human FANCF-targeting sgRNA, human VEGFA-targeting sgRNA, and human HBB-targeting sgRNA). The analysis results through Venn diagrams show that Digenome-seq and DIG-seq predicted multiple different off-target loci. On the other hand, it shows that most of the off-target loci predicted by Extru-seq were identified by at least one of the different off-target prediction methods.
[0515] When tested to confirm whether candidate loci could be validated, Extru-seq showed a higher validation rate than DIG-seq and Digenome-seq. The validation results are shown in Table 5 below.
[0516] In particular, Table 5(a) shows the efficacy confirmation results for the top 10 predicted off-target sites for sgRNAs targeting human FANCF. Table 5(b) shows the efficacy confirmation results for the top 10 predicted off-target sites for sgRNAs targeting human VEGFA. Table 5(c) shows the efficacy confirmation results for the top 10 predicted off-target sites for human HBB-targeting sgRNAs. Off-target sites were validated through targeted deep sequencing. For efficacy confirmation, the indel frequency of the off-target site must be higher than 0.1% and satisfy the formula '(indel frequency at the off-target locus) / (indel frequency in the control)>2'. In Table 5, * indicates that the target was manually confirmed. Results for manually confirmed targets are shown in Figures 101 to 116.
[0517] Table 5. Results of confirming the effectiveness of the top 10 off-target sites predicted by each off-target prediction method. (a) [Table 9] (b) [Table 10] (c) [Table 11]
[0518] Figure 8 shows the efficacy confirmation results for off-target sites predicted by each method (DIG-seq, Digenome-seq, Extru-seq, and GUIDE-seq). In particular, the efficacy confirmation rates for results related to Figures 6 and 7, and Table 6 are shown graphically (ns, not significant in two-sided unpaired Mann-Whitney U test).
[0519] Result 8. Comparison of rank distributions of off-target sites predicted by GUIDE-seq and Extru-seq. The inventors of this application compared the degree of agreement between prediction results obtained by extru-seq, cell-based, in vitro, and in silico methods. The top 10 candidate off-target loci for each prediction method were listed in a table, and the ranking of these loci for each different method was also listed in a table. Subsequently, the ranking of off-target loci in each method was compared. The results are shown in Tables 6 to 9. The top 10 predicted loci for each method pair were counted, and the sharing ratio was calculated. For this purpose, the sharing ratio (%) of the top 10 (ranks) was calculated. That is, the top 10 off-target candidates were extracted in method A, and the ranking of the off-target candidates corresponding to the top 10 off-target candidates in A in another method (e.g., method B) was compared. If the corresponding off-target candidates were in the top 10 in the other method, they were determined to contribute to the sharing ratio. As a result of the calculation, low similarity (overall average sharing ratio of the top 10 = 22%) was observed in most cases. The highest similarity was consistently found between GUIDE-seq and Extru-seq pairwise comparisons (average sharing percentage of the top 10 GUIDE-seq and Extru-seq pairs = 43%).
[0520] Table 6. Off-target prediction results for sgRNAs targeting human PCSK9 (in silico, GUIDE-seq, Digenome-seq, and Extru-seq) [Table 12-1] [Table 12-2] [Table 12-3] [Table 12-4] [Table 12-5]
[0521] Table 7. Off-target prediction results for sgRNAs targeting human albumin (in silico, GUIDE-seq, Digenome-seq, and Extru-seq) [Table 13-1] [Table 13-2] [Table 13-3] [Table 13-4] [Table 13-5]
[0522] Table 8. Off-target prediction results for sgRNA targeting mouse PCSK9 (in silico, GUIDE-seq, Digenome-seq, and Extru-seq) [Table 14-1] [Table 14-2] [Table 14-3] [Table 14-4] [Table 14-5]
[0523] Table 9. Off-target prediction results for sgRNAs targeting mouse albumin (in silico, GUIDE-seq, Digenome-seq, and Extru-seq) [Table 15-1] [Table 15-2] [Table 15-3] [Table 15-4] [Table 15-5]
[0524] A ranking comparison can also be performed with all off-target sites, including those whose effectiveness has not been confirmed. Venn diagrams (Figures 3, 4, 6, and 7) show that there is a statistically significant number of candidate off-target sites in the common region analyzed. To check for median homogeneity of locus scores in the common region of the results from the two methods, the score / read count was min-max normalized and the Wilcoxon rank-sum test was performed. The analysis results are shown in Figure 9. In Figure 9, the dotted line represents p=0.05. Since a sample size of at least 16 was required to use the asymptotic nonparametric Wilcoxon rank test (see references [MUNDRY, ROGER, and JULIA FISCHER. Use of statistical programs for nonparametric tests of small samples often leads to incorrect P values: examples from animal behavior. Animal behavior 56.1(1998):256-259.; and Dwivedi, Alok Kumar, Indika Mallawaarachchi, and Luis A. Alvarado. Analysis of small sample size studies using nonparametric bootstrap test with pooled resampling method. Statistics in medicine 36.14(2017):2187-2205.]), the common area with fewer than 16 samples is not included in this analysis. See Table 10 below. In this test, low p values indicate that the locus scores are distributed differently from the common area between the two populations. With the exception of the GUIDE-seq:Extru-seq and DIG-seq:Digenome-seq pairs, none of the pairs had similar distributions and all had high p-values for N≧3. In Figure 9, p-values were obtained from normalized rank-sum tests for each off-target prediction method pair.Regarding sgRNAs, sgRNAs targeting FANCF, VEGFA, and HBB were used in HeLa cells, while sgRNAs targeting PCSK9 and albumin were used in human and mouse cells (n≧16 were selected for analysis).
[0525] Table 10 below is shown in relation to Figure 9. Table 10 shows the number of samples found in the intersection of the Venn diagram. The table shows the number of samples for each sgRNA (sgRNA targeting human PCSK9, human albumin, mouse PCSK9, mouse albumin, human FANCF, human VEGFA, and human HBB). Cases with n >= 16 (where 16 is the minimum number of samples required for the asymptotic nonparametric Wilcoxon rank test) are underlined.
[0526] Table 10. Number of predicted off-target candidates found in the overlapping regions in Figures 3 and 4, and Figures 6 and 7 (comparison by each method) [Table 16] [Table 17]
[0527] The discrepancies between GUIDE-seq or Extru-seq results and Digenome-seq or in silico prediction results are presumed to be due to the low efficacy confirmation rate of Digenome-seq caused by a large number of false positives and the low efficacy confirmation rate of in silico prediction caused by the difference between machine learning-based prediction scores and real-world experimental values. Furthermore, unlike the Digenome-seq:Extru-seq pair, which showed low p-values, the DIG-seq:Digenome-seq pair consistently showed high p-values, and the results obtained from DIG-seq were analyzed to be similar to the results obtained from in vitro prediction methods (shown here as Digenome-seq). However, the results obtained from Extru-seq were analyzed to be more similar to cell-based prediction methods (shown here as GUIDE-seq) and different from the results obtained from in vitro prediction methods such as Digenome-seq. This is because, as mentioned above, the results for GUIDE-seq:Extru-seq pairs showed high p-values, while the results for Extru-seq:Digenome-seq pairs showed low p-values. In this regard, Extru-seq is distinguished from DIG-seq (which still shows similarity to Digenome-seq) in that it loses its similarity to in vitro Digenome-seq. All experimental processes for Digenome-seq, DIG-seq, and Extru-seq used in this experiment involved WGS, and the analysis through GUIDE-seq was performed based on PCR; therefore, such results showing similarity between Extru-seq and GUIDE-seq are somewhat surprising. This indicates that the conditions for processing genomic DNA with Cas9 are more important than the analytical procedure.
[0528] Result 9. Comparison of missed seq rates between Extru-seq and GUIDE-seq Cell-based methods, including GUIDE-seq, are known to sometimes miss true off-target candidates. The inventors of this application calculated the missed rate (or false negative rate calculated by the following formula: (number of false negatives) / (number of false negatives + number of true positives)) using a Venn diagram showing the overlap between effective and confirmed targets in samples analyzed by extru-seq and GUIDE-seq prediction and deep sequencing. As a result of the investigation, the mean missed rate for extru-seq was confirmed to be 2.3%, and the mean missed rate for GUIDE-seq was confirmed to be 29% (see Figures 10 to 14 and Tables 11 to 17). The combined results for the missed rates of extru-seq and GUIDE-seq are shown in the graph in Figure 14.
[0529] In particular, Figures 10 to 13 show Venn diagrams used to confirm the missed target rate. Specifically, the Venn diagrams shown in Figures 10 to 13 show comparative results regarding off-target candidates predicted by Extru-seq and GUIDE-seq, and off-targets confirmed to be effective (shown as effectiveness confirmations). Off-target prediction and effectiveness confirmation were performed using (a) sgRNA targeting human PCSK9 (Figure 10), (b) sgRNA targeting human albumin (Figure 10), (c) sgRNA targeting mouse PCSK9 (Figure 11), (d) sgRNA targeting mouse albumin (Figure 11), (e) sgRNA targeting human FANCF (Figure 12), (f) sgRNA targeting human VEGFA (Figure 12), and (g) sgRNA targeting human HBB (Figure 13). Effectiveness confirmations represent targets (off-target and on-target) confirmed to be effective by targeted deep sequencing. In Figures 10 to 13, * indicates the number of manually identified off-target sites.
[0530] Figure 14 shows a graph of the missed smear rates investigated for Extru-seq and GUIDE-seq. The missed smear rates are disclosed for each method. Furthermore, the missed smear rates are disclosed for each sgRNA (sgRNA targeting human PCSK9, sgRNA targeting human albumin, sgRNA targeting mouse PCSK9, sgRNA targeting mouse albumin, sgRNA targeting human FANCF, sgRNA targeting human VEGFA, and sgRNA targeting human HBB) (*: P<0.05 in two-sided unpaired Mann-Whitney test).
[0531] Figure 15 shows the distribution of the number of off-target mismatches missed in GUIDE-seq.
[0532] Tables 11 to 17 show the off-target analysis results predicted by each method (Extru-seq or GUIDE-seq) (related to Figures 10 to 14) and the actual off-target analysis results confirmed through deep sequencing. Specifically, the results regarding off-target candidates predicted by Extru-seq or GUIDE-seq were compared with the off-target sites confirmed to be effective by deep sequencing. (a) sgRNA targeting human PCSK9, (b) sgRNA targeting human albumin, (c) sgRNA targeting mouse PCSK9, (d) sgRNA targeting mouse albumin, (e) sgRNA targeting human FANCF, (f) sgRNA targeting human VEGFA, and (g) sgRNA targeting human HBB were investigated. For efficacy confirmation, the indel frequency at off-target sites must be higher than 0.1%, and the formula '((indel frequency at off-target locus) / (indel frequency in control)>2)' must be satisfied.
[0533] In Tables 11 to 17, the Target and Location columns show information about targets confirmed as valid through deep sequencing. A + indicates that the target (valid on-target or off-target) was predicted by the indicated off-target prediction method. A blank space indicates that the target (on-target or off-target) was not predicted by the indicated off-target prediction method (i.e., the indicated off-target prediction method missed a valid target). The miss rate was calculated using the formula '(number of blank spaces) / (total number of columns)'. An * indicates that the target was manually confirmed (see Figures 77 to 116).
[0534] Table 11. Results of examining off-target prediction methods, extru-seq, and GUIDE-seq miss rates (results for sgRNAs targeting human PCSK9) [Table 18-1] [Table 18-2]
[0535] Table 12. Results of examining the missed detection rates of off-target prediction methods, Extru-seq, and GUIDE-seq (results for sgRNAs targeting human albumin). [Table 19]
[0536] Table 13. Results of examining off-target prediction methods and the missed detection rates of Extru-seq and GUIDE-seq (results for sgRNA targeting mouse PCSK9) [Table 20]
[0537] Table 14. Results of examining off-target prediction methods, extru-seq, and GUIDE-seq miss rates (results for sgRNAs targeting mouse albumin). [Table 21]
[0538] Table 15. Results of examining off-target prediction methods, extru-seq, and GUIDE-seq miss rates (results for sgRNAs targeting human FANCF) [Table 22]
[0539] Table 16. Results of examining the missed detection rates of off-target prediction methods, Extru-seq, and GUIDE-seq (results for sgRNAs targeting human VEGFA) [Table 23]
[0540] Table 17. Results of examining the missed detection rates of off-target prediction methods, Extru-seq, and GUIDE-seq (results for sgRNAs targeting human HBBs). [Table 24]
[0541] The results related to missed detection rates in Figures 10 to 15 and Tables 11 to 17 indicate that the sensitivity of Extru-seq is far higher than that of cell-based GUIDE-seq methods, and that Extru-seq rarely misses actual off-target sites.
[0542] The results disclosed in this application indicate that GUIDE-seq missed effective off-target sites, including 1 to 6 mismatches. Therefore, only GUIDE-seq-dependent IND studies (see reference [Stadtmauer, Edward A., et al. "CRISPR-engineered T cells in patients with refractory cancer." Science 367.6481(2020):eaba7365.]) are at risk of missing effective off-target candidates. For CTX001, an in silico method was used to complement GUIDE-seq (see reference [Frangoul, Haydar, et al. "CRISPR-Cas9 gene editing for sickle cell disease and β-thalassemia." New England Journal of Medicine 384.3(2021):252-260.]). However, only genomic sites with three or fewer mismatches, or two or fewer mismatches, and a single DNA or RNA bulge are identified computationally, and there is still a risk of missing effective off-target sites with three or more mismatches.
[0543] Result 10. Extru-seq receiver operating characteristics (ROC) curve One of the most powerful tools for evaluating predictive models is the ROC curve. The ROC curve shows sensitivity and specificity on the y and x axes, respectively. ROC curves were plotted using metrics for predicting efficacy confirmation results through binary classification, including sequence read count (GUIDE-seq), DNA cleavage score (Digenome-seq, DIG-seq, and Extru-seq), CDF score (CDF), or CROP score (CROP). ROC curves for different prediction methods are shown in Figures 16 to 18. Specifically, Figures 16 to 18 show the ROC curves for the GUIDE-seq, Digenome-seq, Extru-seq, CROP, and CFD prediction methods. Figure 16(a) shows the results for a prediction method performed using sgRNA targeting human PCSK9. Figure 16(b) shows the results for a prediction method performed using sgRNA targeting human albumin. Figure 17(c) shows the results of the prediction method performed using sgRNA targeting mouse PCSK9. Figure 17(d) shows the results of the prediction method performed using sgRNA targeting mouse albumin. Figure 18(e) shows the results of the prediction method performed using sgRNA targeting human FANCF. Figure 18(f) shows the results of the prediction method performed using sgRNA targeting human VEGFA. Figure 18(g) shows the results of the prediction method performed using sgRNA targeting human HBB.
[0544] Figure 19 shows the area under the curve calculated using different methods from the results (ROC curve results) disclosed in Figures 16 to 18. Extru-seq shows an area under the curve value of 0.83, GUIDE-seq shows an area under the curve value of 0.81, DIG-seq shows an area under the curve value of 0.80, Digenome-seq shows an area under the curve value of 0.72, CROP shows an area under the curve value of 0.69, and CFD shows an area under the curve value of 0.68. Error bars indicate the standard deviation.
[0545] As described above, the area under the ROC curve was calculated, and Extru-seq showed the highest area. Specifically, Extru-seq showed an ROC area under the curve of 0.83, GUIDE-seq showed an ROC area under the curve of 0.81, DIG-seq showed an ROC area under the curve of 0.80, Digenome-seq showed an ROC area under the curve of 0.72, CROP showed an ROC area under the curve of 0.69, and CFD showed an ROC area under the curve of 0.68. The closer the ROC area under the curve is to 1, the better the model is in predicting the efficacy confirmation results. The highest ROC area under the curve for Extru-seq suggests high performance in DNA cleavage scoring. Furthermore, the use of different threshold or cutoff values may affect the predicted number of off-target sites. A high ROC area under the curve indicates a higher probability of discovering significant thresholds for Extru-seq than other methods.
[0546] Result 11. Use of Extru-seq in primary cultured cells The inventors of this application have confirmed that Extru-seq can be applied to primary cultured cells with less optimization. The GUIDE-seq method requires a high insertion rate of double-stranded oligodeoxynucleotides (dsODNs) at double-strand break (DSB) sites, which can be difficult to achieve experimentally under some cell types and experimental conditions. For example, the inventors of this application have not been able to obtain a high insertion rate of dsODNs in primary mesenchymal stem cells (MSCs) derived from bone marrow. In contrast, Extru-seq does not require dsODN insertion. Considering these advantages of Extru-seq, the inventors of this application performed Extru-seq in MSCs using the aforementioned indiscriminate sgRNAs targeting human PCSK9 and albumin.
[0547] The Venn diagrams shown in Figures 74 and 75 illustrate the difference between extru-seq results obtained from MSCs and extru-seq results obtained from HEK293T cells.
[0548] In particular, Figure 74 shows a comparison of predicted off-target sites for sgRNAs targeting human PCSK9. It shows that only 1213 off-target sites overlap between those predicted through extru-seq performed on MSCs and those predicted through extru-seq performed on HEK293T cells (showing some differences depending on the cell type).
[0549] Figure 75 shows a comparison of off-target candidates predicted using sgRNA targeting human albumin. It shows that only 26 off-target sites overlap between those predicted by extru-seq performed on MSCs and those predicted by extru-seq performed on HEK293T cells (showing some differences depending on cell type).
[0550] Tables 18 and 19 further show the results of each prediction method in MSCs and HEK293T cells. The aforementioned sgRNAs targeting PCSK9 and albumin were used as sgRNAs. In particular, Table 18 shows the top 10 candidate off-target loci predicted by extru-seq performed in MSCs for sgRNAs targeting human PCSK9. Furthermore, the ranking of the top 10 loci from extru-seq predicted by other prediction methods is shown. The top 10 loci from extru-seq were derived based on DNA cleavage scores.
[0551] Table 18 Comparison of predicted off-target candidates by cell type and method (sgRNAs targeting human PCSK9) [Table 25]
[0552] Table 19 shows the top 10 candidate off-target loci predicted by extru-seq for human albumin-targeting sgRNAs in HEK293T cells. Furthermore, Table 19 shows the ranking of the top 10 loci from extru-seq predicted by other prediction methods. The top 10 loci from extru-seq were derived based on DNA cleavage scores.
[0553] Table 19. Comparison of predicted off-target candidates by cell type and method (human albumin-targeting sgRNA) [Table 26]
[0554] As shown in Tables 18 and 19, a comparison based on the top 10 extru-seq results in MSCs, derived from sgRNAs targeting human PCSK9, confirmed that the top 10 extru-seq results (MSCs) and the top 10 extru-seq results (HEK293T) were 30% identical.
[0555] A comparison based on the top 10 extru-seq loci in MSCs, derived from results associated with sgRNAs targeting human albumin, revealed that the top 10 extru-seq loci (MSCs) and the top 10 extru-seq loci (HEK293T) were 70% identical. These results indicate that genome-wide off-target loci predicted by extru-seq vary depending on the cell type.
[0556] Analysis of the intersection of Venn diagrams using a normalized rank-sum test revealed high p-values in the studies targeting albumin-dependent sgRNAs. On the other hand, low p-values were observed in the studies targeting PCSK9-dependent sgRNAs (see Figure 76).
[0557] In particular, Figure 76 shows the p-values obtained from pairwise normalized rank-sum tests for off-target prediction methods for PCSK9 and albumin-targeting indiscriminate sgRNAs in MSCs and HEK293T cells. These results indicate that the rank of off-targets may or may not vary depending on the cell type.
[0558] Cell-based methods such as GUIDE-seq are known to miss more effective off-target candidates than in vitro and in silico methods. The inventors of this application confirmed that the miss rate of Extru-seq (2.33%) is 12.6 times lower than that of GUIDE-seq (29.5%). Furthermore, like other in vitro methods, Extru-seq was confirmed to be universally applicable to various cell types of different origins. This is because, unlike GUIDE-seq which requires dsODN insertion into DSB sites, Extru-seq does not require dsODN insertion. Extru-seq overcomes the major limitations of cell-based methods (high miss rate and need for optimization for different cell types) and the major limitations of in vitro methods (low efficacy confirmation rate and loss of cell type-specific information). In addition, the intensive execution of Extru-seq as a binary classifier of efficacy confirmation results is supported by the area under the ROC curve for Extru-seq. Therefore, extru-seq is expected to be a strong candidate as a balanced method for obtaining a comprehensive list of off-target sites in various cell types and patient-specific clinical safety studies.
[0559] Most cell-based methods use "surrogate" cell lines to predict genome-wide off-target sites in human clinical samples. However, as can be seen by comparing extru-seq results for HEK293T cells and MSCs, differences in chromatin and epigenetic state can exist between dividing in vitro cell lines and most non-dividing in vivo cells. Therefore, off-target prediction is preferably performed by extru-seq in clinically relevant cells rather than surrogate cell lines. In recent years, two cell-based methods, DISCOVER-seq (see reference [Wienert, Beeke, et al. "Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq." Science 364.6437(2019):286-289.]) and GUIDE-tag (see reference [Liang, Shun-Qing, et al. "Genome-wide detection of CRISPR editing in vivo using GUIDE-tag." Nature communications 13.1(2022):1-14.]), have been directly performed in vivo in mouse models. However, directly performing these methods in human organs for preclinical studies concerning human therapeutics is virtually impossible. Extru-seq has the advantage of being performed in primary human cells isolated from a specific patient or organ. In the experimental examples of this application, whole-genome sequencing (WGS) was used as the genomic analysis method after extrusion by extru-seq.This application is not limited to the methods disclosed in the experimental examples herein, and it is strongly anticipated that other methods may be used for post-extrusion genomic analysis (e.g., analysis at DNA cleavage sites), such as PCR-based amplification protocols (e.g., PCR-based amplification protocols used in SITE-Seq; see reference [Cameron, Peter, et al. "Mapping the genomic landscape of CRISPR-Cas9 cleavage." Nature methods 14.6(2017):600-606.]). Furthermore, it is strongly anticipated that optimization of the algorithm used (e.g., optimization of the algorithm to increase analytical sensitivity) and optimization of the extruder (e.g., optimization of the extruder size, cost, and throughput) may be performed with respect to Extru-seq optimization. Inventions developed after the filing date of this application that succeed the inventive concept of Extru-seq disclosed herein are included in the scope of this application. Furthermore, by using Extru-seq in combination with tools such as CAST-seq (see reference [Turchiano G, Andrieux G, Klermund J, Blattner G, Pennucci V, El Gaz M, Monaco G, Poddar S, Mussolino C, Cornu TI et al: Quantitative evaluation of chromosomal rearrangements in gene-edited human stem cells by CAST-Seq. Cell Stem Cell 2021, 28(6):1136-1147 e1135.]), which have been developed in recent years, Cas9-mediated large-scale deletions, chromosomal losses, and translocations can be detected.
[0560] Some of the references used herein are disclosed below. References used herein may or may not be mentioned in the paragraphs relating to the corresponding references.
[0561] References 1. Mullard A:Gene-editing pipeline takes off. Nat Rev Drug Discov 2020、19(6):367-372. 2. Tsai SQ、Zheng Z、Nguyen NT、Liebers M、Topkar VV、Thapar V、Wyvekens N、Khayter C、Iafrate AJ、Le LP et al:GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat Biotechnol 2015、33(2):187-197. 3. Liang SQ、Liu P、Smith JL、Mintzer E、Maitland S、Dong X、Yang Q、Lee J、Haynes CM、Zhu LJ et al:Genome-wide detection of CRISPR editing in vivo using GUIDE-tag. Nat Commun 2022、13(1):437. 4. Wienert B、Wyman SK、Richardson CD、Yeh CD、Akcakaya P、Porritt MJ、Morlock M、Vu JT、Kazane KR、Watry HL et al:Unbiased detection of CRISPR off-targets in vivo using DISCOVER-Seq. Science 2019、364(6437):286-289. 5. Yan WX、Mirzazadeh R、Garnerone S、Scott D、Schneider MW、Kallas T、Custodio J、Wernersson E、Li Y、Gao L et al:BLISS is a versatile and quantitative method for genome-wide profiling of DNA double-strand breaks. Nat Commun 2017、8:15058. 6. Crosetto N、Mitra A、Silva MJ、Bienko M、Dojer N、Wang Q、Karaca E、Chiarle R、Skrzypczak M、Ginalski K et al:Nucleotide-resolution DNA double-strand break mapping by next-generation sequencing. Nat Methods 2013、10(4):361-365. 7. Wang X、Wang Y、Wu X、Wang J、Qiu Z、Chang T、Huang H、Lin RJ、Yee JK:Unbiased detection of off-target cleavage by CRISPR-Cas9 and TALENs using integrase-defective lentiviral vectors. Nat Biotechnol 2015、33(2):175-178. 8. Chiarle R、Zhang Y、Frock RL、Lewis SM、Molinie B、Ho YJ、Myers DR、Choi VW、Compagno M、Malkin DJ et al:Genome-wide translocation sequencing reveals mechanisms of chromosome breaks and rearrangements in B cells. Cell 2011、147(1):107-119. 9. Petri K、Kim DY、Sasaki KE、Canver MC、Wang X、Shah H、Lee H、Horng JE、Clement K、Iyer S et al:Global-scale CRISPR gene editor specificity profiling by ONE-seq identifies population-specific、variant off-target effects. bioRxiv 2021:2021.2004.2005.438458. 10. Kim HS、Hwang GH、Lee HK、Bae T、Park SH、Kim YJ、Lee S、Park JH、Bae S、Hur JK:CReVIS-Seq:A highly accurate and multiplexable method for genome-wide mapping of lentiviral integration sites. Mol Ther Methods Clin Dev 2021、20:792-800. 11. Breton C、Clark PM、Wang L、Greig JA、Wilson JM:ITR-Seq、a next-generation sequencing assay、identifies genome-wide DNA editing sites in vivo following adeno-associated viral vector-mediated genome editing. BMC Genomics 2020、21(1):239. 12. Huang H、Hu Y、Huang G、Ma S、Feng J、Wang D、Lin Y、Zhou J、Rong Z:Tag-seq:a convenient and scalable method for genome-wide specificity assessment of CRISPR / Cas nucleases. Commun Biol 2021、4(1):830. 13. Kim D、Bae S、Park J、Kim E、Kim S、Yu HR、Hwang J、Kim JI、Kim JS:Digenome-seq:genome-wide profiling of CRISPR-Cas9 off-target effects in human cells. Nat Methods 2015、12(3):237-243、231 p following 243. 14. Kim D、Kim JS:DIG-seq:a genome-wide CRISPR off-target profiling method using chromatin DNA. Genome Res 2018、28(12):1894-1900. 15. Cameron P、Fuller CK、Donohoue PD、Jones BN、Thompson MS、Carter MM、Gradia S、Vidal B、Garner E、Slorach EM et al:Mapping the genomic landscape of CRISPR-Cas9 cleavage. Nat Methods 2017、14(6):600-606. 16. Tsai SQ、Nguyen NT、Malagon-Lopez J、Topkar VV、Aryee MJ、Joung JK:CIRCLE-seq:a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets. Nat Methods 2017、14(6):607-614.17. Lazzarotto CR、Malinin NL、Li Y、Zhang R、Yang Y、Lee G、Cowley E、He Y、Lan X、Jividen K et al:CHANGE-seq reveals genetic and epigenetic effects on CRISPR-Cas9 genome-wide activity. Nat Biotechnol 2020、38(11):1317-1327. 18. Bae S、Park J、Kim JS:Cas-OFFinder:a fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided endonucleases. Bioinformatics 2014、30(10):1473-1475. 19. Montague TG、Cruz JM、Gagnon JA、Church GM、Valen E:CHOPCHOP:a CRISPR / Cas9 and TALEN web tool for genome editing. Nucleic Acids Res 2014、42(Web Server issue):W401-407. 20. Concordet JP、Haeussler M:CRISPOR:intuitive guide selection for CRISPR / Cas9 genome editing experiments and screens. Nucleic Acids Res 2018、46(W1):W242-W245. 21. Shapiro J、Iancu O、Jacobi AM、McNeill MS、Turk R、Rettig GR、Amit I、Tovin-Recht A、Yakhini Z、Behlke MA et al:Increasing CRISPR Efficiency and Measuring Its Specificity in HSPCs Using a Clinically Relevant System. Mol Ther Methods Clin Dev 2020、17:1097-1107. 22. Gillmore JD、Gane E、Taubel J、Kao J、Fontana M、Maitland ML、Seitzer J、O'Connell D、Walsh KR、Wood K et al:CRISPR-Cas9 In Vivo Gene Editing for Transthyretin Amyloidosis. N Engl J Med 2021、385(6):493-502. 23. Maeder ML、Stefanidakis M、Wilson CJ、Baral R、Barrera LA、Bounoutas GS、Bumcrot D、Chao H、Ciulla DM、DaSilva JA et al:Development of a gene-editing approach to restore vision loss in Leber congenital amaurosis type 10. Nat Med 2019、25(2):229-233. 24. Poirot L、Philip B、Schiffer-Mannioui C、Le Clerre D、Chion-Sotinel I、Derniame S、Potrel P、Bas C、Lemaire L、Galetto R et al:Multiplex Genome-Edited T-cell Manufacturing Platform for "Off-the-Shelf"Adoptive T-cell Immunotherapies. Cancer Res 2015、75(18):3853-3864. 25. MacLeod DT、Antony J、Martin AJ、Moser RJ、Hekele A、Wetzel KJ、Brown AE、Triggiano MA、Hux JA、Pham CD et al:Integration of a CD19 CAR into the TCR Alpha Chain Locus Streamlines Production of Allogeneic Gene-Edited CAR T Cells. Mol Ther 2017、25(4):949-961. 26. Stadtmauer EA、Fraietta JA、Davis MM、Cohen AD、Weber KL、Lancaster E、Mangan PA、Kulikovskaya I、Gupta M、Chen F et al:CRISPR-engineered T cells in patients with refractory cancer. Science 2020、367(6481). 27. Goh WJ、Zou S、Ong WY、Torta F、Alexandra AF、Schiffelers RM、Storm G、Wang JW、Czarny B、Pastorin G:Bioinspired Cell-Derived Nanovesicles versus Exosomes as Drug Delivery Systems:a Cost-Effective Alternative. Sci Rep 2017、7(1):14322. 28. Kim D、Kim S、Park J、Kim JS:Genome-wide target specificities of CRISPR-Cas9 nucleases revealed by multiplex Digenome-seq. Genome Res 2016、26(3):406-415. 29. Chu VT、Weber T、Wefers B、Wurst W、Sander S、Rajewsky K、Kuhn R:Increasing the efficiency of homology-directed repair for CRISPR-Cas9-induced precise gene editing in mammalian cells. Nat Biotechnol 2015、33(5):543-548. 30. Akcakaya P、Bobbin ML、Guo JA、Malagon-Lopez J、Clement K、Garcia SP、Fellows MD、Porritt MJ、Firth MA、Carreras A et al:In vivo CRISPR editing with no detectable genome-wide off-target mutations. Nature 2018、561(7723):416-419. 31. Liu Q、Cheng X、Liu G、Li B、Liu X:Deep learning improves the ability of sgRNA off-target propensity prediction. BMC Bioinformatics 2020、21(1):51. 32. Doench JG、Fusi N、Sullender M、Hegde M、Vaimberg EW、Donovan KF、Smith I、Tothova Z、Wilen C、Orchard R et al:Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nat Biotechnol 2016、34(2):184-191. 33. Mundry R、Fischer J:Use of statistical programs for nonparametric tests of small samples often leads to incorrect P values:examples from animal behaviour. . Animal Behaviour 1998、56:256-259. 34. Dwivedi AK、Mallawaarachchi I、Alvarado LA:Analysis of small sample size studies using nonparametric bootstrap test with pooled resampling method. Stat Med 2017、36(14):2187-2205. 35. Frangoul H、Altshuler D、Cappellini MD、Chen YS、Domm J、Eustace BK、Foell J、de la Fuente J、Grupp S、Handgretinger R et al:CRISPR-Cas9 Gene Editing for Sickle Cell Disease and beta-Thalassemia. N Engl J Med 2020、384(3):252-260. 36. Turchiano G、Andrieux G、Klermund J、Blattner G、Pennucci V、El Gaz M、Monaco G、Poddar S、Mussolino C、Cornu TI et al:Quantitative evaluation of chromosomal rearrangements in gene-edited human stem cells by CAST-Seq. Cell Stem Cell 2021、28(6):1136-1147 e1135. 37. Park J、Bae S、Kim JS:Cas-Designer:a web-based tool for choice of CRISPR-Cas9 target sites. Bioinformatics 2015、31(24):4014-4016. 38. Cho SW、Kim S、Kim JM、Kim JS:Targeted genome engineering in human cells with the Cas9 RNA-guided endonuclease. Nat Biotechnol 2013、31(3):230-232. 39. Kim E、Koo T、Park SW、Kim D、Kim K、Cho HY、Song DW、Lee KJ、Jung MH、Kim S et al:In vivo genome editing with a small Cas9 orthologue derived from Campylobacter jejuni. Nat Commun 2017、8:14500. 40. Kim D、Kang BC、Kim JS:Identifying genome-wide off-target sites of CRISPR RNA-guided nucleases and deaminases with Digenome-seq. Nat Protoc 2021、16(2):1170-1192. 41. Park J、Childs L、Kim D、Hwang GH、Kim S、Kim ST、Kim JS、Bae S:Digenome-seq web tool for profiling CRISPR specificity. Nat Methods 2017、14(6):548-549. 42. DiGiusto DL、Cannon PM、Holmes MC、Li L、Rao A、Wang J、Lee G、Gregory PD、Kim KA、Hayward SB et al:Preclinical development and qualification of ZFN-mediated CCR5 disruption in human hematopoietic stem / progenitor cells. Mol Ther Methods Clin Dev 2016、3:16067. 43. A Safety and Efficacy Study Evaluating CTX110 in Subjects With Relapsed or Refractory B-Cell Malignancies(CARBON). https: / / clinicaltrials.gov / ct2 / show / NCT04035434. Accessed 15 Dec 2022. 44. Safety、Tolerability、and PK of LBP-EC01 in Patients With Lower Urinary Tract Colonization Caused by E. Coli. https: / / clinicaltrials.gov / ct2 / show / NCT04191148. Accessed 15 Dec 2022. 45. Miller JC、Paschon D、Rebar EJ:METHODS AND COMPOSITIONS FOR TREATING HEMOPHILIA. World Intellectual Property Organization 2015、WO:2015 / 089046. 46. Kwon J、Kim M、Lee J:Extru-seq:A method for predicting genome-wide off-target sites with high sensitivity. NCBI Bioproject 2022、PRJNA796642. https: / / www.ncbi.nlm.nih.gov / bioproject / ?term=PRJNA796642 [Other possible items] [Item 1] A method for identifying information about off-targets that occur during genome editing processes using the CRISPR / Cas genome editing system, (i) A step of preparing an initiation composition comprising Cas protein, guide RNA, and cells; (ii) A step of obtaining the test composition by physically destroying the cells, Here, through the physical disruption of the cells, the genomic DNA comes into contact with the Cas / gRNA complex formed by the Cas protein and the guide RNA, thereby causing the genomic DNA to cleave at one or more cleavage sites; and (iii) A step of analyzing the test composition to obtain information about one or more cleavage sites. A method that includes [a certain feature]. [Item 2] The method according to item 1, wherein the physical destruction of the cells comprises passing the cells through a filter having pores, the average diameter of the pores in the filter being smaller than the size of the cells. [Item 3] The method according to item 2, wherein the force that causes the cells to pass through the filter is pressure. [Item 4] The method according to item 2, wherein the average diameter of the pores of the filter is 5 to 15 μm. [Item 5] The method according to item 1, wherein the physical destruction of the cells is achieved through the use of an extruder equipped with a filter having pores. [Item 6] The method according to item 5, wherein the average diameter of the pores in the filter contained in the extruder is smaller than the size of the cells. [Item 7] The method according to item 5, wherein the average diameter of the pores in the filter is 5 to 15 μm. [Item 8] The information regarding the aforementioned rupture site is, The genomic DNA location of each of the one or more cleavage sites; The crack score for each of the one or more crack sites; and Number of rupture sites The method described in item 1, which includes one or more of the following. [Item 9] The aforementioned method, (iv)(iii) Step of identifying off-target candidate information from the information about the fracture site obtained from (iv)(iii) The method described in item 1, further comprising the features described in item 1. [Item 10] Information regarding the aforementioned off-target candidates is available at: Genomic DNA location of each off-target candidate with respect to one or more off-target candidates; The off-target prediction score for each off-target candidate with respect to one or more off-target candidates; and Number of predicted off-target candidates The method described in item 9, which includes one or more of the following. [Item 11] The method according to item 1, wherein the step of analyzing the test composition comprises the step of analyzing the cleaved genomic DNA contained in the test composition through sequencing. [Item 12] The method according to item 1, wherein the step of analyzing the test composition comprises the step of analyzing the cleaved genomic DNA contained in the test composition by a PCR-based method. [Item 13] The method according to item 1, wherein the membrane structure including the cell membrane is destroyed by physically destroying the cells, thereby creating an environment in which the Cas / gRNA complex can come into contact with the genomic DNA from the cells. [Item 14] The method according to item 1, wherein the membrane structure of the cell, including ...
Claims
1. A method for predicting off-target events that may occur in a genome editing process using the CRISPR / Cas genome editing system, (i) A step of preparing an initiation composition comprising Cas protein, guide RNA, and cells; (ii) A step of obtaining the test composition by physically destroying the cells, where, The physical destruction of the cells includes passing the cells through a filter having pores, wherein the average diameter of the pores in the filter is smaller than the size of the cells. Through the physical destruction of the cells, the genomic DNA comes into contact with the Cas / gRNA complex formed by the Cas protein and the guide RNA, thereby causing the genomic DNA to cleave at one or more cleavage sites, and the DNA repair mechanism of the cells is inactivated through the physical destruction of the cells; and (iii) A step of analyzing the test composition to obtain information about the cleavage site. A method that includes [a certain feature].
2. The method according to claim 1, wherein the average diameter of the holes in the filter is 10 μm or less.
3. The method according to claim 1, wherein the physical destruction of the cells is achieved by using an extruder equipped with the filter having pores.
4. The aforementioned information regarding the site of the rupture is, Number of rupture sites; The genomic DNA location of each cleavage site with respect to one or more cleavage sites; and Crack score for each of the one or more crack sites The method according to claim 1, comprising one or more of the above.
5. The method according to claim 4, wherein each of the one or more rupture sites is related to on-target or off-target.
6. The aforementioned method, Steps to identify off-target candidates from the information about the cleavage site obtained from (iv) and (iii). The method according to claim 1, further comprising:
7. Information regarding the aforementioned off-target candidates is available at: Number of off-target candidates; The genomic DNA location of each off-target candidate with respect to one or more off-target candidates; and Off-target prediction score for each off-target candidate for one or more off-target candidates The method according to claim 6, comprising one or more of the above.
8. The method according to claim 1, wherein the step of analyzing the test composition comprises the step of analyzing the cleaved genomic DNA contained in the test composition.
9. The method according to claim 8, wherein the step of analyzing the test composition comprises the step of analyzing the cleaved genomic DNA contained in the test composition through a sequencing method or a PCR-based analytical method.
10. The method according to claim 1, wherein the membrane structure including the cell membrane is destroyed by physically destroying the cells.
11. The method according to claim 1, wherein the membrane structure including the nuclear membrane of the cell is destroyed by physically destroying the cell.
12. The aforementioned method, A step to identify a predetermined CRISPR / Cas genome editing system, wherein the step to identify the predetermined CRISPR / Cas genome editing system is performed before (i). The method according to claim 1, further comprising:
13. The concentration of the Cas protein contained in the initial composition is between 10 nM and 100 μM. The method according to claim 1.
14. The concentration of the guide RNA contained in the initiation composition is between 10 nM and 100 μM. The method according to claim 1.
15. The concentration of the Cas / gRNA complex contained in the initial composition is between 10 nM and 100 μM. The method according to claim 1.
16. The concentration of the cells contained in the initial composition is 1 x 10 5 Cells / mL to 1x10 9 The method according to claim 1, wherein the amount is cells / mL.
17. The step of obtaining the test composition further comprises incubating the composition obtained through the physical disruption of the cells to accumulate the cleavage rate with respect to the genomic DNA. The method according to any one of claims 1 to 16.
18. The step of obtaining the aforementioned test composition is: The step of removing RNA from the composition obtained by destroying the aforementioned cells. It further possesses The method according to any one of claims 1 to 16.
19. The step of obtaining the aforementioned test composition is: A step of purifying DNA from the composition obtained by destroying the aforementioned cells. It further possesses The method according to any one of claims 1 to 16.
20. A method for predicting off-target events that may occur in a genome editing process using the CRISPR / Cas genome editing system, (i) A step of filling a first vessel of an extruder with an initiation composition comprising Cas protein, guide RNA, and cells; (ii) Using the extruder, (a) A step of applying pressure to the first container to move the components of the starting composition from the first container of the extruder to the second container of the extruder, Here, the components of the initial composition pass through a filter with holes located between the first and second containers due to the applied pressure, thereby filling the second container with the mixture. Here, the cells, which are components larger in size than the diameter of the pores of the filter, are destroyed by the applied pressure and pass through the pores of the filter. Here, by destroying the cells, an environment is created in which the genomic DNA can come into contact with the Cas protein and the guide RNA. As a result, the genomic DNA comes into contact with the Cas / gRNA complex, As a result, the genomic DNA is cleaved at one or more cleavage sites, Here, the DNA repair mechanism of the cells is inactivated by destroying them. A step of obtaining the test composition by performing an extrusion process including the following steps; and (iii) A step of analyzing the test composition to obtain information about the cleavage site. A method that includes [a certain feature].
21. The extrusion process described above is (b) A step of applying pressure to the second container to transfer the components of the mixture from the second container to the first container, Here, the components of the mixture pass through a filter with holes located between the first and second containers due to the applied pressure, thereby filling the first container with the mixture that has moved from the second container by passing through the filter. The method according to claim 20, further comprising:
22. The extrusion process described above is (c) The step of repeating process (a) and / or (b) a predetermined number of times. Here, the predetermined number of times is counted in increments of 0.5, where 0.5 represents the execution of a single process of (a) or (b). The method according to claim 21, further comprising:
23. The method according to claim 22, wherein the predetermined number of times is 0.5 to 10.
24. The aforementioned information regarding the site of the rupture is, Number of rupture sites; The genomic DNA location of each of the one or more cleavage sites; and Crack score for each of the one or more crack sites The method according to claim 20, comprising one or more of the above.
25. The method according to claim 24, wherein each of the one or more rupture sites is related to on-target or off-target.
26. The aforementioned method, (iv)(iii) The step of identifying information about off-target candidates from the information about the cleavage site obtained from (iv)(iii) The method according to claim 20, further comprising:
27. Information regarding the aforementioned off-target candidates is available at: Number of off-target candidates; The genomic DNA location of each off-target candidate with respect to one or more off-target candidates; and Off-target prediction score for each off-target candidate for one or more off-target candidates The method according to claim 26, comprising one or more of the above.
28. The method according to claim 20, wherein the average diameter of the pores of the filter is 10 μm or less.
29. The method according to claim 20, wherein the membrane structure including the cell membrane is destroyed by destroying the aforementioned cells.
30. The method according to claim 20, wherein the membrane structure including the nuclear membrane of the cell is destroyed by destroying the cell.
31. The method according to claim 20, wherein the concentration of the Cas protein contained in the initial composition is 10 nM to 100 μM.
32. The method according to claim 20, wherein the concentration of the guide RNA contained in the initiation composition is 10 nM to 100 μM.
33. The concentration of the cells contained in the initial composition is 1 x 10 5 From 1x10 9 The method according to claim 20, wherein the amount is cells / mL.
34. The method involves incubating a composition comprising disrupted cellular components, the Cas protein, and the guide RNA to accumulate the cleavage rate of the genomic DNA. The method according to claim 20, further comprising:
35. The method according to claim 20, further comprising the step of removing the RNA component from a composition comprising the destroyed cellular component, the Cas protein, and the guide RNA.
36. The method according to claim 20, further comprising the step of purifying the DNA of a composition comprising the disrupted cellular components, the Cas protein, and the guide RNA.
37. (iii) The method according to any one of claims 20 to 36, wherein the step of analyzing the test composition comprises the step of analyzing the cleaved genomic DNA contained in the test composition.
38. (iii) The method according to any one of claims 20 to 36, wherein the step of analyzing the test composition comprises analyzing the cleaved genomic DNA contained in the test composition through a sequencing method or a PCR-based analytical method.
Citation Information
Patent Citations
Method for detecting off-target positions of programmable nucleases in the genome
JP2017533724A
Protein Delivery in Primitive Hematopoietic Cells
JP2018503387A
Methods for assessing nuclease cleavage
JP2020503056A
A method for identifying non-targeted positions in base editing by DNA single-strand breaks
JP2020505062A
Homologous recombination in mismatch repair inactivated eukaryotic cells
US20030221208A1