Artificial guide RNAs and their use for optimized CRISPR / Cas12f1 systems
The CRISPR/Cas12f1 system is enhanced with a U-rich tail sequence to improve gene editing efficiency, addressing its low activity against double-stranded DNA, achieving efficient editing of both single- and double-stranded DNA.
Patent Information
- Application Number
- JP2022520279
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-10-29
- Filing Date
- 2020-10-29
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2040-10-29
AI Technical Summary
The CRISPR/Cas12f1 system, particularly the CRISPR/Cas14a system, exhibits low or no cleavage activity against double-stranded DNA, limiting its application in gene editing technology.
An artificial CRISPR/Cas12f1 system is developed with a U-rich tail sequence in the guide RNA, enhancing its gene editing efficiency by improving the cleavage activity on both single-stranded and double-stranded DNA.
The artificial CRISPR/Cas12f1 system with a U-rich tail sequence demonstrates significantly higher gene editing efficiency compared to the wild-type system, effectively editing nucleic acids including both single-stranded and double-stranded DNA.
Smart Images

Figure 0007799912000032 
Figure 0007799912000033 
Figure 0007799912000034
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to technology in the area of using CRISPR / Cas systems, particularly the CRISPR / Cas12f1 system, for gene editing. [Background technology]
[0002] The CRISPR / Cas12f1 system is a Class 2, Type V CRISPR / Cas system. Previous research (Harrington et al., Programmed DNA destruction by CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)) first reported a CRISPR system, the archaeal CRISPR / Cas system CRISPR / Cas14. Since then, subsequent research (Karvelis et al., Nucleic Acids Research, Vol. 48, No. 9, 5017 (2020)) has classified the CRISPR / Cas14 system as a CRISPR / Cas12f system. The CRISPR / Cas12f1 system belongs to the V-F1 system, a subtype of Class 2, Type V CRISPR / Cas systems, and includes the CRISPR / Cas14a system, which utilizes Cas14a as an effector protein. The CRISPR / Cas12f1 system is characterized by the smaller size of its effector protein compared to the CRISPR / Cas9 system. However, as revealed in previous studies (Harington et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018), U.S. Patent Application Publication No. 2020 / 0190494), the CRISPR / Cas12f1 system, particularly the CRISPR / Cas14a system, exhibits the ability to cleave single-stranded DNA but has no or very low cleavage activity against double-stranded DNA, limiting its application to gene editing technology. Summary of the Invention [Problem to be solved by the invention]
[0003] One aspect of the present disclosure provides an artificial CRISPR / Cas12f1 system with improved gene editing efficiency.
[0004] One aspect of the present disclosure provides an artificial CRISPR / Cas12f1 system having a uracil (U)-rich tail sequence.
[0005] One aspect of the present disclosure provides an artificial crRNA with a U-rich tail sequence for an artificial CRISPR / Cas12f1 system.
[0006] One aspect of the present disclosure provides an artificial guide RNA having a U-rich tail sequence for an artificial CRISPR / Cas12f1 system.
[0007] One aspect of the present disclosure provides an artificial CRISPR / Cas12f1 complex having a U-rich tail sequence for an artificial CRISPR / Cas12f1 system.
[0008] One aspect of the present disclosure provides a vector having nucleic acid sequences encoding each component of the artificial CRISPR / Cas12f1 system.
[0009] One aspect of the present disclosure provides a gene editing method using an artificial CRISPR / Cas12f1 system.
[0010] One aspect of the present disclosure provides for the use of an artificial CRISPR / Cas12f1 system. [Means for solving the problem]
[0011] In one embodiment, one aspect of the present disclosure provides an artificial CRISPR RNA (crRNA) for a CRISPR / Cas12f1 system capable of editing a nucleic acid comprising a target sequence, the artificial CRISPR RNA (crRNA) comprising: CRISPR RNA repeat sequence and a spacer sequence complementary to the target sequence; a U-rich tail sequence linked to the 3' end of the spacer sequence; Including, The U-rich tail sequence is (U a N) n U b is expressed as n is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G); a is an integer of 1 to 4, n is an integer selected from 0, 1 and 2, and b is an integer of 1 to 10; The artificial CRISPR RNA contains a CRISPR RNA repeat sequence, a spacer sequence, and a U-rich tail sequence, which are linked sequentially in the 5'→3' direction.
[0012] In one embodiment, one aspect of the present disclosure provides a DNA having a sequence encoding an artificial crRNA. In one embodiment, the U-rich tail sequence can be UUUUUU, UUUUAUUUUUU, or UUUUGUUUUUU. In one embodiment, the CRISPR RNA repeat sequence can be the sequence of SEQ ID NO: 58.
[0013] In one embodiment, one aspect of the present disclosure provides an artificial guide RNA for a CRISPR / Cas12f1 system capable of editing a nucleic acid comprising a target sequence, the artificial guide RNA comprising: tracrRNA sequence, and a crRNA sequence comprising a CRISPR RNA repeat sequence and a spacer sequence complementary to the target sequence; a U-rich tail sequence linked to the 3' end of the spacer sequence; Including, The U-rich tail sequence is (U a N) n U b is expressed as N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G); a is an integer of 1 or more and 4 or less; n is an integer selected from 0, 1 and 2; and b is an integer of 1 or more and 10 or less.
[0014] In one embodiment, the artificial guide RNA further comprises a linker sequence through which the tracrRNA sequence and the crRNA sequence are linked. In one embodiment, the linker sequence can be 5'-gaaa-3'. In one embodiment, the U-rich tail sequence can be UUUUUU, UUUUAUUUUUU, or UUUUGUUUUUU.
[0015] In one embodiment, one aspect of the present disclosure provides a DNA having a sequence encoding an artificial guide RNA.
[0016] In one embodiment, one aspect of the present disclosure provides an artificial CRISPR / Cas12f1 complex capable of editing a nucleic acid comprising a target sequence, comprising: Cas12f1 protein belonging to the CRISPR14 system, An artificial guide RNA containing a scaffold sequence that interacts with the Cas12f1 protein, a spacer sequence complementary to the target sequence, and a U-rich tail sequence. Including, The U-rich tail sequence is (U a N) n U b is expressed as N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G); a is an integer of 1 or more and 4 or less; n is an integer selected from 0, 1 and 2; and b is an integer of 1 or more and 10 or less.
[0017] In one embodiment, one aspect of the present disclosure provides a vector for expressing an artificial CRISPR / Cas12f1 complex capable of editing a nucleic acid comprising a target sequence, the vector comprising: a first sequence comprising a sequence encoding a Cas12f1 protein; a first promoter sequence operably linked to the first sequence; a second sequence comprising a sequence encoding an artificial guide RNA, The artificial guide RNA has a scaffold sequence that interacts with the Cas12f1 protein, a spacer sequence that is complementary to a target sequence, and a U-rich tail sequence linked to the 3' end of the spacer sequence, and the U-rich tail sequence has a target sequence that is complementary to the scaffold sequence that interacts with the Cas12f1 protein; The U-rich tail sequence is (U a N) n U b is expressed as N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G); a second sequence, wherein a is an integer between 1 and 4, n is an integer selected from 0, 1, and 2, and b is an integer between 1 and 10, inclusive; a second promoter sequence operably linked to the second sequence; Including, The vector is constructed to express the Cas12f1 protein and the artificial guide RNA, so that the Cas12f1 protein and the artificial guide RNA can form a CRISPR / Cas12f1 complex in the cell; Nucleic acids containing target sequences can be edited by the CRISPR / Cas12f1 complex.
[0018] In one embodiment, the second promoter sequence may be a U6 promoter sequence. In one embodiment, the vector may be a plasmid vector. In one embodiment, the vector may be a viral vector.
[0019] In one embodiment, the vector may be one or more selected from the group consisting of retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, and herpes simplex virus, hi one embodiment, the vector may be a linear PCR amplicon.
[0020] One aspect of the present disclosure provides a vector for expressing an artificial CRISPR / Cas12f1 complex capable of editing a nucleic acid comprising a first target sequence and a second target sequence, the vector comprising: a first sequence comprising a sequence encoding a Cas12f1 protein; a first promoter sequence operably linked to the first sequence; a second sequence comprising a sequence encoding the first artificial guide RNA, The first artificial guide RNA has a first scaffold sequence that interacts with the Cas12f1 protein, a first spacer sequence that is complementary to the first target sequence, and a first U-rich tail sequence; The first U-rich tail sequence is (U a N) n U b is expressed as N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). a second sequence, wherein a is an integer between 1 and 4, n is an integer selected from 0, 1, and 2, and b is an integer between 1 and 10, inclusive; a second promoter sequence operably linked to the second sequence; a third sequence comprising a sequence encoding a second artificial guide RNA, the second artificial guide RNA has a second scaffold sequence capable of interacting with the Cas12f1 protein, a second spacer sequence complementary to the second target sequence, and a second U-rich tail sequence; The second U-rich tail sequence is (U a N) n U b is expressed as N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G); a third sequence, wherein a is an integer between 1 and 4, n is an integer selected from 0, 1, and 2, and b is an integer between 1 and 10, inclusive; a third promoter sequence operably linked to a third sequence; Includes.
[0021] In one embodiment, the second promoter sequence and the third promoter sequence may be the same promoter sequence.
[0022] In one embodiment, the second promoter sequence may be an H1 promoter sequence and the third promoter sequence may be a U6 promoter sequence.
[0023] In one embodiment, an aspect of the present disclosure provides a method for editing a nucleic acid comprising a target sequence in a cell, the method comprising: Delivering a Cas12f1 protein or a nucleic acid encoding a Cas12f1 protein and an artificial guide RNA or a nucleic acid encoding an artificial guide RNA into a cell; Then, the CRISPR / Cas12f1 complex is formed inside the cell, The nucleic acid containing the target sequence is edited by the CRISPR / Cas12f1 complex, The artificial guide RNA has a scaffold sequence that interacts with the Cas12f1 protein, a spacer sequence that is complementary to the target sequence, and a U-rich tail sequence linked to the 3' end of the spacer sequence; The U-rich tail sequence is (U a N) n U b is expressed as N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G); a is an integer of 1 or more and 4 or less; n is an integer selected from 0, 1 and 2; and b is an integer of 1 or more and 10 or less.
[0024] In one embodiment, during delivery, the Cas12f1 protein and the artificial guide RNA may be introduced into the cell in the form of a CRISPR / Cas12f1 complex.
[0025] In one embodiment, for delivery, a vector containing a nucleic acid encoding the Cas12f1 protein and a nucleic acid encoding an artificial guide RNA may be introduced into the cell.
[0026] In one embodiment, the vector may be selected from the group consisting of retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, and herpes simplex virus.
[0027] In one embodiment, an aspect of the present disclosure provides a method for editing a nucleic acid comprising a target sequence in a cell, the method comprising: contacting a CRISPR / Cas12f1 complex with a nucleic acid comprising a target sequence; The CRISPR / Cas12f1 complex contains the Cas12f1 protein and an artificial guide RNA. The artificial guide RNA comprises a scaffold sequence that interacts with the Cas12f1 protein, a spacer sequence that is complementary to the target sequence, and a U-rich tail sequence; The U-rich tail sequence is (U a N) n U b is expressed as N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). a is an integer of 1 or more and 4 or less; n is an integer selected from 0, 1 and 2; and b is an integer of 1 or more and 10 or less. [Effects of the Invention]
[0028] When used for gene editing, the artificial CRISPR / Cas12f1 system having a U-rich tail sequence provided in the present disclosure exhibits higher gene editing efficiency than the wild-type CRISPR / Cas12f1 system. [Brief explanation of the drawings]
[0029] [Figure 1]Figure 1 shows an example of the structure of an artificial guide RNA provided in the present disclosure. (a) shows an example of the structure of a dual guide RNA containing a portion of tracrRNA and a repeat sequence of CRISPR RNA (crRNA). The complementary sequences bind to each other to form a double-stranded RNA. The names of the crRNA parts (CRISPR RNA repeat sequence portion, spacer sequence portion, and U-rich tail sequence portion) are shown. (b) shows an example of the structure of a single guide RNA in which tracrRNA and crRNA are linked via a linker. The positions of the CRISPR RNA repeat sequence portion, spacer sequence portion, and U-rich tail sequence portion are shown. Each sequence shown in the figure is an example of a sequence.
[0030] [Figure 2] Graphs showing the results of Experimental Example 3. (a) Graph showing the indel efficiency of the CRISPR / Cas12f1 system with a 3'-U4AU6 U-rich tail crRNA sequence against DYTarget2, DYtarget10, and DYtarget13 in HEK293T cells. In this regard, the indel efficiency of the artificial CRISPR / Cas12f1 system with a U-rich tail sequence is significantly higher than that of the wild-type CRISPR / Cas12f1 system. (b) Graph showing the indel efficiency of Cas12f1 guide RNAs with either a U6 or U4AU6 sequence against Intergenic-22 in HEK293T cells.
[0031] [Figure 3] 10 is a graph showing the results of Experimental Example 4. This shows the indel efficiency of the artificial CRISPR / Cas12f1 system for target 1 (DYtarget2), target 2 (DYtarget10), and target 3 (Intergenic-22) in HEK293T cells. No crRNA, canonical gRNA, and gRNA (MS1) refer to the indel efficiency of the control without crRNA, the wild-type CRISPR / Cas12f1 system, and the artificial CRISPR / Cas12f1 system with crRNA containing the U4AU6 structure, respectively.
[0032] [Figure 4] 10 is a graph showing the results of Experimental Example 7. This shows the indel efficiency of the CRISPR / Cas12f1 system for DYtarget2 or DYtarget10 in HEK293T cells. (a) shows the indel efficiency of DYtarget2 in HEK293T cells. WT indicates the indel efficiency of the experimental group without sgRNA treatment. Ux indicates the indel efficiency of the experimental group in which U was not added to the 3' end of the spacer sequence of the sgRNA. U6, U4AU6, U3AU6, and U3AU3AU6 refer to the 3' tail sequence of the sgRNA.
[0033] [Figure 5] 10 is a graph showing the results of Experimental Example 8. The graph shows the indel efficiency of the artificial CRISPR / Cas12f1 system against DYtarget2 in HEK293T cells. "None" indicates the indel efficiency of the experimental group in which U was not added to the sgRNA. "T," "T2," "T3," "T4," "T4A," "T4AT," "T4AT2," "T4AT3," "T4AT4," "T4AT5," "T4AT6," "T4AT7," "T4AT8," "T4AT4AT," and "T4AT4AT2" refer to the indel efficiency of the experimental group in which U, U2, U3, U4, U4A, U4AU, U4AU2, U4AU3, U4AU4, U4AU5, U4AU6, U4AU7, U4AU8, U4AU4AU, and U4AU4AU2 structures were added to the 3'-end sequence, respectively.
[0034] [Figure 6] 10 is a graph showing the results of Experimental Example 9. This shows the indel efficiency of the CRISPR / Cas12f1 system for CSMD1 in HEK293T cells. "None" indicates the indel efficiency of the experimental group in which U was not added to the sgRNA. "U," "U2," "U3," "U4," "U5," "U6," "U4A," "U4AU," "U4AU2," "U4AU3," "U4AU4," "U4AU5," "U4AU6," "U4AU7," "U4AU8," "U4AU9," and "U4AU10" indicate the indel efficiency of the experimental group in which the sgRNA contains the corresponding sequence at the 3' end.
[0035] [Figure 7] 10 is a graph showing the results of Experimental Example 10. The graph shows the indel efficiency of the CRISPR / Cas12f1 system against DYtarget2 in HEK293T cells. "None" refers to the indel efficiency of the experimental group in which no U was added to the sgRNA. "T4AT6" and "T4GT6" refer to the indel efficiency of the experimental group in which an artificial sgRNA with a U-rich tail sequence of the U4AU6 structure was used and the experimental group in which an artificial sgRNA with a U-rich tail sequence of the U4GU6 structure was used, respectively.
[0036] [Figure 8] Graphs showing the results of Experimental Example 11, which show the indel efficiency of the artificial CRISPR / Cas12f1 system against DYtarget2 and DYtarget10 in HEK293T cells. (a) shows the indel efficiency of all experimental groups of sgRNAs with U4AU6 U-rich tail sequences against DYtarget2 in HEK293T cells. In this regard, the 18-mer, 19-mer, 20-mer, 21-mer, 25-mer, and 30-mer on the horizontal axis of the graph indicate the length of the spacer sequence of the sgRNA. The spacer sequences are listed in Table 25 of Experimental Example 11. (b) shows the indel efficiency of all experimental groups of sgRNAs with U4AU6 U-rich tail sequences against DYtarget10 in HEK293T cells. In this regard, the 17-mer, 18-mer, 19-mer, 20-mer, 21-mer, 25-mer, and 30-mer on the horizontal axis of the graph indicate the length of the spacer sequence of the sgRNA. The spacer sequences are shown in Table 24 of Example 11.
[0037] [Figure 9] The results of Western blot analysis using the HA antibody in Experimental Example 12-1 are shown. Expression of the AsCas12f1 and SpCas12f1 effector proteins in HEK293T cells was confirmed. The AsCas12f1 and SpCas12f1 proteins exhibit molecular weights of 56.55 kDa and 64.65 kDa, respectively. DETAILED DESCRIPTION OF THE INVENTION
[0038] Definition of Terms about
[0039] As used herein, the term "about" refers to an amount, level, value, number, frequency, percentage, dimension, size, quantity, weight or length that varies to the extent of 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1% of the reference amount, level, value, number, frequency, percentage, dimension, size, quantity, weight or length. Operatively linked
[0040] As used herein, the term "operably linked" refers to a state in gene expression technology where certain components are linked so as to allow other components to function in an intended manner. For example, when a promoter sequence is operably linked to a coding sequence, it means that the promoter is linked so as to affect the intracellular transcription and / or expression of the coding sequence. Furthermore, this term includes all meanings that can be recognized by those skilled in the art and can be interpreted appropriately depending on the context. Target gene or target nucleic acid
[0041] As used herein, "target gene" or "target nucleic acid" basically refers to a gene or nucleic acid in a cell that can be targeted by a CRISPR system. Target gene or target nucleic acid can be used interchangeably and can refer to the same object. Unless otherwise specified, target gene or target nucleic acid can refer to a gene or nucleic acid that is inherent to a target cell, or to a gene or nucleic acid of external origin, and is not particularly limited as long as it can be a target for gene editing. Target gene or target nucleic acid can be single-stranded DNA, double-stranded DNA, and / or RNA. Furthermore, this term includes all meanings that can be recognized by those skilled in the art and can be interpreted appropriately according to the context. Target sequence
[0042] As used herein, the term "target sequence" refers to a specific sequence that a CRISPR / Cas complex recognizes to cleave a target gene or target nucleic acid. The target sequence can be appropriately selected depending on the purpose. Specifically, the term "target sequence" refers to a sequence contained in a target gene or target nucleic acid sequence that is complementary to a spacer sequence contained in a guide RNA or artificial guide RNA provided herein. Generally, the spacer sequence is determined by considering the target gene or target nucleic acid sequence and the PAM sequence recognized by the effector protein of the CRISPR / Cas system. The term "target sequence" may refer to only the specific strand that complementarily binds to the guide RNA of the CRISPR / Cas complex, or it may refer to the entire target duplex containing the specific strand portion, depending on the context. Furthermore, this term encompasses all meanings that can be recognized by those skilled in the art and can be interpreted appropriately depending on the context. vector
[0043] As used herein, unless otherwise specified, the term "vector" refers collectively to any material capable of transporting genetic material into a cell. For example, a vector may be, but is not limited to, a DNA molecule containing a desired genetic material, such as a nucleic acid encoding an effector protein of a CRISPR / Cas system and / or a nucleic acid encoding a guide RNA. This term includes all meanings that can be recognized by those skilled in the art and can be interpreted appropriately depending on the context. artificial
[0044] As used herein, "artificial" is a term used to distinguish substances, molecules, etc. that have a structure that already exists in nature, and refers to substances, molecules, etc. that have been artificially modified. For example, a guide RNA whose composition has been artificially modified from a guide RNA that exists in nature may be referred to as an "artificial guide RNA." Furthermore, this term includes all meanings that can be recognized by those skilled in the art and can be interpreted appropriately depending on the context. Nuclear localization sequence or signal (NLS)
[0045] As used herein, "NLS" refers to a peptide of a certain length that binds to a protein or its sequence to serve as a kind of "tag" when transporting extranuclear substances into the nucleus through nuclear transport. Specifically, the NLS may be an NLS from the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 237), an NLS from nucleoplasmin (e.g., a nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 238)), a c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 239) or RQRRNELKRSP (SEQ ID NO: 240), an hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 241), an IBB domain from importin alpha having the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 242), an IBB domain from importin alpha having the sequence VSRKRPRP (SEQ ID NO: 243) and PPKKARED (SEQ ID NO: 244) of the sarcoma T protein, an NLS from human p53 having the sequence PQPKKKPL (SEQ ID NO: 245), a mouse c-abl The NLS sequence may be, but is not limited to, the NLS sequence derived from the sequence SALIKKKKKMAP (SEQ ID NO: 246) of SEQ ID NO: IV, the sequences DRLRR (SEQ ID NO: 247) and PKQKKRK (SEQ ID NO: 248) of influenza virus NS1, the sequence RKLKKKIKKL (SEQ ID NO: 249) of hepatitis virus delta antigen, the sequence REKKKFLKRR (SEQ ID NO: 250) of mouse Mx1 protein, the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 251) of human poly(ADP-ribose) polymerase, or the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 252) of steroid hormone receptor (human) glucocorticoid. As used herein, the term "NLS" includes all meanings that can be recognized by those skilled in the art and can be interpreted appropriately depending on the context. Nuclear export sequence or signal (NES)
[0046] As used herein, "NES" refers to a peptide of a certain length that binds to a protein or its sequence to serve as a kind of "tag" when transporting a substance in the cell nucleus out of the nucleus by nuclear transport. As used herein, the term "NES" includes all meanings that can be recognized by those skilled in the art and can be interpreted appropriately depending on the context. tag
[0047] As used herein, the term "tag" collectively refers to a functional domain added to a peptide or protein to facilitate tracking and / or isolation / purification. Specifically, tags include, but are not limited to, tag proteins such as histidine (His) tag, V5 tag, FLAG tag, influenza hemagglutinin (HA) tag, Myc tag, VSV-G tag, and thioredoxin (Trx); autofluorescent proteins such as green fluorescent protein (GFP), yellow fluorescent protein (YFP), cyan fluorescent protein (CFP), blue fluorescent protein (BFP), HcRED, and DsRed; and reporter genes such as glutathione-S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, and luciferase. As used herein, the term "tag" encompasses all meanings recognized by those skilled in the art and may be interpreted appropriately depending on the context. Background Technology - CRISPR / Cas12f1 System overview
[0048] Various types of CRISPR / Cas systems exist in nature, and new CRISPR / Cas systems are still being discovered. The CRISPR / Cas12f1 system provided in this disclosure specifically belongs to the Class 2, Type V classification and is a CRISPR / Cas system belonging to the Cas14 family (Harrington et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)). The following describes the classification of the CRISPR / Cas12f1 system provided in this disclosure. Class 2 CRISPR / Cas systems
[0049] CRISPR / Cas systems are broadly divided into class 1 and class 2. Class 2 CRISPR / Cas systems are characterized by their effector complexes containing a large single protein with multiple domains. The most representative class 2 CRISPR / Cas system is the type II CRISPR / Cas9 system, and CRISPR / Cas systems actively studied for gene editing, such as the CRISPR / Cpf1 system, generally belong to class 2. Type V CRISPR / Cas system
[0050] Class 2 CRISPR / Cas systems are divided into types II, V, and VI. The CRISPR / Cas12f1 system provided herein belongs to the type V CRISPR / Cas system. The effector protein of type V CRISPR / Cas systems is named Cas12, and depending on the specific classification, it is named Cas12a, Cas12b, etc. Cas12 protein has one nuclease domain (RuvC-like nuclease), which distinguishes it from type II effector proteins (e.g., Cas9 protein) that have two nuclease domains (HNH and RuvC domains). To date, type V CRISPR / Cas systems have been divided into 11 subtypes, of which the CRISPR / Cas12f1 system provided herein belongs to V-F1, a variant of the VF subtype (Makarova et al., Nature Reviews, Microbiology, Volume 18, 67 (2020)). CRISPR / Cas12f1 system
[0051] The CRISPR / Cas12f system belongs to the VF subtype of type V CRISPR / Cas systems, which is further subdivided into variants V-F1 to V-F3. CRISPR / Cas12f systems include CRISPR / Cas14 systems containing variants of the effector protein Cas14, previously named Cas14 (Harrington et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)). Among these, CRISPR / Cas14a systems containing the Cas14a effector protein are classified as CRISPR / Cas12f1 systems (Makarova et al., Nature Reviews, Microbiology volume 18, 67 (2020)). However, CRISPR / Cas12f1 includes not only the Cas14a family (Harrington et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)), but also the CRISPR / Cas system whose effector system is named c2c10 (Karvelis et al., Nucleic Acids Research, Vol. 48, No. 9 5017 (2020), Makarova et al., Nature Reviews, Microbiology volume 18, 67 (2020)). Therefore, in the present disclosure, unless otherwise specified, the CRISPR / Cas12f1 system is a concept that includes both the CRISPR / Cas14a system and the CRISPR / c2c10 system, and when the CRISPR / Cas14a system is mentioned, it refers to a "CRISPR / Cas14a system" or a "CRISPR / Cas12f1 system belonging to the Cas14 family." This term has a meaning that can be appropriately interpreted by those skilled in the art depending on the context. Background Art - Design of CRISPR / Cas System Expression Vectors overview
[0052] To use the CRISPR / Cas system for gene editing, a widely used method is to express each component of the CRISPR / Cas system in cells by introducing a vector containing a sequence encoding each component of the CRISPR / Cas system into cells. Below, we will explain the components of the vector that express the CRISPR / Cas system in cells. Nucleic acids encoding components of the CRISPR / Cas system
[0053] Since the purpose of the vector is to express each component of the CRISPR / Cas system in cells, it is desirable that the vector sequence essentially contains one or more nucleic acid sequences encoding each component of the CRISPR / Cas system. Specifically, the vector sequence includes a nucleic acid sequence encoding a guide RNA and / or a Cas protein contained in the CRISPR / Cas system to be expressed. In this regard, depending on its purpose, the vector sequence may include not only a nucleic acid sequence encoding a wild-type guide RNA and a wild-type Cas protein, but also a nucleic acid sequence encoding an artificial guide RNA and a codon-optimized Cas protein, or a nucleic acid sequence encoding an artificial Cas protein. Regulatory / Control Components
[0054] To express the vector in cells, it is desirable to include one or more regulatory / controlling components. Specifically, regulatory / controlling components may include, but are not limited to, a promoter, an enhancer, an intron, a polyadenylation signal, a Kozak consensus sequence, an internal ribosome entry site (IRES), a splice acceptor, a 2A sequence, and / or a replication origin. In this regard, the replication origin may be, but is not limited to, an f1 replication origin, an SV40 replication origin, a pMB1 replication origin, an adenovirus replication origin, an AAV replication origin, and / or a BBV replication origin. promoter
[0055] In order to express the expression target of the vector in cells, it is desirable to activate the RNA transcription factor in cells by operably linking a promoter sequence to the sequence encoding each component.The promoter sequence can be designed differently depending on the corresponding RNA transcription factor or expression environment, and is not limited as long as the promoter sequence can appropriately express the components of the CRISPR / Cas system in cells.The promoter sequence can be a promoter that promotes the transcription of RNA polymerase (e.g., RNA Pol I, Pol II, or Pol III). For example, the promoter may be, but is not limited to, one of the following: SV40 early promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter, adenovirus major late promoter (Ad MLP), herpes simplex virus (HSV) promoter, cytomegalovirus (CMV) promoter such as the CMV immediate early promoter region (CMVIE), Rous sarcoma virus (RSV) promoter, human U6 micronucleus promoter (U6) (Miyagishi et al., Nature Biotechnology 20, 497-500 (2002)), enhanced U6 promoter (e.g., Xia et al., Nucleic Acids Res. 2003 Sep 1; 31 (17)), and human H1 promoter (H1). Termination signal
[0056] When a vector sequence contains a promoter sequence, the transcription of the sequence operably linked to the promoter is induced by an RNA transcription factor, and the sequence that induces the transcription termination of the RNA transcription factor is called a termination signal. The termination signal may vary depending on the type of promoter sequence. For example, if the promoter is a U6 promoter or an H1 promoter, the promoter recognizes a series of thymidine sequences (e.g., a TTTTTT (T6) sequence) as a termination signal. Additional Expression Components
[0057] In addition to the wild-type CRISPR / Cas system construct and / or the artificial CRISPR / Cas system, the vector may optionally contain a nucleic acid sequence encoding additional expression components intended to be expressed by those skilled in the art. For example, the additional expression component may be, but is not limited to, one of the tags described in the "Tags" section of the <Explanation of Terms> section. For example, the additional expression component may be, but is not limited to, herbicide resistance genes such as glyphosate, glufosinate ammonium, and phosphinothricin, or antibiotic resistance genes such as ampicillin, kanamycin, G418, bleomycin, hygromycin, and chloramphenicol. Expression vector form The expression vector can be designed in the form of a linear vector or a circular vector. Limitations of related technologies
[0058] A major limitation of the CRISPR / Cas12f1 system, especially the CRISPR / Cas14a system, which belongs to the Cas14 family, is that its gene editing activity is either absent or too low. In the first published paper on the CRISPR / Cas14 system (Harrington et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)), previous researchers found that the CRISPR / Cas14 system cleaves single-stranded DNA in prokaryotic cells but not double-stranded DNA. Given that the CRISPR / Cas14 system was discovered in archaea, previous researchers assumed that early CRISPR / Cas systems only cleaved single-stranded DNA and subsequently acquired double-stranded DNA cleavage activity (e.g., the CRISPR / Cas9 system) over time. In another prior publication (US Patent Application Publication No. 2020 / 0190494), previous researchers reported that the CRISPR / Cas14 system exhibited cleavage activity against double-stranded DNA in eukaryotic cells, but the cleavage efficiency was 0.1% or less, which appears to be significantly low. Therefore, according to the prior publication, the CRISPR / Cas12f1 system, which belongs to the Cas14 family, either did not exhibit cleavage activity against double-stranded DNA in cells, or even if the CRISPR / Cas12f1 system did exhibit cleavage activity, the efficiency was very low, making it difficult to actively use the CRISPR / Cas12f1 system for gene editing. U-rich tail sequence U-rich tail sequences - Overview
[0059] One aspect of the present disclosure provides a U-rich tail sequence for improving the gene editing efficiency of the CRISPR / Cas12f1 system. The U-rich tail sequence is characterized by being essentially uridine-rich and containing one or more uridines. In one embodiment, the U-rich tail sequence may contain a sequence of 1 to 10 repeated uridines. The U-rich tail sequence may further contain additional nucleotides in addition to uridines depending on the actual usage situation and the expression environment of the artificial CRISPR / Cas12f1 system (e.g., the internal environment of a eukaryote or prokaryote). In one embodiment, the U-rich tail sequence may contain a sequence of one or more repeated UV, UUV, UUUV, and / or UUUUV. In this regard, V is any one of adenosine (A), cytidine (C), and guanosine (G). The U-rich tail sequence is characterized by being linked to the 3' end of the crRNA sequence contained in the CRISPR / Cas12f1 system. In the present disclosure, the U-rich tail sequence serves to improve the indel efficiency of the artificial CRISPR-Cas12f1 system against target nucleic acids. In this regard, the target nucleic acid may be single-stranded DNA, double-stranded DNA, and / or RNA. As used herein, the term "U-rich tail sequence" refers to an RNA sequence containing abundant uridines or a DNA sequence encoding the same, and is interpreted appropriately according to the context. In the present disclosure, the structure and effect of the U-rich tail sequence are described in detail, which will be further explained in the following paragraphs. Background of the introduction of U-rich tail sequences - Related research on the CRISPR / Cpf1 system
[0060] In a previous study, the present inventors found that adding a U-rich tail sequence to the 3' end of the crRNA sequence contained in the CRISPR / Cpf1 system significantly improved the nucleic acid cleavage efficiency of the CRISPR / Cpf1 system (Moon et al., Highly efficient genome editing by CRISPR-Cpf1 using CRISPR RNA with a uridinylate-rich 3'-overhang. Nature Commun 9, 3651 (2018)). The present inventors focused on the fact that Cpf1 belongs to the type V CRISPR / Cas system (hereinafter referred to as the CRISPR / Cas12 system) and that CRISPR / Cas12 systems share several characteristics. Therefore, we investigated whether adding a consecutive uridine sequence to the 3' end of the crRNA contained in the CRISPR / Cas12 system could improve the gene editing efficiency of the CRISPR / Cas12 complex. Background of U-rich tail sequence introduction - Effective in the CRISPR / Cas12f1 system
[0061] Through experiments, the present inventors have found that ligating a U-rich tail sequence to the 3' end of the crRNA of Cas12f1, particularly the Cas14 family of CRISPR / Cas12f1 systems, significantly improves gene editing efficiency. Furthermore, various configurations of U-rich tail sequences were investigated to determine their effects on indel efficiency, and the optimal structure of the U-rich tail sequence is provided in the present disclosure. Background of U-rich tail sequence introduction: Not applicable to all CRISPR / Cas12 systems
[0062] However, introduction of U-rich tail sequences into other CRISPR / Cas12 systems has shown that simply belonging to a Type V CRISPR / Cas system does not improve the gene editing efficiency of the CRISPR / Cas12 complex (see Experimental Example 00). Even within the same Type V group, differences in the presence of tracrRNA, Cas protein structure, and other factors may lead to different results with respect to U-rich tail sequences. To the best of our knowledge, the function and mechanism by which U-rich tail sequences added to the 3' end of crRNA in CRISPR-Cas systems improve nucleic acid cleavage efficiency remains unclear. Furthermore, it is difficult to determine whether U-rich tail sequences added to the 3' end of crRNA affect gene cleavage activity based solely on the structure of the guide RNA and / or Cas protein. Therefore, clarifying the role of U-rich tail sequences and their impact on CRISPR / Cas systems remains a topic of future research. Below, we provide a more detailed description of U-rich tail sequences introduced into CRISPR / Cas12f1 systems, particularly those belonging to the Cas14 family. U-rich tail sequence - consecutive uridine sequences
[0063] One of the important points in designing a U-rich tail sequence is that it contains a sequence rich in one or more consecutive uridines. Through experiments, the present inventors have found that introducing a U-rich tail sequence, which is a sequence of one or more consecutive uridines, into the CRISPR / Cas12f1 system improves the gene editing efficiency of the CRISPR / Cas12f1 complex. Therefore, the U-rich tail sequence provided in the present disclosure contains a sequence of one or more consecutive uridines. In one embodiment, the U-rich tail sequence may contain a sequence of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive uridines.
[0064] Structure of U-rich tail sequences - Potential problems when a single consecutive uridine sequence is included
[0065] If a U-rich tail sequence contains a long, consecutive uridine sequence (e.g., a sequence of five or more consecutive uridines), the following problems may occur when expressing crRNA by introducing a vector into cells. When designing a vector to express the aforementioned consecutive uridine sequence, consecutive thymidine sequences corresponding to the consecutive uridine sequence will inevitably be included in the vector. In this regard, some promoters used to express vectors in eukaryotic cells (e.g., U6 promoter) use consecutive thymidine sequences, such as the T6 sequence, as a termination signal, so consecutive thymidine sequences corresponding to the consecutive uridine sequence may be recognized as a termination signal. If the consecutive uridine sequence contains one to four uridines, one to four consecutive thymidine sequences will be included in the vector, and these will not function as termination signals. Therefore, there is no particular problem with expressing a U-rich tail sequence containing up to four consecutive thymidines. However, if the consecutive uridine sequence contains six or more uridines, the vector contains a T6 sequence that acts as a termination signal, and therefore, a U-rich tail sequence containing five or more consecutive uridines may not be expressed as intended. Therefore, to include a U-rich tail sequence rich in uridines while using a promoter that recognizes consecutive thymidine sequences as a termination sequence, it is necessary to avoid a structure with six or more consecutive uridines. Structure of U-rich tail sequence - Modified consecutive uridine sequence
[0066] To address these issues, the U-rich tail sequence provided herein may include a modified uridine repeat sequence in which one ribonucleoside (A, C, G) is incorporated for every one to five uridines. In one embodiment, the U-rich tail sequence may include a sequence in which one or more of UV, UUV, UUUV, UUUUV, UUUUV, and / or UUUUUV are repeated. In this context, V is one of adenosine (A), cytidine (C), and guanosine (G). The modified consecutive uridine sequence is useful in designing vectors that express artificial crRNAs. Example of U-rich tail sequence - U x form
[0067] In one embodiment, the U-rich tail sequence is U x In one embodiment, x can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In one embodiment, x can be an integer within a range between two of these values. For example, x can be an integer from 1 to 6. In one embodiment, x can be an integer from 1 to 20. In one embodiment, x can be an integer of 20 or greater. Example of U-rich tail sequence - (U a N) n U b form
[0068] In one embodiment, the U-rich tail sequence is (U a N) n U b In this regard, N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). In this regard, a is an integer from 1 to 5, and n is an integer of 0 or greater. In one embodiment, n can be an integer from 0 to 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In one embodiment, b can be an integer within a numerical range between two of these values. For example, b can be an integer from 1 to 6. Exemplary embodiments of U-rich tail sequences—(U a V) n form
[0069] In one embodiment, the U-rich tail sequence is (U a V) n In this regard, V is one of adenosine (A), cytidine (C), and guanosine (G). In this regard, a can be an integer from 1 to 4, and n can be an integer of 1 or greater.
[0070] In one embodiment, the U-rich tail sequence is (U a M) n In this regard, M is adenosine (A) or cytidine (C). In this regard, a can be an integer from 1 to 4, and n can be an integer of 1 or greater.
[0071] In one embodiment, the U-rich tail sequence can be represented by (UaR)n, where R is adenosine (A) or guanosine (G), where a can be an integer from 1 to 4, and n can be an integer of 1 or greater.
[0072] In one embodiment, the U-rich tail sequence can be represented as (UaS)n, where S is cytidine (C) or guanosine (G), where a can be an integer from 1 to 4, and n can be an integer of 1 or greater.
[0073] In one embodiment, the U-rich tail sequence is (U a A) n In this regard, a can be an integer from 1 to 4, and n can be an integer of 1 or greater.
[0074] In one embodiment, the U-rich tail sequence is (U a C) n In this regard, a can be an integer from 1 to 4, and n can be an integer of 1 or greater.
[0075] In one embodiment, the U-rich tail sequence is (U a G) n In this regard, a can be an integer from 1 to 4, and n can be an integer of 1 or greater. An example of a U-rich tail sequence is -(U a V) n U b form
[0076] In one embodiment, the U-rich tail sequence is (U a V) n U b In this regard, V is one of adenosine (A), cytidine (C), and guanosine (G). In this regard, a is an integer from 1 to 4, and n is an integer of 0 or greater. In one embodiment, n can be 1 or 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In one embodiment, b can be an integer within a numerical range between two of these values. For example, b can be an integer from 1 to 6. In one embodiment, b can be an integer from 1 to 20. In one embodiment, b can be an integer of 20 or greater.
[0077] In one embodiment, a U-rich tail sequence can be represented by (UaM)n, where M is adenosine (A) or cytidine (C). In this regard, a is an integer from 1 to 4, and n is an integer of 0 or greater. In one embodiment, n can be 1 or 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In one embodiment, b can be an integer within a range between two of these values. In one embodiment, b can be an integer from 1 to 6. In one embodiment, b can be an integer from 1 to 20. In one embodiment, b can be an integer of 20 or greater.
[0078] In one embodiment, the U-rich tail sequence is (Ua R) n In this regard, R is adenosine (A) or guanosine (G). In this regard, a is an integer from 1 to 4, and n is an integer of 0 or greater. In one embodiment, n can be 1 or 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In one embodiment, b can be an integer within a numerical range between two of these values. In one embodiment, b can be an integer from 1 to 6. In one embodiment, b can be an integer from 1 to 20. In one embodiment, b can be an integer of 20 or greater.
[0079] In one embodiment, the U-rich tail sequence is (U a S) n In this regard, S is cytidine (C) or guanosine (G). In this regard, a is an integer from 1 to 4, and n is an integer of 0 or greater. In one embodiment, n can be 1 or 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In this regard, b can be an integer within a numerical range between two of these values. In one embodiment, b can be an integer from 1 to 6. In one embodiment, b can be an integer from 1 to 20. In one embodiment, b can be an integer of 20 or greater.
[0080] In one embodiment, a U-rich tail sequence can be represented as (UaA)nUb, where a is an integer from 1 to 4 and n is an integer of 0 or greater. In one embodiment, n can be 1 or 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In this regard, b can be an integer within a range between two of these values. In one embodiment, b can be an integer from 1 to 6. In one embodiment, b can be an integer from 1 to 20. In one embodiment, b can be an integer of 20 or greater.
[0081] In one embodiment, a U-rich tail sequence can be represented as (UaC)nUb, where a is an integer from 1 to 4 and n is an integer of 0 or greater. In one embodiment, n can be 1 or 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In this regard, b can be an integer within a range between two of these values. In one embodiment, b can be an integer from 1 to 6. In one embodiment, b can be an integer from 1 to 20. In one embodiment, b can be an integer of 20 or greater.
[0082] In one embodiment, the U-rich tail sequence is (U a G) n U b In this regard, a is an integer from 1 to 4, and n is an integer of 0 or greater. In one embodiment, n can be 1 or 2. In one embodiment, b can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In this regard, b can be an integer within a numerical range between two of these values. In one embodiment, b can be an integer from 1 to 6. In one embodiment, b can be an integer from 1 to 20. In one embodiment, b can be an integer of 20 or greater. Example of U-rich tail sequence - U x and (U a V) n Combination with
[0083] In one embodiment, the U-rich tail sequence is U x and the array represented by (U a V) n In one embodiment, the U-rich tail sequence is (U) n1 -V1-(U) n2 -V2-U x In this regard, V1 and V2 may each be one of adenosine (A), cytidine (C), and guanosine (G). In this regard, n1 and n2 may each be an integer from 1 to 4. In this regard, x may be an integer from 1 to 20. U-rich tail sequence - full length
[0084] In one embodiment, the U-rich tail sequence can be 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt in length, hi one embodiment, the U-rich tail sequence can be 20 nt or more in length. U-rich tail sequence - sequence example
[0085] In one embodiment, the U-rich tail sequence can be U, UU, UUU, UUUU, UUUUU, UUUUUU, UUUAUUU, UUUAUUUUU (SEQ ID NO: 11), UUUUAU, UUUUAUU, UUUUAUUU, UUUUAUUUU, UUUUAUUUUU (SEQ ID NO: 4), UUUUAUUUUUUU (SEQ ID NO: 5), UUUGUUU, UUUGUUUGUUU (SEQ ID NO: 29), UUUUGU, UUUUGUUU, UUUUGUUUU, UUUUGUUUU, UUUUGUUUUU (SEQ ID NO: 22) or UUUUGUUUUUUU (SEQ ID NO: 23).
[0086] In one embodiment, the U-rich tail sequence may be any one selected from the group consisting of SEQ ID NOs: 3 to 57. In one embodiment, the U-rich tail sequence may be UUUUUU, UUUUAUUUUUU, or UUUUGUUUUUU. Advantage 1 of Introducing a U-Rich Tail Sequence—Improved Gene Editing Efficiency
[0087] Introducing the U-rich tail sequence provided by the present disclosure into the CRISPR / Cas12f1 system can significantly improve the gene editing efficiency of the resulting artificial CRISPR / Cas12f1 complex. Specifically, the artificial CRISPR / Cas12f1 complex exhibits double-stranded DNA cleavage activity in eukaryotic cells that cannot be cleaved by wild-type CRISPR / Cas12f1 complexes, or exhibits significantly improved double-stranded DNA cleavage activity in eukaryotic cells compared to wild-type CRISPR / Cas12f1 complexes. This has been experimentally demonstrated by the present inventors and is described in detail in the experimental examples section of the present disclosure. Advantages of introducing U-rich tail sequence 2 - Advantages of Cas12f1 can be used
[0088] As mentioned above, the Cas12f1 protein is significantly smaller than the Cas9 and Cpf1 proteins currently being actively studied. Therefore, the length of the sequence encoding the Cas12f1 protein and its guide RNA is very short, making it possible to insert the entire sequence into a single AAV vector. Furthermore, even if functional domains such as base editors or prime editors are added to the Cas12f1 protein, the entire sequence can be contained in a single AAV vector. Because AAV vectors have high intracellular delivery efficiency and are approved for human gene therapy, including the CRISPR / Cas system or a CRISPR / Cas system with appropriate functional domains added to it in a single AAV vector would enable the use of AAV vectors for a wider range of gene editing techniques. Because the entire sequence of the CRISPR / Cas9 system is very long, it is not possible to insert the CRISPR / Cas9 system into a single AAV vector. In this aspect, the CRISPR / Cas12f1 system may be a very attractive gene editing tool. However, previous studies have shown that Cas12f1 proteins, particularly those belonging to the Cas14 family, are unable to cleave double-stranded DNA or exhibit significantly reduced cleavage efficiency. However, the present inventors have demonstrated that introducing a U-rich tail sequence into the CRISPR / Cas12f1 system significantly improves the gene editing efficiency of the CRISPR / Cas12f1 complex. Therefore, the artificial CRISPR / Cas12f1 complex provided in the present disclosure, incorporating a U-rich tail sequence, exhibits high nucleic acid cleavage efficiency while retaining the advantages of the Cas12f1 protein described above, making this artificial CRISPR / Cas12f1 complex applicable to a variety of gene editing technologies. Artificial CRISPR RNA (crRNA) for the CRISPR / Cas12f1 system Artificial crRNA - Overview
[0089] One aspect of the present disclosure provides an artificial CRISPR RNA (hereinafter, crRNA) for a CRISPR / Cas12f1 system. The artificial crRNA is a crRNA in which a U-rich tail sequence is linked to the 3'-end of the crRNA sequence of the CRISPR / Cas12f1 system. In this regard, when "CRISPR RNA" or "crRNA" is described in the present disclosure without any other modification, it is a concept distinct from the artificial crRNA, and basically refers to a CRISPR RNA having a naturally occurring structure, particularly the CRISPR RNA contained in the CRISPR / Cas12f1 system. The U-rich tail sequence has the same characteristics and structure as those described in the section of <U-rich tail sequence>. The crRNA includes a CRISPR RNA repeat sequence and a spacer sequence. The spacer sequence is a sequence related to the target sequence contained in the target nucleic acid to be edited by the artificial CRISPR / Cas12f1 complex. The sequence of the artificial crRNA provided in the present disclosure includes a CRISPR RNA repeat sequence, a spacer sequence, and a U-rich tail sequence.
[0090] In one embodiment, in the sequence of the artificial crRNA, the CRISPR RNA repeat sequence, the spacer sequence, and the U-rich tail sequence may be sequentially linked in the 5'→3' direction.
[0091] In one embodiment, the spacer sequence may be a sequence complementary to the target sequence of the target nucleic acid.
[0092] In one embodiment, the U-rich tail sequence is (U a N) n U b and can be represented by. In this regard, N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). In this regard, a is an integer of 1 or more and 4 or less, n is an integer selected from 0, 1, and 2, and b is an integer of 1 or more and (end) 10 or less.
[0093] In one embodiment, the U-rich tail sequence is (U a V) n U bIn this regard, a, n, and b can be integers, where a can be 1 to 4, n can be 0 or greater, and b can be 1 to 10. Furthermore, one embodiment of the present disclosure provides DNA encoding an artificial crRNA. The construction of the artificial crRNA will be described in more detail below. CRISPR RNA Repeats – Definition
[0094] The CRISPR RNA repeat sequence is a sequence determined by the type of Cas12f1 protein in the CRISPR / Cas12f1 system and refers to the sequence linked to the 5' end of the spacer sequence. The CRISPR RNA repeat sequence may vary depending on the type of Cas12f1 protein contained in the CRISPR / Cas12f1 system. The CRISPR RNA repeat sequence may interact with at least a portion of the tracrRNA of the CRISPR / Cas12f1 system, and may also interact with the Cas12f1 protein. CRISPR RNA Repeat Sequence - Example Sequence In one embodiment, the repeat sequence of the CRISPR RNA may be the sequence of SEQ ID NO: 58. CRISPR RNA repeat sequence - sequence can be modified
[0095] As used herein, the term "CRISPR RNA repeat sequence" refers to a CRISPR RNA repeat sequence normally found in nature, but is not limited to this, and may refer to a CRISPR RNA repeat sequence in which part or all of a CRISPR RNA repeat sequence found in nature has been modified without impairing the function of the CRISPR / Cas12f1 system.
[0096] In one embodiment, the CRISPR RNA repeat sequence may be a CRISPR RNA repeat sequence in which all or part of a wild-type CRISPR RNA has been modified.
[0097] In one embodiment, the CRISPR RNA repeat sequence can be 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51% or 50% identical or homologous to a wild-type CRISPR RNA repeat sequence. In one embodiment, the CRISPR RNA repeat sequence may be identical to the wild-type CRISPR RNA repeat sequence within a numerical range between two of these values. For example, the CRISPR RNA repeat sequence may be 90% to 100% identical to the wild-type CRISPR RNA repeat sequence. In one embodiment, the CRISPR RNA repeat sequence may be 60% to 80% identical to the wild-type CRISPR RNA repeat sequence. In one embodiment, the wild-type CRISPR RNA repeat sequence may be the sequence of SEQ ID NO: 58. Spacer Sequence - Definition
[0098] The spacer sequence plays a central role in enabling the CRISPR / Cas complex to exhibit target-specific gene editing activity. As used herein, unless otherwise specified, the spacer sequence refers to the spacer sequence in the crRNA (or artificial crRNA) contained in the CRISPR / Cas12f1 complex. The spacer sequence is designed to be located at the target sequence contained in the sequence of the target nucleic acid to be edited using the CRISPR / Cas12f1 complex. That is, the artificial crRNA sequences provided in the present disclosure may have various spacer sequences depending on the target sequence. In this regard, the target nucleic acid may be single-stranded DNA, double-stranded DNA, or RNA. Spacer sequence - Relationship between target nucleic acid and target sequence
[0099] The spacer sequence is a sequence complementary to the target sequence and is linked to the 3' end of the CRISPR RNA repeat sequence. The spacer sequence is a sequence homologous to the protospacer adjacent motif (PAM) sequence recognized by the Cas12f1 protein, and has a sequence in which thymidines in the protospacer sequence are replaced with uridines. In this regard, the target sequence and the protospacer sequence are determined in the sequence adjacent to the PAM sequence contained in the target nucleic acid, and the spacer sequence is determined accordingly.
[0100] In one embodiment, the spacer sequence portion of the crRNA may bind to the target nucleic acid in a complementary manner. In one embodiment, the spacer sequence portion of the crRNA may bind to the target sequence portion of the target nucleic acid in a complementary manner. In one embodiment, when the target nucleic acid is double-stranded DNA, the spacer sequence may be a sequence complementary to a target sequence contained in the target strand of the double-stranded DNA. In one embodiment, when the target nucleic acid is double-stranded DNA, the spacer sequence may be a sequence homologous to a protospacer sequence contained in the non-target strand of the double-stranded DNA. Specifically, the spacer sequence may have the same base sequence as the protospacer sequence, but each thymidine (T) contained in the base sequence may be replaced with uracil (U). In one embodiment, the spacer sequence may be an RNA sequence corresponding to the DNA sequence of the protospacer. Spacer sequence - sequence length
[0101] In one embodiment, the length of the spacer sequence can be 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, 30 nt, 31 nt, 32 nt, 33 nt, 34 nt, 35 nt, 36 nt, 37 nt, 38 nt, 39 nt, or 40 nt. In one embodiment, the spacer sequence can have a length within a range between two of these values. For example, the length of the spacer sequence can be 17 nt to 23 nt. In one embodiment, the length of the spacer sequence can be 17 nt to 30 nt. U-rich tail sequence
[0102] The artificial crRNA is characterized in that the U-rich tail sequence is linked to the 3'-terminal portion of the spacer sequence contained in the crRNA sequence. The U-rich tail sequence has the same composition and characteristics as those described in the section of <U-rich tail sequence>. DNA encoding artificial crRNA
[0103] One aspect of the present disclosure provides DNA encoding artificial crRNA. In one embodiment, the DNA sequence may have a sequence in which all uridines are substituted with thymidines in the sequence of the artificial crRNA. Use of artificial crRNA
[0104] The artificial crRNA, together with the tracrRNA for Cas12f1, may constitute an artificial guide RNA for the CRISPR / Cas12f1 system. The artificial guide RNA containing the artificial crRNA may bind to the Cas12f1 protein to form a CRISPR / Cas12f1 complex. Artificial guide RNA for the CRISPR / Cas12f1 system Artificial guide RNA - Overview
[0105] One aspect of the present disclosure provides an artificial guide RNA for the CRISPR / Cas12f1 system. The sequence of the artificial guide RNA includes a tracrRNA sequence, a crRNA sequence, and a U-rich tail sequence. In this regard, the U-rich tail sequence is linked to the 3'-end of the crRNA sequence. That is, the sequence of the artificial guide RNA includes a tracrRNA sequence and an artificial crRNA sequence. The crRNA includes a repeat sequence and a spacer sequence of CRISPR RNA. In this regard, the CRISPR RNA repeat sequence and the spacer sequence have the same characteristics and structures as each component described in the section of <Artificial CRISPR RNA for the CRISPR / Cas12f1 system>. The U-rich tail has the same characteristics and structures as each component described in the section of <U-rich tail sequence>.
[0106] In one embodiment, the U-rich tail sequence is (U a N) n U b In this regard, N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). In this regard, a is an integer from 1 to 4, n is an integer selected from 0, 1, and 2, and b is an integer from 1 to 10.
[0107] In one embodiment, the U-rich tail sequence is (U a V) n U b where a, n, and b are integers, a can be 1 to 4, inclusive, n can be 0 or greater, and b can be 1 to 10, inclusive.
[0108] At least a portion of the tracrRNA and at least a portion of the crRNA contained in the artificial guide RNA can have complementary sequences to form a double-stranded RNA. Specifically, at least a portion of the tracrRNA and at least a portion of the CRISPR RNA repeat sequence contained in the crRNA sequence can be complementary to each other to form a double-stranded RNA. The artificial guide RNA can bind to the Cas12f1 protein to form a CRISPR / Cas12f1 complex, recognize a target sequence complementary to the spacer sequence contained in the crRNA sequence, and enable the CRISPR / Cas12f1 complex to edit a target nucleic acid containing the target sequence. tracrRNA-definition
[0109] The tracrRNA, together with the crRNA, is an important component that constitutes the guide RNA of the CRISPR / Cas system. As used herein, unless otherwise specified, the term tracrRNA refers to the tracrRNA that constitutes the guide RNA of the CRISPR / Cas12f1 system. At least a portion of the tracrRNA can bind to at least a portion of the crRNA in a complementary manner to form a duplex. More specifically, a partial sequence of the tracrRNA has a sequence complementary to all or a portion of the CRISPR RNA repeat sequence contained in the crRNA. tracrRNA - Complementarity between tracrRNA and CRISPR RNA repeats
[0110] In one embodiment, the sequence of the tracrRNA may comprise a sequence that is 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% complementary to the repeat sequence of the CRISPR RNA. In one embodiment, the sequence of the tracrRNA may comprise a sequence that is complementary to the repeat sequence of the CRISPR RNA within a numerical range between two of these values. For example, the sequence of the tracrRNA may comprise a sequence that is 60% to 90% complementary to the repeat sequence of the CRISPR RNA. In one embodiment, the sequence of the tracrRNA may comprise a sequence that is 70% to 100% complementary to the repeat sequence of the CRISPR RNA. Number of mismatches in tracrRNA-tracrRNA and CRISPR RNA repeat sequences
[0111] In one embodiment, the tracrRNA sequence may comprise a complementary sequence with 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 mismatches to the CRISPR RNA repeat sequence. In one embodiment, the tracrRNA sequence may comprise a complementary sequence with 0 to 8 mismatches to the CRISPR RNA repeat sequence within the numerical range selected in the immediately preceding sentence. For example, the tracrRNA may comprise a complementary sequence with 0 to 8 mismatches to the CRISPR RNA repeat sequence. In one embodiment, the tracrRNA sequence may comprise a complementary sequence with 8 to 12 mismatches to the CRISPR RNA repeat sequence. Example tracrRNA-sequences
[0112] In one embodiment, the tracrRNA sequence may be the sequence of SEQ ID NO:60. tracrRNA-sequence can be modified
[0113] As used herein, "tracrRNA" refers to, but is not limited to, a tracrRNA normally found in nature, and may refer to a tracrRNA in which part or all of the naturally occurring tracrRNA sequence has been modified without impairing the function of the CRISPR / Cas12f1 system.
[0114] In one embodiment, the tracrRNA sequence may be a tracrRNA sequence in which all or part of the wild-type tracrRNA sequence has been modified.
[0115] In one embodiment, the tracrRNA repeat sequence can be 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51%, or 50% identical or homologous to the wild-type tracrRNA repeat sequence. In one embodiment, the tracrRNA sequence can be identical to the wild-type tracrRNA sequence within a range between two of these values. For example, the tracrRNA sequence can be 90% to 100% identical to the wild-type tracrRNA sequence. In one embodiment, the tracrRNA sequence can be 60% to 80% identical to the wild-type tracrRNA sequence. In one embodiment, the wild-type tracrRNA sequence can be the sequence of SEQ ID NO: 60. Dual guide RNAs - definition and structure
[0116] The artificial guide RNA provided in the present disclosure may be a dual guide RNA. In other words, the tracrRNA and the crRNA form separate RNA molecules. A portion of the tracrRNA and a portion of the crRNA bind complementary to each other to form a double strand. Specifically, in the dual guide RNA, a portion of the tracrRNA containing the 3' end and a portion of the crRNA containing the CRISPR RNA repeat sequence may form a double strand.
[0117] In one embodiment, the artificial guide RNA can be a dual guide RNA. In one embodiment, when the artificial guide RNA is a dual guide RNA, the tracrRNA and the crRNA are separate RNA molecules. Dual-guide RNA sequence examples
[0118] In one embodiment, the tracrRNA may have the sequence of SEQ ID NO: 60. In one embodiment, the CRISPR RNA repeat sequence contained in the crRNA may have the sequence of SEQ ID NO: 58. In one embodiment, at least a portion of the trancrRNA having the sequence of SEQ ID NO: 60 and the CRISPR RNA repeat sequence having the sequence of SEQ ID NO: 58 may bind to each other in a complementary manner to form a duplex. Single-guide RNA-defined
[0119] A single molecule of tracrRNA and crRNA designed to facilitate the formation of a guide RNA is called a single-guide RNA. The artificial guide RNA provided in the present disclosure may be a single-guide RNA. When the artificial guide RNA is a single-guide RNA, the single-guide RNA is a single RNA molecule that includes all of the tracrRNA, crRNA, and U-rich tail. In one embodiment, when the artificial guide RNA is a single-guide RNA, the sequence of the tracrRNA and the sequence of the crRNA may both be included in the sequence of a single RNA molecule. Examples of single-guide RNA linkers and sequences
[0120] In one embodiment, when the artificial guide RNA is a single-guide RNA, the sequence of the artificial guide RNA may further comprise a linker sequence, and the tracrRNA sequence and the crRNA sequence may be linked via the linker sequence. In one embodiment, the linker sequence may be 5'-gaaa-3'. Single-guide RNA structure
[0121] The single-guide RNA sequence consists of a tracrRNA sequence, a linker sequence, a crRNA sequence, and a U-rich tail sequence linked sequentially from 5' to 3'. A portion of the tracrRNA sequence and all or part of the CRISPR RNA repeat sequence contained in the crRNA sequence are complementary to each other. Therefore, a portion of the tracrRNA and the repeat sequence portion of the CRISPR RNA in the crRNA bind complementary to each other to form a double-stranded RNA. Examples of single-guide RNA sequences
[0122] The single guide RNA may have a SEQ ID NO: selected from the group consisting of SEQ ID NOs: 75-125 and 261-280. Scaffold sequence
[0123] The sequence of the artificial guide RNA provided in the present disclosure can be functionally divided into 1) a sequence portion that interacts with the Cas12f1 protein to form a CRISPR / Cas12f1 complex, 2) a sequence portion that enables the CRISPR / Cas12f1 complex to find the target nucleic acid, and 3) a U-rich tail sequence portion. In this regard, the sequence portion that interacts with the Cas12f1 protein to form the CRISPR / Cas12f1 complex can be referred to as a scaffold sequence. In this regard, the scaffold sequence collectively refers to a portion that interacts with the Cas12f1 protein and is not limited to a single RNA molecule sequence. Specifically, the scaffold sequence can include sequences of two or more RNA molecules.
[0124] In one embodiment, when the artificial guide RNA is a dual guide RNA, the scaffold sequence can include a tracrRNA sequence and a CRISPR RNA repeat sequence contained in the crRNA among the artificial guide RNA sequences.In one embodiment, the tracrRNA sequence can be a tracrRNA sequence obtained by modifying all or part of a naturally occurring tracrRNA sequence.In one embodiment, the CRISPR RNA repeat sequence can be a CRISPR RNA sequence obtained by modifying all or part of a CRISPR RNA repeat sequence found in nature.
[0125] In one embodiment, when the artificial guide RNA is a single-guide RNA, the scaffold sequence may include a tracrRNA sequence, a linker sequence, and a CRISPR RNA repeat sequence contained in the crRNA sequence. In one embodiment, the tracrRNA sequence may be a tracrRNA sequence in which all or part of a naturally occurring tracrRNA sequence has been modified. In one embodiment, the CRISPR RNA repeat sequence may be a CRISPR RNA sequence in which all or part of a CRISPR RNA repeat sequence found in nature has been modified.
[0126] In one embodiment, the scaffold sequence may be a sequence selected from the group consisting of SEQ ID NOs: 253 and 254. In one embodiment, the scaffold sequence can be a sequence that is 100%, 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, 75%, 74%, 73%, 72%, 71%, 70%, 69%, 68%, 67%, 66%, 65%, 64%, 63%, 62%, 61%, 60%, 59%, 58%, 57%, 56%, 55%, 54%, 53%, 52%, 51% or 50% identical or homologous to a sequence selected from the group consisting of SEQ ID NOs: 253 and 254. In one embodiment, the scaffold sequence can be a sequence that is identical to a sequence selected from the group consisting of SEQ ID NOs: 253 and 254, within a numerical range between two of these numbers. For example, the scaffold sequence can be a sequence that is 90% to 100% identical to a scaffold sequence selected from the group consisting of SEQ ID NOs: 253 and 254. In one embodiment, the scaffold sequence can be a sequence that is 60% to 80% identical to a scaffold sequence selected from the group consisting of SEQ ID NOs: 253 and 254. Artificial CRISPR / Cas12f1 complex CRISPR / Cas12f1 complex - overview
[0127] One aspect of the present disclosure provides an artificial CRISPR / Cas12f1 complex. The artificial CRISPR / Cas12f1 complex includes a Cas12f1 protein and an artificial guide RNA. In one embodiment, the Cas12f1 protein may be derived from the Cas14 family. The artificial guide RNA sequence includes a tracrRNA sequence, a crRNA sequence, and a U-rich tail sequence. That is, the artificial guide RNA sequence includes a scaffold sequence, a spacer sequence, and a U-rich tail sequence. The U-rich tail sequence may be linked to the 3'-end of the crRNA sequence. The U-rich tail has the same characteristics and structure as those described in the section of <U-rich tail sequence>.
[0128] In one embodiment, the U-rich tail sequence may be represented by (U a N) n U b where N is one of adenosine (A), uracil (U), cytosine (C), and guanine (G). Here, a is an integer from 1 to 4, n is an integer selected from 0, 1, and 2, and b is an integer from 1 or more to 10 or less.
[0129] In one embodiment, the U-rich tail sequence may be represented by (U a V) n U b where a, n, and b are integers, a can be from 1 to 4, n can be 0 or greater, and b can be from 1 to 10.
[0130] The scaffold sequence can interact with the Cas12f1 protein and plays an important role when the artificial guide RNA and the Cas12f1 protein form a complex. Cas12f1 Protein - Overview
[0131] The artificial CRISPR / Cas12f1 complex may comprise a Cas12f1 protein. Essentially, the Cas12f1 protein may be a wild-type Cas12f1 protein found in nature. The sequence encoding the Cas12f1 protein may be a Cas12f1 sequence that has been human codon-optimized relative to the wild-type Cas12f1 protein. Furthermore, the Cas12f1 protein may have the same function as the wild-type Cas12f1 protein found in nature. However, unless otherwise specified, the term "Cas12f1 protein" as used herein may refer not only to wild-type or codon-optimized Cas12f1 proteins, but also to modified Cas12f1 proteins or Cas12f1 fusion proteins. Furthermore, the term "Cas12f1 protein" may refer not only to those having the same function as wild-type Cas12f1 protein, but also to Cas12f1 proteins with all or part of their functions modified, Cas12f1 proteins with all or part of their functions lost, and / or Cas12f1 proteins with additional functions. The meaning of the Cas12f1 protein may be interpreted appropriately depending on the context, and unless otherwise specified, is interpreted broadly. The structure and function of the Cas12f1 protein are described in detail below. Cas12f1 protein - wild-type Cas12f1 protein
[0132] The artificial CRISPR / Cas12f1 complex may comprise a Cas12f1 protein. In one embodiment, the Cas12f1 protein may be a wild-type Cas12f1 protein. In one embodiment, the Cas12f1 protein may be derived from the Cas14 family (Harrington et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)). In one embodiment, the Cas12f1 protein may be a Cas14a protein derived from uncultured archaea (Harrington et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)). In one embodiment, the Cas12f1 protein may be a Cas14a1 protein. Cas12f1 protein - modified Cas12f1 protein
[0133] The artificial CRISPR / Cas12f1 complex may include a modified Cas12f1 protein. In this regard, modified Cas12f1 refers to Cas12f1 in which at least a portion of the sequence of a wild-type or codon-optimized Cas12f1 protein has been modified. The modification of the Cas12f1 protein may be performed on an individual amino acid basis or on a functional domain basis. In one embodiment, the protein may be modified by individually substituting, deleting, and / or adding one or more amino acids, peptides, polypeptides, proteins, and / or domains into the wild-type or codon-optimized Cas12f1 protein sequence. In one embodiment, the Cas12f1 protein may be a Cas12f1 protein in which one or more amino acids, peptides, and / or polypeptides in the RuvC domain contained in the wild-type Cas12f1 protein have been substituted, deleted, or added. Cas12f1 protein-Cas12f1 fusion protein
[0134] The artificial CRISPR / Cas12fl complexes provided herein can include a Cas12fl fusion protein. In this context, a Cas12fl fusion protein refers to a protein in which a wild-type or modified Cas12fl protein is fused with additional amino acids, peptides, polypeptides, proteins, or domains.
[0135] In one embodiment, the Cas12fl protein can be a fusion of a wild-type Cas12fl protein with a base editor and / or a reverse transcriptase. In one embodiment, the base editor can be adenosine deaminase and / or cytosine deaminase. In one embodiment, the reverse transcriptase can be Moloney murine leukemia virus (M-MLV) reverse transcriptase or a variant thereof. In this regard, the Cas12fl protein fused with the reverse transcriptase can function as a prime editor.
[0136] In one embodiment, the Cas12f1 protein can be a fusion product of wild-type Cas12f1 protein and various enzymes that can participate in intracellular gene expression process.In this regard, the Cas12f1 protein fused with enzymes can cause various quantitative and qualitative changes in intracellular gene expression.In one embodiment, the enzymes can be DNMT, TET, KRAB, DHAC, LSD and / or p300. Cas12f1 protein - Altered function
[0137] The Cas12fl protein contained in the artificial CRISPR / Cas12fl complex provided herein may have the same function as a wild-type Cas12fl protein. The Cas12fl protein contained in the artificial CRISPR / Cas12fl complex provided herein may be a Cas12fl protein with an altered function compared to a wild-type Cas12fl protein. Specifically, the alteration may involve the alteration of all or part of a function, the loss of all or part of a function, or the addition of an additional function. In one embodiment, the Cas12fl protein is not particularly limited as long as it has a modification that can be applied to Cas proteins in CRISPR / Cas systems by those skilled in the art. In this regard, it can be modified using known techniques. In one embodiment, the Cas12fl protein may be modified to cleave only one strand of a double-stranded target nucleic acid. Furthermore, the Cas12fl protein may be modified to cleave only one strand of a double-stranded target nucleic acid and perform base editing or prime editing on the uncleaved strand. In one embodiment, the Cas12fl protein may be modified so that it cannot cleave both strands of a target nucleic acid. Furthermore, the Cas12fl protein may be modified so that it cannot cleave both strands of a target nucleic acid, and may perform base editing, prime editing, or gene expression regulation functions on the target nucleic acid. Cas12f1 protein - other examples of modifications
[0138] In one embodiment, the Cas12fl protein may comprise a nuclear localization sequence (NLS) or a nuclear export sequence (NES). Specifically, the NLS may be any one of those exemplified in the NLS section of the <Definitions>, but is not limited to these. In one embodiment, the Cas12fl protein may comprise a tag. Specifically, the tag may be any one of those exemplified in the Tag section of the <Definitions>, but is not limited to these. Cas12f1 protein-PAM sequence
[0139] Two conditions are necessary for the CRISPR / Cas12f1 complex to cleave a target gene or nucleic acid. First, the target gene or nucleic acid must contain a base sequence of a certain length that can be recognized by the Cas12f1 protein. Second, a sequence that can complementarily bind to the spacer sequence contained in the guide RNA must be present around the base sequence of a certain length. When these two conditions are met—1) the Cas12f1 protein recognizes the base sequence of a certain length, and 2) the spacer sequence portion complementarily binds to the sequence portion surrounding the base sequence of a certain length—the target gene or nucleic acid is cleaved. In this regard, the base sequence of a certain length recognized by the Cas12f1 protein is called a PAM (protospacer adjacent motif) sequence. The PAM sequence is a unique sequence determined by the Cas12f1 protein. When determining the target sequence for the CRISPR / Cas12f1 complex, the target sequence must be determined in the sequence adjacent to the PAM sequence. Example of Cas12f1 protein-PAM sequence
[0140] In one embodiment, the PAM sequence of the Cas12fl protein can be a T-rich sequence. In one embodiment, the PAM sequence of the Cas12fl protein can be TTTN in the 5'→3' direction, where N is one of deoxythymidine (T), deoxyadenosine (A), deoxycytidine (C), or deoxyguanosine (G). In one embodiment, the PAM sequence of the Cas12fl protein can be TTTA, TTTT, TTTC, or TTTG in the 5'→3' direction. In one embodiment, the PAM sequence of the Cas12fl protein can be TTTA or TTTG in the 5'→3' direction. In one embodiment, the PAM sequence of the Cas12fl protein can be different from the PAM sequence of a wild-type Cas12fl protein. Cas12f1 Protein - Example Sequence
[0141] In one embodiment, the Cas12f1 protein may have an amino acid sequence selected from the group consisting of SEQ ID NOs: 281, 317, 319, 321, and 323.
[0142] In one embodiment, the DNA sequence encoding the Cas12f1 protein may be a human codon-optimized sequence.
[0143] In one embodiment, the DNA sequence encoding the Cas12f1 protein may be a DNA sequence selected from the group consisting of SEQ ID NOs: 1, 2, 316, 318, 320, and 322. Artificial guide RNA
[0144] The artificial guide RNA constituting the CRISPR / Cas12f1 complex provided in the present disclosure has the same characteristics and structure as those described in the section <Artificial guide RNA for the CRISPR / Cas12f1 system>. CRISPR / Cas12f1 complex - Example of composition
[0145] In one embodiment, the Cas12f1 protein may have the amino acid sequence of SEQ ID NO: 281, the artificial guide RNA may have an amino acid sequence selected from the group consisting of SEQ ID NOs: 75 - 125, and the Cas12f1 protein and the artificial guide RNA may bind to each other to form a CRISPR / Cas12f1 complex.
[0146] In one embodiment, the Cas12f1 protein may have an amino acid sequence selected from the group consisting of SEQ ID NOs: 317, 319, 321, and 323, the artificial guide RNA may have a sequence selected from the group consisting of SEQ ID NOs: 75 - 125 and 261 - 280, and the CRISPR / Cas12f1 complex formed by the binding of the Cas12f1 protein and the artificial guide RNA may have a base editing function. Vector for expressing the CRISPR / Cas12f1 system Overview
[0147] One aspect of the present disclosure provides a vector for expressing components of the CRISPR / Cas12f1 system. The vector is constructed to express the Cas12f1 protein and / or artificial guide RNA. The sequence of the vector may include a nucleic acid sequence encoding one of the components of the CRISPR / Cas12f1 system, and may include nucleic acid sequences encoding two or more components. The sequence of the vector includes a nucleic acid sequence encoding the Cas12f1 protein and / or a nucleic acid sequence encoding the artificial guide RNA. The sequence of the vector includes one or more promoter sequences. The promoter is operably linked to a nucleic acid sequence encoding the Cas12f1 protein and / or a nucleic acid sequence encoding the artificial guide RNA to facilitate transcription of the nucleic acid sequence in a cell. The Cas12f1 protein has similar characteristics and configurations as the Cas12f1 protein, modified Cas12f1 protein or Cas12f1 fusion protein described in the section of <CRISPR / Cas12f1 complex>. The artificial guide RNA has similar characteristics and configurations as the artificial guide RNA described in the section of <Artificial guide RNA for CRISPR / Cas12f1 system>.
[0148] The sequence of the vector may include a nucleic acid sequence encoding the Cas12f1 protein and a nucleic acid sequence encoding the artificial guide RNA. In one embodiment, the sequence of the vector may include a first sequence including a nucleic acid sequence encoding the Cas12f1 protein and a second sequence including a nucleic acid sequence encoding the artificial guide RNA. The sequence of the vector includes a promoter sequence for expressing the nucleic acid sequence encoding the Cas12f1 protein in a cell and a promoter sequence for expressing the nucleic acid sequence encoding the artificial guide RNA in a cell, and these promoters are operably linked to each expression target. In one embodiment, the sequence of the vector may include a first promoter sequence operably linked to the first sequence and a second promoter sequence operably linked to the second sequence.
[0149] The vector sequence may include a nucleic acid sequence encoding the Cas12f1 protein and a nucleic acid sequence encoding two or more different artificial guide RNAs. In one embodiment, the vector sequence may include a first sequence including a nucleic acid sequence encoding the Cas12f1 protein, a second sequence including a nucleic acid sequence encoding the artificial guide RNA, and a third sequence including a nucleic acid sequence encoding the second artificial guide RNA. Furthermore, the vector sequence may include a first promoter sequence operably linked to the first sequence, a second promoter sequence operably linked to the second sequence, and a third promoter sequence operably linked to the third sequence. Expression target - Cas12f1 protein
[0150] The vector may be constructed to express a Cas12fl protein. In this regard, the Cas12fl protein has the same structure and characteristics as those described in the section <Artificial CRISPR / Cas12fl Complex>. In one embodiment, the vector may be constructed to express a wild-type Cas12fl protein. In this regard, the wild-type Cas12fl protein may be Cas14al. In one embodiment, the vector may be constructed to express a Cas12fl protein modified to cleave only one strand of the double-stranded target nucleic acid. Furthermore, the modified Cas12fl protein may be modified to cleave only one strand of the double-stranded target nucleic acid and to perform base editing or prime editing on the uncleaved strand. In one embodiment, the Cas12fl protein may be modified to prevent cleavage of both double-stranded target nucleic acid strands. Additionally, the Cas12fl protein may be modified so that it cannot cleave both strands of the target nucleic acid, and may perform base-editing, prime-editing, or gene expression regulation functions on the target nucleic acid. Expression target - artificial guide RNA
[0151] The vector may be constructed to express an artificial guide RNA. The artificial guide RNA has the same characteristics and configuration as the artificial guide RNA described in the section "<Artificial guide RNA for CRISPR / Cas12f1 system>". In one embodiment, the vector may be constructed to express an artificial guide RNA. In one embodiment, the artificial guide RNA sequence may include a scaffold sequence, a spacer sequence, and a U-rich tail sequence. In one embodiment, the artificial guide RNA sequence may include a tracrRNA sequence, a crRNA sequence, and a U-rich tail sequence. In one embodiment, the U-rich tail sequence may be represented by (U a N) n U b . In this regard, N is one of adenosine (A), uracil (U), cytosine (C), and guanosine (G). In this regard, a is an integer of 1 or more and 4 or less, n is an integer selected from 0, 1, and 2, and b is an integer of 1 or more and 10 or less. In one embodiment, the U-rich tail sequence may be represented by (U a V) n U b . In this regard, a, n, and b may be integers, a may be 1 to 4, n may be 0 or greater, and b may be 1 to 10.
[0152] The vector may be constructed to express two or more artificial guide RNAs. In one embodiment, the vector may be constructed to express a first artificial guide RNA and a second artificial guide RNA. In one embodiment, the first artificial guide RNA sequence may include a first scaffold sequence, a first spacer sequence, and a first U-rich tail sequence, and the second artificial guide RNA sequence may include a second scaffold sequence, a second spacer sequence, and a second U-rich tail sequence. Expression target - additional components
[0153] In addition to the above-mentioned expression targets, the vector may be constructed to express additional components such as NLSs and tag proteins. In one embodiment, the additional components may be expressed independently of Cas12f1, modified Cas12f1, and / or artificial guide RNA. In one embodiment, the additional components may be expressed linked to Cas12f1, modified Cas12f1, and / or artificial guide RNA. In this regard, the additional components may be components that are generally expressed when expressing a CRISPR / Cas system, and publicly known techniques may be referred to. The additional components may be one or more of the components described in the section <Related Art - Design of CRISPR / Cas System Expression Vectors>. Vector components - Cas12f1 protein expression sequence
[0154] The vector sequence may include a nucleic acid sequence encoding a Cas12f1 protein. In this regard, the Cas12f1 protein has the same structure and characteristics as those described in the section <Artificial CRISPR / Cas12f1 Complex>. In one embodiment, the vector sequence may include a sequence encoding a wild-type Cas12f1 protein. In this regard, the wild-type Cas12f1 protein may be Cas14a1. In one embodiment, the vector sequence may include a human codon-optimized nucleic acid sequence encoding a Cas12f1 protein. In this regard, the human codon-optimized nucleic acid sequence encoding a Cas12f1 protein may be a human codon-optimized nucleic acid sequence encoding a Cas14a1 protein. In one embodiment, the vector sequence may be a sequence encoding a Cas12f1 protein or a Cas12f1 fusion protein. In one embodiment, the vector sequence may include a sequence encoding a Cas12f1 protein that has been modified so that only one strand of the double strand of the target nucleic acid can be cleaved and base editing or prime editing can be performed on the uncleaved strand. In one embodiment, the Cas12fl protein may comprise a sequence encoding a Cas12fl protein that has been modified so that it cannot cleave both strands of a target nucleic acid and can perform base editing, prime editing, or gene expression regulation functions on the target nucleic acid. Vector components - artificial guide RNA expression sequences
[0155] In one embodiment, the sequence of the vector may include a sequence encoding an artificial guide RNA. In one embodiment, the sequence of the artificial guide RNA may include a scaffold sequence, a spacer sequence, and a U-rich tail sequence. In one embodiment, the sequence of the artificial guide RNA sequence may include a tracrRNA sequence, a crRNA sequence, and a U-rich tail sequence. In one embodiment, the U-rich tail sequence is (U a N) n U bIn this regard, N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). In this regard, a is an integer between 1 and 4, n is an integer selected from 0, 1, and 2, and b is an integer between 1 and 10. In one embodiment, the U-rich tail sequence is (U a V) n U b In this regard, a, n, and b may be integers, where a may be 1 to 4, n may be 0 or greater, and b may be 1 to 10.
[0156] In one embodiment, the vector sequence may comprise a sequence encoding two or more artificial guide RNAs. In one embodiment, the vector sequence may comprise a sequence encoding a first artificial guide RNA and a sequence encoding a second artificial guide RNA. In this regard, the spacer sequence contained in the first artificial guide RNA sequence may be different from the spacer sequence contained in the second artificial guide RNA sequence. Furthermore, the U-rich tail sequence contained in the first artificial guide RNA sequence may be different from the U-rich tail sequence contained in the second artificial guide RNA sequence. In one embodiment, the first artificial guide RNA sequence may comprise a first scaffold sequence, a first spacer sequence, and a first U-rich tail sequence, and the second artificial guide RNA sequence may comprise a second scaffold sequence, a second spacer sequence, and a second U-rich tail sequence. Vector Components - Promoter Sequences
[0157] The vector sequence includes a promoter sequence operably linked to the sequence encoding each component. Specifically, the promoter sequence may be, but is not limited to, any one of the promoters disclosed in the promoter portion of the section <Related Art - Design of CRISPR / Cas System Expression Vector>.
[0158] In one embodiment, the vector sequence may comprise a sequence encoding the Cas12f1 protein and a promoter sequence. In this regard, the promoter sequence is operably linked to the sequence encoding the Cas12f1 protein. In one embodiment, the vector sequence may comprise a sequence encoding an artificial guide RNA and a promoter sequence. In this regard, the promoter sequence is operably linked to the sequence encoding the artificial guide RNA. In one embodiment, the vector sequence may comprise a sequence encoding the Cas12f1 protein, a sequence encoding the artificial guide RNA, and a promoter sequence. In this regard, the promoter sequence is operably linked to the sequence encoding the Cas12f1 protein and the sequence encoding the artificial guide RNA, and a transcription factor activated by the promoter sequence expresses the Cas12f1 protein and the artificial guide RNA. Vector construction - can contain two or more promoter sequences
[0159] In one embodiment, the vector sequence may include a first sequence encoding a Cas12f1 protein, a first promoter sequence, a second sequence encoding an artificial guide RNA, and a second promoter sequence. In this regard, the first promoter sequence is operably linked to the first sequence, the second promoter sequence is operably linked to the second sequence, and the transcription of the first sequence is induced by the first promoter sequence, and the transcription of the second sequence is induced by the second promoter sequence. In this regard, the first promoter and the second promoter may be the same type of promoter. In this regard, the first promoter and the second promoter may be different types of promoters.
[0160] In one embodiment, the vector sequence may include a first sequence encoding a Cas12f1 protein, a first promoter sequence, a second sequence encoding a first artificial guide RNA, a third sequence including a sequence encoding a second artificial guide RNA, and a third promoter sequence. In this regard, the first promoter sequence is operably linked to the first sequence, the second promoter sequence is operably linked to the second sequence, and the third promoter sequence is operably linked to the third sequence, with transcription of the first sequence being induced by the first promoter sequence, transcription of the second sequence being induced by the second promoter sequence, and transcription of the third sequence being induced by the third promoter sequence. In this regard, the second promoter and the third promoter may be the same type of promoter. Specifically, the second promoter and the third promoter may be, but are not limited to, a U6 promoter sequence. In this regard, the second promoter and the third promoter may be different types of promoters. Specifically, the second promoter sequence and the third promoter sequence may be, but are not limited to, a U6 promoter sequence and an H1 promoter sequence, respectively. Vector Components - Termination Signals
[0161] The vector may include a termination signal operably linked to the promoter sequence. In this regard, the termination signal may be one of the termination signals disclosed in the termination signal section of the section <Related Art - Design of CRISPR / Cas System Expression Vectors>, but is not limited thereto. The termination signal may vary depending on the type of promoter sequence. In one embodiment, when the vector sequence includes a U6 promoter sequence, a sequence of consecutive thymidines operably linked to the U6 promoter sequence may act as a termination signal. In one embodiment, the thymidine sequence may be a sequence of five or more consecutive thymidines. In one embodiment, when the vector sequence includes an H1 promoter sequence, a sequence of consecutive thymidines operably linked to the H1 promoter sequence may act as a termination signal. In one embodiment, the thymidine sequence may be a sequence of five or more consecutive thymidines. Vector components - relationship between the sequence encoding the U-rich tail and the termination signal
[0162] The sequences of the artificial guide RNAs provided in the present disclosure contain a U-rich tail sequence at their 3'-end. Therefore, the sequence encoding the artificial guide RNA will contain a T-rich sequence corresponding to the U-rich tail sequence at its 3'-end. As mentioned above, some promoter sequences recognize consecutive thymidine sequences, for example, a sequence in which five or more thymidines are linked consecutively, as a termination signal, and therefore, in some cases, a T-rich sequence may be recognized as a termination signal. That is, when a vector sequence according to the present disclosure includes a sequence encoding an artificial guide RNA, the sequence encoding the U-rich tail sequence included in the artificial guide RNA sequence can be used as a termination signal. In one embodiment, when the vector sequence is a U6 or H1 promoter sequence and includes a sequence encoding an artificial guide RNA operably linked to the U6 or H1 promoter sequence, the sequence portion encoding the U-rich tail sequence included in the artificial guide RNA sequence can be recognized as a termination signal. In this regard, the U-rich tail sequence includes a sequence in which five or more uridines are linked consecutively. Vector Component - Other Components
[0163] In addition to the structure, the vector sequence may contain components necessary for the purpose. In one embodiment, the vector sequence may include a regulatory / control component sequence and / or an additional component sequence. In one embodiment, the additional component may be added for the purpose of distinguishing transfected cells from non-transfected cells. In this regard, the regulatory / control component sequence and the additional component may be one of those disclosed in <Related Art - Design of CRISPR / Cas System Expression Vector>, but are not limited thereto. Types of vectors - viral vectors
[0164] The vector may be a viral vector. In one embodiment, the viral vector may be one or more selected from the group consisting of retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, and herpes simplex virus. In one embodiment, the viral vector may be an adeno-associated virus (AAV). Vector type - non-viral vector
[0165] The vector may be a non-viral vector. In one embodiment, the non-viral vector may be one or more selected from the group consisting of a plasmid, a phage, naked DNA, a DNA complex, and mRNA. In one embodiment, the plasmid may be selected from the group consisting of the pcDNA series, pSC101, pGV1106, pACYC177, ColE1, pKT230, pME290, pBR322, pUC8 / 9, pUC6, pBD9, pHC79, pIJ61, pLAFR1, pHV14, the pGEX series, the pET series, and pUC19. In one embodiment, the phage may be selected from the group consisting of λgt4λB, λ-Charon, λΔz1, and M13. In one embodiment, the vector may be a PCR amplicon. Vector form - circular or linear vector
[0166] A vector may be circular or linear. When a vector is linear, RNA transcription terminates at the 3' end of the vector even if a termination signal is not separately included in the linear vector sequence. In contrast, when a vector is circular, RNA transcription does not terminate unless a termination signal is separately included in the circular vector sequence. Therefore, when using a circular vector, termination signals corresponding to transcription factors associated with each promoter sequence must be included to express the intended target. In one embodiment, the vector may be a linear vector. In one embodiment, the vector may be a linear amplicon. In one embodiment, the vector may be a linear vector comprising a sequence selected from the group consisting of SEQ ID NOs: 1, 2, and 127-177. In one embodiment, the vector may be a circular vector. In one embodiment, the vector may be a circular vector comprising a sequence selected from the group consisting of SEQ ID NOs: 1, 2, and 127-177. Vector-sequence example
[0167] In one embodiment, the vector sequence may comprise one or more sequences selected from the group consisting of SEQ ID NOs: 1, 2, and 127 to 177. In one embodiment, the vector sequence may comprise one or more sequences selected from the group consisting of SEQ ID NOs: 1, 2, 316, 318, 320, and 322. Chemical modification of nucleic acids
[0168] One aspect of the present disclosure provides a component comprising or consisting of a nucleic acid such as an artificial crRNA or a nucleic acid encoding an artificial crRNA, an artificial guide RNA or a nucleic acid encoding an artificial guide RNA, and / or a vector for expressing components of the CRISPR / Cas12f1 system. In this regard, the "nucleic acid" used herein can be naturally occurring DNA or RNA, or a modified nucleic acid in which part or all of the nucleic acid is chemically modified. In one embodiment, the nucleic acid can be naturally occurring DNA and / or RNA. In one embodiment, the nucleic acid can be a nucleic acid in which one or more nucleotides are chemically modified. In this regard, the chemical modification includes all nucleic acid modifications known in the art. Specifically, examples of the chemical modification include, but are not limited to, all of the nucleic acid modifications described in International Publication No. WO 2019 / 089820. Gene editing method using artificial crRNA Overview
[0169] One aspect of the present disclosure provides a method for editing a target gene or a target nucleic acid in a target cell using an artificial crRNA. The target gene or the target nucleic acid contains a target sequence. The target nucleic acid can be single-stranded DNA, double-stranded DNA, and / or RNA. The gene editing method includes delivering to a target cell containing the target gene or the target nucleic acid a nucleic acid encoding each of an artificial guide RNA, a Cas12f1 protein, or an artificial guide RNA and a Cas12f1 protein. As a result, an artificial CRISPR / Cas12f1 complex is injected into the target cell or the formation of the artificial CRISPR / Cas12f1 complex is induced, and the target gene is edited by the artificial CRISPR / Cas12f1 complex. The artificial guide RNA has the same characteristics and structure as those described in the section of <Artificial guide RNA for CRISPR / Cas12f1 system>. The Cas12f1 protein has the same characteristics and composition as the Cas12f1 protein and / or the modified Cas12f1 protein described in the section of <Artificial CRISPR / Cas12f1 complex>.
[0170] In one embodiment, the gene editing method may include delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into a target cell. In this regard, the artificial guide RNA sequence comprises a scaffold sequence, a spacer sequence and a U-rich tail sequence. In this regard, the spacer sequence can be complementary to a target gene or target nucleic acid contained in a target cell. In one embodiment, the U-rich tail sequence is (U a V) n U b In this regard, a, n, and b are integers, a is 1 to 4, n is 0 or greater, and b is 1 to 10. In one embodiment, the U-rich tail sequence is (U a N) n U b In this regard, N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). In this regard, a is an integer between 1 and 4, n is an integer selected from 0, 1, and 2, and b is an integer between 1 and 10. target cell
[0171] In one embodiment, the target cell may be a prokaryotic cell. In one embodiment, the target cell may be a eukaryotic cell. Specifically, the eukaryotic cell may be, but is not limited to, a plant cell, an animal cell, and / or a human cell. Determination of target sequence
[0172] The target gene or nucleic acid and target sequence to be edited using the CRISPR Cas12f1 complex can be determined taking into consideration the purpose of gene editing, the environment of the target cell, the PAM sequence recognized by the Cas12f1 protein, and / or other variables. In this regard, the method is not particularly limited, and known techniques can be used as long as a target sequence having an appropriate length can be determined. Determining spacer sequences according to target sequences
[0173] Once the target sequence is determined, a corresponding spacer sequence is designed. The spacer sequence is designed as a sequence capable of binding complementarily to the target sequence. In one embodiment, the spacer sequence is designed as a sequence capable of binding complementarily to a target gene. In one embodiment, the spacer sequence is designed to be capable of binding complementarily to a target nucleic acid. In one embodiment, the spacer sequence is designed as a sequence complementary to a target sequence contained in the target strand sequence of the target nucleic acid. In one embodiment, the spacer sequence is designed as an RNA sequence corresponding to the DNA sequence of a protospacer contained in the non-target strand sequence of the target nucleic acid. Specifically, the spacer sequence has the same base sequence as the protospacer sequence, but each thymidine contained in the base sequence is substituted with a uridine. Complementarity between the target sequence and the spacer sequence
[0174] In one embodiment, the spacer sequence can be 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% complementary to the target sequence. In one embodiment, the spacer sequence can be a sequence that is complementary to the target sequence within the numerical range selected in the immediately preceding sentence. For example, the spacer sequence can be a sequence that is 60% to 90% complementary to the target sequence. In one embodiment, the spacer sequence may be a sequence that is 90% to 100% complementary to the target sequence. The number of mismatches between the target sequence and the spacer sequence
[0175] In one embodiment, the spacer sequence may be a complementary sequence having 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches with the target sequence. In one embodiment, the spacer sequence may have mismatches within a range of two of these values. For example, the spacer sequence may have 1 to 5 mismatches with the target sequence. In one embodiment, the spacer sequence may have 6 to 10 mismatches with the target sequence. Use of the CRISPR / Cas12f1 complex
[0176] The gene editing method provided in the present disclosure utilizes the fact that the artificial CRISPR / Cas12f1 complex has the activity of cleaving genes or nucleic acids in a target-specific manner. The artificial CRISPR / Cas12f1 complex has the same characteristics and composition as the artificial CRISPR / Cas12f1 complex described in the section <Artificial CRISPR / Cas12f1 Complex>. Delivery of each component of the CRISPR / Cas12f1 complex into cells
[0177] In the gene editing method provided herein, it is assumed that the artificial CRISPR / Cas12f1 complex is contacted with a target gene or nucleic acid in a target cell. Thus, to induce the artificial CRISPR / Cas12f1 complex to contact the target gene or target nucleic acid, the gene editing method includes a step of delivering each component of the artificial CRISPRCas12f1 complex into the target cell. In one embodiment, the gene editing method may include a step of delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA, and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into the target cell. In one embodiment, the gene editing method may include a step of delivering the artificial guide RNA and the Cas12f1 protein into the target cell. In one embodiment, the gene editing method may include a step of delivering the artificial guide RNA and the nucleic acid encoding the Cas12f1 protein into the target cell. In one embodiment, the gene editing method may include a step of delivering the artificial guide RNA and the nucleic acid encoding the Cas12f1 protein into the target cell. In one embodiment, the gene editing method may include a step of delivering the artificial guide RNA and the nucleic acid encoding the Cas12f1 protein into the target cell. In one embodiment, the gene editing method may include delivering a nucleic acid encoding an artificial guide RNA and a nucleic acid encoding a Cas12f1 protein into a target cell. The artificial guide RNA or the nucleic acid encoding the artificial guide RNA and the Cas12f1 protein or the nucleic acid encoding the Cas12f1 protein can be delivered into the target cell in various delivery forms using various delivery methods. Mode of delivery - RNP
[0178] As a delivery form, a ribonucleoprotein particle (RNP) in which an artificial guide RNA and a Cas12f1 protein are bound may be used. In one embodiment, the gene editing method may include injecting a CRISPR / Cas12f1 complex in which an artificial guide RNA and a Cas12f1 protein are bound into a target cell. Delivery mode - non-viral vector
[0179] As another delivery method, a non-viral vector containing a nucleic acid sequence encoding an artificial guide RNA and a nucleic acid sequence encoding a Cas12f1 protein can be used. In one embodiment, the gene editing method can include injecting a non-viral vector containing a nucleic acid sequence encoding an artificial guide RNA and a nucleic acid sequence encoding a Cas12f1 protein into a target cell. Specifically, the non-viral vector can be, but is not limited to, a plasmid, naked DNA, a DNA complex, or mRNA. In one embodiment, the gene editing method can include injecting a first non-viral vector containing a nucleic acid sequence encoding an artificial guide RNA and a second non-viral vector containing a nucleic acid sequence encoding a Cas12f1 protein into a target cell. Specifically, the first non-viral vector and the second non-viral vector can be, but are not limited to, one selected from the group consisting of a plasmid, naked DNA, a DNA complex, and mRNA, respectively. Delivery mode - viral vector
[0180] As yet another delivery method, a viral vector containing a nucleic acid sequence encoding an artificial guide RNA and a nucleic acid sequence encoding a Cas12f1 protein can be used. In one embodiment, the gene editing method can include injecting a nucleic acid sequence encoding an artificial guide RNA and a nucleic acid sequence encoding a Cas12f1 protein into a target cell. Specifically, the viral vector can be, but is not limited to, one selected from the group consisting of retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, and herpes simplex virus. In one embodiment, the viral vector can be an adeno-associated virus (AAV).
[0181] In one embodiment, the gene editing method may include injecting into a target cell a first viral vector comprising a nucleic acid sequence encoding an artificial guide RNA and a second viral vector comprising a nucleic acid sequence encoding a Cas12f1 protein. Specifically, the first viral vector and the second viral vector may each be one selected from the group consisting of, but not limited to, retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, and herpes simplex virus. Method of delivery - Common means of delivery
[0182] The delivery method is not particularly limited, as long as it can deliver the artificial guide RNA or the nucleic acid encoding the artificial guide RNA and the Cas12f1 protein or the nucleic acid encoding the Cas12f1 protein into cells in an appropriate delivery form. In one embodiment, the delivery method can be electroporation, gene gun method, ultrasonic perforation, magnetofection, and / or transient cell compression or squeezing. Delivery Method – Nanoparticles
[0183] The delivery method may be to deliver at least one component of the CRISPR / Cas12f1 system using nanoparticles. In this regard, the delivery method may be a known method that can be appropriately selected by a person skilled in the art. For example, the nanoparticle delivery method may be, but is not limited to, the method disclosed in (WO 2019 / 089820).
[0184] In one embodiment, the delivery method may be to deliver the Cas12fl protein or a nucleic acid encoding the Cas12fl protein and / or the artificial guide RNA or a nucleic acid encoding the artificial guide RNA using nanoparticles. In one embodiment, the delivery method may be to deliver the Cas12fl protein or a nucleic acid encoding the Cas12fl protein, the first artificial guide RNA or a nucleic acid encoding the first artificial guide RNA and / or the second artificial guide RNA or a nucleic acid encoding the second artificial guide RNA using nanoparticles. In this regard, delivery methods can be, but are not limited to, cationic liposomes, lithium acetate-DMSO, lipid-mediated transfection, calcium phosphate precipitation, lipofection, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, and / or nanoparticle-mediated nucleic acid delivery (see Panyam et al. Adv Drug Deliv Rev. 2012 Sep 13. pii:S0169-409X(12)00283-9. doi:10.1016 / j.addr.2012.09.023). In this regard, components of the CRISPR / Cas12f1 system can be in the form of RNPs, non-viral vectors, and / or viral vectors. For example, components of the CRISPR / Cas12f1 system can be in the form of mRNA encoding each component, but are not limited thereto. Delivery forms and methods - possible combinations
[0185] The gene editing method of the present invention comprises a step of delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into a cell, wherein the delivery form and / or delivery method of the components may be the same or different. In one embodiment, the gene editing method may comprise a step of delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA using a first delivery form, and a step of delivering a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein using a second delivery form. In this regard, the first delivery form and the second delivery form may each be any one of the delivery forms described above. In one embodiment, the gene editing method may comprise a step of delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA using a first delivery method, and a step of delivering a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein using a second delivery method. In this regard, the first delivery method and the second delivery method may each be any one of the delivery methods described above. delivery order
[0186] The gene editing method includes a step of delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into a cell, wherein the components may be delivered into the cell simultaneously or sequentially with a time interval therebetween.
[0187] In one embodiment, the gene editing method may include simultaneously delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein. In one embodiment, the gene editing method may include delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA into a cell, and then, after a time interval, delivering a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into the cell. In one embodiment, the gene editing method may include delivering a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into a cell, and then, after a time interval, delivering an artificial guide RNA into the cell. In one embodiment, the gene editing method may include delivering a nucleic acid encoding a Cas12f1 protein into a cell, and then, after a time interval, delivering an artificial guide RNA into the cell.
[0188] Delivery of multiple artificial guide RNAs or nucleic acids encoding multiple artificial guide RNAs
[0189] The gene editing method provided herein may include delivering into a target cell a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein and two or more artificial guide RNAs or a nucleic acid encoding two or more artificial guide RNAs. By this method, two or more CRISPR / Cas12f1 complexes targeting different sequences may be injected into the target cell or formed within the target cell. As a result, two or more target genes or target nucleic acids contained in the cell may be edited. In one embodiment, the gene editing method includes delivering into a target cell containing a target gene or target nucleic acid a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein, a first artificial guide RNA or a nucleic acid encoding the first artificial guide RNA, and a second artificial guide RNA or a nucleic acid encoding the second artificial guide RNA. In this regard, each component may be delivered into the cell using one or more of the delivery forms and methods described above. In this regard, the two or more components may be delivered into the cell simultaneously or sequentially with a time interval therebetween. The CRISPR / Cas12f1 complex and the target nucleic acid are brought into contact with each other.
[0190] In the gene editing method provided herein, editing of a target gene or target nucleic acid is carried out while an artificial CRISPR / Cas12f1 complex is contacted with the target gene or target nucleic acid in the target cell. Thus, the gene editing method may include contacting or inducing contact of the artificial CRISPR / Cas12f1 complex in the target cell. In one embodiment, the gene editing method may include contacting the artificial CRISPR / Cas12f1 complex with the target nucleic acid in the target cell. In one embodiment, the gene editing method may include inducing contact of the artificial CRISPR / Cas12f1 complex with the target nucleic acid in the target cell. In this regard, the induction is not particularly limited as long as it is a method of contacting the artificial CRISPR / Cas12f1 complex with the target nucleic acid in the cell. In one embodiment, the induction may include delivering an artificial guide RNA or a nucleic acid encoding the artificial guide and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into the cell. Gene editing results – indels
[0191] The gene editing method provided herein may result in the generation of indels in a target gene or target nucleic acid. In this regard, indels may occur inside and / or outside the target sequence portion and / or protospacer sequence portion. An indel refers to a mutation in a nucleotide sequence that deletes a portion of the nucleotide sequence of a nucleic acid, inserts any nucleotide, and / or combines insertion and deletion from the sequence before editing. Generally, the occurrence of an indel in a target gene or target nucleic acid sequence inactivates the corresponding gene or nucleic acid. In one embodiment, the gene editing method may result in the deletion and / or addition of one or more nucleotides in the target gene or target nucleic acid. Gene Editing Results – Base Editing
[0192] The gene editing method provided in the present disclosure can result in base editing in a target gene or target nucleic acid. That is, unlike indels, in which any base in a target gene or target nucleic acid is deleted or added, one or more specific bases in a nucleic acid are changed as intended. That is, base editing causes a predetermined point mutation at a specific position in a target gene or nucleic acid. In one embodiment, the gene editing method can result in one or more bases in a target gene or target nucleic acid being replaced with other bases. Gene editing results – insertion
[0193] The gene editing method provided herein can result in knock-in of a target gene or target nucleic acid. Knock-in refers to the insertion of an additional nucleic acid sequence into a target gene or target nucleic acid sequence. For knock-in to occur, in addition to the CRISPR / Cas12f1 complex, a donor containing an additional nucleic acid sequence is also required. When the CRISPR / Cas12f1 complex cleaves the target gene or target nucleic acid in a cell, repair of the cleaved target gene or target nucleic acid occurs. In this regard, the donor is involved in the repair process so that the additional nucleic acid sequence is inserted into the target gene or target nucleic acid. In one embodiment, the gene editing method can further include delivering a donor into a target cell. In this regard, the donor contains an additional target nucleic acid, and the donor induces the insertion of the additional target nucleic acid into the target gene or target nucleic acid. In this regard, when delivering the donor to a target cell, the above-mentioned delivery form and / or delivery method can be used. Gene editing results in deletion
[0194] The gene editing method provided in the present disclosure can result in the deletion of all or part of a target gene or target nucleic acid sequence. Deletion refers to the deletion of a portion of the base sequence in the target gene or target nucleic acid. In one embodiment, the gene editing method includes introducing a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein, a first artificial guide RNA or a nucleic acid encoding the first artificial guide RNA, and a second artificial guide RNA or a nucleic acid encoding the second artificial guide RNA into a cell containing the target gene or target nucleic acid. Thus, gene editing results in the deletion of a specific sequence portion in the target gene or target nucleic acid. Examples of gene editing methods
[0195] In one embodiment, the gene editing method may include delivering a CRISPR / Cas12f1 complex in the form of a ribonucleoprotein particle containing an artificial guide RNA and a Cas12f1 protein into a eukaryotic cell. In this regard, delivery may be performed by electroporation or lipofection.
[0196] In one embodiment, the gene editing method can include delivering a nucleic acid encoding an artificial guide RNA and a nucleic acid encoding a Cas12f1 protein into a eukaryotic cell. In this regard, delivery can be performed by electroporation or lipofection.
[0197] In one embodiment, the gene editing method may include delivering into a eukaryotic cell an adeno-associated virus (AAV) vector comprising a nucleic acid sequence encoding an artificial guide RNA and a nucleic acid sequence encoding a Cas12f1 protein.
[0198] In one embodiment, the gene editing method may include delivering into a eukaryotic cell an adeno-associated virus (AAV) comprising a nucleic acid sequence encoding a first artificial guide RNA, a nucleic acid sequence encoding a second artificial guide RNA, and a nucleic acid sequence encoding a Cas12f1 protein. Use of artificial crRNA Use of artificial crRNA in gene editing
[0199] In one embodiment, the artificial crRNA can be used for gene editing of a nucleic acid containing a target gene or target sequence.
[0200] In one embodiment, the artificial guide RNA can be used for gene editing of a nucleic acid containing a target gene or target sequence.
[0201] In one embodiment, the artificial CRISPR / Cas12f1 complex can be used for gene editing of a nucleic acid containing a target gene or target sequence.
[0202] In this regard, the artificial crRNA, artificial guide RNA and artificial CRISPR / Cas12f1 complex have the same characteristics and configurations as described in each section.
[0203] In one embodiment, the use of gene editing of target genes or nucleic acids can be used to perform the methods described in the section <Gene editing methods using artificial crRNA>.
[0204] In one embodiment, the use may include a step of delivering an artificial guide RNA or a nucleic acid encoding the artificial guide RNA and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into a target cell. In this regard, the artificial guide RNA sequence comprises a scaffold sequence, a spacer sequence, and a U-rich tail sequence. In this regard, the spacer sequence can be complementary to a target gene or target nucleic acid contained in a target cell. In one embodiment, the U-rich tail sequence is (U a V) n U b In this regard, a, n, and b are integers, a is 1 to 4, n is 0 or more, and b is 1 to 10. In one embodiment, the U-rich tail sequence is (U a N) n U bIn this regard, N is one of adenosine (A), uracil (U), cytidine (C), and guanosine (G). In this regard, a is an integer between 1 and 4, n is an integer selected from 0, 1, and 2, and b is an integer between 1 and 10.
[0205] In one embodiment, the use may include contacting the artificial CRISPR / Cas12f1 complex with a target nucleic acid in a target cell.
[0206] In this regard, the artificial CRISPR / Cas12f1 complex has the same characteristics and structure as those described in the section <Artificial CRISPR / Cas12f1 complex>.
[0207] In one embodiment, the use may include inducing the artificial CRISPR / Cas12f1 complex to contact with the target nucleic acid in the target cell. In this regard, the induction is not particularly limited as long as it is a method of inducing the artificial CRISPR / Cas12f1 complex to contact with the target nucleic acid in the cell. In one embodiment, the induction may include delivering an artificial guide RNA or a nucleic acid encoding the artificial guide and a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein into the cell. Use of artificial crRNA in preparing compositions for gene editing In one embodiment, artificial crRNA can be used in preparing compositions for gene editing.
[0208] In one embodiment, artificial guide RNAs may be used in the preparation of gene editing compositions.
[0209] In one embodiment, an artificial CRISPR / Cas12f1 complex may be used to prepare a gene editing composition.
[0210] In one embodiment, a gene-editing composition may include a Cas12f1 protein or a nucleic acid encoding a Cas12f1 protein, an artificial crRNA or a nucleic acid encoding an artificial crRNA, and a tracrRNA or a nucleic acid encoding a tracrRNA.
[0211] In one embodiment, the gene editing composition may comprise a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein, and an artificial guide RNA or a nucleic acid encoding the artificial guide RNA. In one embodiment, the gene editing composition may comprise a CRISPR / Cas12f1 complex.
[0212] In one embodiment, the gene editing composition may comprise a vector having a nucleic acid sequence encoding the Cas12f1 protein and a nucleic acid sequence encoding the artificial guide RNA.
[0213] [Mode for Carrying Out the Invention] The present invention will be described in more detail below through experimental examples and examples. It will be obvious to those skilled in the art that these examples are merely for illustrating the contents disclosed in the present disclosure and should not be construed as limiting the scope of the contents disclosed in the present disclosure. Experimental Example 1 Preparation of experimental materials and experimental method Experimental Example 1-1 Human codon-optimized Cas14a1 sequence
[0214] The Cas12f1 protein (hereafter referred to as Cas14a1; Doudna et al., Programmed DNA destruction by miniature CRISPR-Cas14 enzymes, Science 362, 839-842 (2018)), a member of the Cas14 family, was codon-optimized using a codon optimization program. NLS sequences were added to the 5' and 3' ends of the gene, resulting in a CDS. The synthesized gene was then PCR-amplified and cloned into a vector containing a promoter and polyA signal sequence compatible with eukaryotic systems using the Gibson assembly method according to the desired cloning sequence. The sequence of the resulting plasmid vector was then confirmed by Sanger sequencing. The amino acid sequence of the Cas14a1 protein and the human codon-optimized nucleic acid sequence encoding the Cas14a1 protein are shown in Table 1. [Table 1] Sequence information of NLS-Cas14a1-NLS protein [Table 1] Experimental Example 1-2 Preparation of Plasmid Vector
[0215] Oligonucleotides containing the human codon-optimized Cas14a1 DNA sequence used in this experiment were synthesized (Bionics). The Cas14a1 protein sequence oligonucleotides were cloned into a plasmid containing a chicken β-actin promoter, 5' and 3' NLSs, and T2A-linked eGFP. Template DNA for the canonical guide RNA used in this experiment was synthesized (Twist Bioscience) and cloned into the pTwist Amp plasmid for replication. Template DNA for the artificial guide RNA was constructed using enzymatic cloning techniques and cloned into the pTwist Amp plasmid for replication. Using this plasmid as a template, guide RNA or artificial guide RNA amplicons were prepared using a U6-complementary forward primer and a protospacer sequence-complementary reverse primer. The prepared amplicons were then cloned into the T-blunt plasmid (Biofact) for replication as needed. To prepare the (artificial) dual guide RNAs, oligonucleotides encoding the tracrRNA and the (artificial) crRNA were cloned into pSilencer 2.0 (ThermoFisher Scientific) using BamHI and HindIII restriction enzymes (New England Biolabs). Experimental Example 1-3 Preparation of viral vector
[0216] An adeno-associated virus (AAV) inverted terminal repeat (PTR) vector containing a guide RNA (or artificial guide RNA) sequence and a Cas12f1 protein sequence linked to enhanced green fluorescent protein (eGFP) via an NLS and a self-cleaving T2A sequence was prepared. Transcription of Cas12f1 and the guide RNA was driven by the chicken β-actin and U6 promoters, respectively. To prepare rAAV2 vectors, pAAV-ITR-sgRNA-Cas12f1, pAAVED29, and helper plasmids were transfected into HEK293T cells. Transfected HEK293T cells were cultured in DMEM medium containing 2% FBS. Recombinant pseudotyped AAV vector stocks were generated using triple transfection with Polyplus-transfection (PEIpro) and PEI coprecipitation at the same molar ratio relative to the plasmid. After 72 hours of culture, the cells were lysed and the lysate was purified by iodixanol (Sigma-Aldrich) step gradient ultracentrifugation. Experimental Example 1-4 Preparation of guide RNA amplicon
[0217] To prepare the guide RNA and artificial guide RNA amplicons, PCR amplification was performed on the template DNA plasmid for the canonical guide RNA and the artificial guide RNA using a U6-complementary forward primer and a protospacer sequence-complementary reverse primer with KAPA HiFi HotStart DNA polymerase (Roche) or Pfu DNA polymerase (Biofact). The PCR amplification products were purified using the Higene Gel & PCR Purification System (Biofact) to obtain the guide RNA and artificial guide RNA amplicons. Experimental Example 1-5 Preparation of Guide RNA Guide RNA (or artificial guide RNA) was prepared using one of the following methods: Guide RNA was prepared by chemically synthesizing a pre-designed guide RNA.
[0218] A PCR amplicon containing a predesigned guide RNA sequence and a T7 promoter sequence was prepared. Using the PCR amplicon as a template, in vitro transcription was performed using NEB T7 polymerase. The product obtained by in vitro transcription was treated with NEB DNase I, and then purified using the Monarch RNA Cleanup Kit (NEB). Guide RNA was then isolated.
[0219] A plasmid vector containing a predesigned guide RNA sequence and a T7 promoter sequence was prepared using the Tblunt plasmid cloning method of Experimental Example 1-2. The guide RNA sequence containing the T7 promoter sequence was double-cut at both ends, and the vector was purified. The resulting product was then subjected to in vitro transcription using NEB T7 polymerase. The resulting product was treated with NEB DNase I, and purified using the Monarch RNA Cleanup Kit (NEB). Guide RNA was then isolated. Experimental Example 1-6 Preparation of recombinant Cas12f1 protein
[0220] The human codon-optimized Cas14a1 DNA sequence used in this experiment was cloned into the pMAL-c2 plasmid vector. BL21(DE3) E. coli cells were transformed with this plasmid vector. Transformed E. coli colonies were grown in LB broth at 37°C until an optical density of 0.7 was reached. Transformed E. coli cells were cultured overnight at 18°C in the presence of 0.1 mM isopropylthio-β-D-galactoside. Cells were then harvested by centrifugation at 3,500g for 30 minutes. The cells were resuspended in 20mM Tris-HCl (pH 7.6), 500mM NaCl, 5mM β-mercaptoethanol, and 5% glycerol. The cells were lysed and disrupted by sonication. The sample containing the disrupted cells was centrifuged at 15,000g for 30 minutes, and the supernatant was filtered through a 0.4µm syringe filter (Millipore). The filtered supernatant was purified using an FPLC purification system (AKT Aurora, GE Healthcare). 2+ The protein was loaded onto an affinity column. The bound fraction was eluted with a gradient of 80–400 mM imidazole and 20 mM Tris-HCl (pH 7.5). The eluted protein was treated with TEV protease for 16 hours. The cleaved protein was purified on a heparin column with a linear gradient of 0.15–1.6 M NaCl. The purified recombinant Cas12f1 protein was dialyzed against 20 mM Tris pH 7.6, 150 mM NaCl, 5 mM β-mercaptoethanol, and 5% glycerol. The dialyzed protein was then passed through an MBP column and repurified on a monoS column (GE Healthcare) or Enrich S with a linear gradient of 0.5–1.2 M NaCl. The purified protein was collected and dialyzed against 20mM Tris pH 7.6, 150mM NaCl, 5mM β-mercaptoethanol, and 5% glycerol to prepare the Cas12f1 protein. The concentration of the produced protein was quantified by the Bradford assay using bovine serum albumin as a standard and electrophoretically measured on a Coomassie blue-stained SDS-PAGE gel. Experimental Example 1-7 Preparation of RNP
[0221] Ribonucleoprotein particles (RNPs) were prepared by incubating the recombinant Cas12f1 protein prepared according to Experimental Examples 1-6 with the guide RNA (or artificial guide RNA) prepared according to Experimental Examples 1-5 at 300 nM and 900 nM at room temperature for 10 minutes, respectively. Experimental Example 1-8 Cell culture
[0222] HEK-293T cells (ATCC CRL-11269) were cultured in DMEM medium supplemented with 10% heat-inactivated FBS (Corning) and 1% penicillin / streptomycin under conditions of 37°C and 5% CO2. Experimental Example 1-9: Transfection of Plasmid Vectors or PCR Amplicons into Cells
[0223] Cell transfection was performed by electroporation or lipofection. For electroporation, 4 × 10 cells were transfected using the Neon transfection system (Invitrogen). 5 HEK-293T cells were transfected with 2-5 μg of the plasmid vector encoding the Cas12f1 protein prepared in Experiment 1-2 and DNA encoding the guide RNA (and artificial guide RNA). Electroporation was performed at 1300 V, 10 mA, and three pulses. For lipofection, 3-15 μL of FuGene reagent (Promega) was mixed with 1-5 μg of the plasmid vector encoding the Cas12f1 protein and 0.3-2 μg of the PCR amplicon for 15 minutes. This mixture was added to 1 x 10 cells one day before transfection. 6The mixture was added to 1 ml of DMEM medium containing 1000 cells. The cells were cultured for 72–96 hours. After culture, the cells were harvested and genomic DNA was manually isolated using the PureHelix™ genomic DNA preparation kit (NanoHelix) or the Maxwell RSC Cultured cells DNA Kit (Promega). Experimental Example 1-10 Intracellular RNP transfection
[0224] The ribonucleoprotein particles (RNPs) prepared in Experimental Example 1-7 were transfected using the electroporation method of Experimental Example 1-9, or the plasmid encoding the Cas12f1 protein prepared in Experimental Example 1-2 was transfected using the lipofection method of Experimental Example 1-9 one day before transfection of the guide RNA (and artificial guide RNA), and one day later, the guide RNA (and artificial guide RNA) prepared in Experimental Example 1-5 was transfected using the electroporation method of Experimental Example 1-9. Experimental Example 1-11 Transfection of viral vectors into cells
[0225] Human HEK293T cells were infected with rAAV2-Cas12f1-sgRNA at different multiplicities of infection (MOI) of 1, 5, 10, 50, and 100, as determined by quantitative PCR. Transfected HEK293T cells were cultured in DMEM medium containing 2% FBS. At different time points, cells were harvested and genomic DNA was isolated. Experimental Example 1-12 Quantitative Real-Time PCR (Quantitative rtPCR)
[0226] Guide RNA (or artificial guide RNA) or genomic DNA was extracted from HEK293T cells using the RNeasy Miniprep Kit (Qiagen), Maxwell RSC miRNA Tissue Kit (Promega), or DNeasy Blood & Tissue Kit (Qiagen), respectively. To quantify guide RNA, RNA-specific primers were ligated and cDNA was synthesized using crRNA-specific primers. The cDNA was used as a template for quantitative real-time PCR. Real-time PCR was performed using the KAPA SYBR FAST qPCR Master Mix (2X) Kit (KapaBiosystems). Experimental Example 1-13 Western blotting to confirm Cas protein expression
[0227] A plasmid vector encoding the Cas protein prepared in Experimental Example 1-2 was prepared. In this regard, the plasmid further contains a sequence encoding an HA tag (5'-TACCCATACGATGTTCCAGATTACGCTTATCCCTACGACGTGCCTGATTATGCATACCCATATGATGTCCCCGACTATGCC-3' (SEQ ID NO: 315) such that the HA tag can be linked and expressed when the Cas protein is expressed in cells. HEK293T cells cultured in Experimental Example 1-8 were transfected with the plasmid vector using the method disclosed in Experimental Example 1-9. HEK293T cells were seeded in DMEM medium containing penicillin / streptomycin and cultured statically in an incubator at 37°C and 5% CO2 for 48 hours. The HEK293T cells were then diluted with PBS and then disrupted by sonication. The sample containing the disrupted cells was centrifuged at 15,000g and 4°C for 10 minutes, and the supernatant was quantified by Bradford assay. Whether the quantified HA antibody was expressed in vivo was verified by Western blotting. Experimental Example 1-14 Indel Efficiency Analysis
[0228] Genomic DNA isolated from HEK-293T cells was subjected to PCR using target-specific primers in the presence of KAPA HiFi HotStart DNA polymerase (Roche). The amplification method followed the manufacturer's instructions. The resulting PCR amplicon, containing the Illumina TruSeq HT Dual Index, was subjected to 150-bp paired-end sequencing using an Illumina iSeq 100. Indel frequencies were calculated using MAUND.MAUND. MAUND is available at https: / / github.com / ibscge / maund. Experimental Example 1-15 Analysis of T7 endonuclease I (T7E1)
[0229] PCR products were obtained using BioFACT™ Lamp Pfu DNA polymerase. PCR products (100–300 μg) were reacted with 10 units of T7E1 enzyme (New England Biolabs) in a 25 μg reaction mixture at 37°C for 30 minutes. 20 μl of the reaction mixture was directly loaded onto a 10% arylamide gel, and the cleaved PCR products were run in a TBE buffer system. The gel images were stained with ethidium bromide solution and then digitized using a Printgraph 2M gel imaging system (Atto). Gene editing efficiency was assessed by analyzing the digitized results. Experimental Example 1-16 Statistical Analysis
[0230] Statistical significance was verified using a two-tailed Student's t-test using Sigma Plot software (ver. 14.0). A p-value of less than 0.05 was considered statistically significant, and the p-value is indicated in each figure. In box and dot plots, data points represent the full range of values, with each box spanning the interquartile range (25th to 75th percentile). The median and mean are indicated by black and red parallel lines, respectively. Error bars for all dot and bar plot data were plotted using Sigmaplot (v. 41.0) and represent the standard deviation of each data point. In this regard, sample size was not predetermined based on statistical methods.
[0231] Experimental Example 2: Comparison of indel efficiency of artificial CRISPR / Cas14a1 systems with U-rich tails 1
[0232] We conducted experiments to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system with a U-rich tail sequence using various gene sequences as target nucleic acids. The protospacer sequences and PAM sequences of the target nucleic acids are listed in Table 2. [Table 2] Target nucleic acid of Experimental Example 2 [Table 2]
[0233] 1) According to Experimental Example 1-2, a plasmid vector containing a human codon-optimized sequence encoding the Cas14a1 protein was prepared. 2) Amplicons capable of expressing wild-type or (artificial) single-guide RNAs with U-rich tail sequences were prepared according to Experimental Examples 1-4. The DNA sequences encoding the single-guide RNAs and the primer sequences used to prepare the amplicons are shown in Tables 3-5. [Table 3] DNA sequence encoding wild-type (WT) sgRNA used in Experimental Example 2 [Table 3]
[0234] In the above table, for the DNA sequence encoding the sgRNA for each target nucleic acid, the sgRNA consensus sequence, protospacer sequence, and U-rich tail sequence are linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 254. The amplicon sequence capable of expressing an (artificial) single guide RNA is an amplicon sequence in which a U6 promoter is operably linked to the DNA sequence encoding the sgRNA. [Table 4] DNA sequence encoding the artificial sgRNA used in Experimental Example 2 [Table 4]
[0235] In the above table, for the DNA sequence encoding the sgRNA for each target nucleic acid, the sgRNA consensus sequence, protospacer sequence, and U-rich tail sequence are linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 254. The amplicon sequence capable of expressing an (artificial) single guide RNA is an amplicon sequence in which a U6 promoter is operably linked to the DNA sequence encoding the sgRNA. [Table 5] Primer sequences used in Experimental Example 2 [Table 5]
[0236] 3) HEK-293T cells cultured according to Experimental Example 1-8 were transfected with a plasmid vector containing the Cas14a1 sequence and an amplicon capable of expressing a single guide RNA using the method disclosed in Experimental Example 1-9. 4) The indel efficiency for the target nucleic acid of each example was analyzed by the method disclosed in Experimental Examples 1-12 to 1-15. The results of the indel efficiency analysis are shown in Table 6 below. [Table 6] Indel efficiency analysis results for Experimental Example 2 [Table 6]
[0237] In this regard, a control was a transfectant containing only a plasmid vector containing the Cas14a1 protein sequence without sgRNA.
[0238] Experimental Example 3: Comparison of indel efficiency of artificial CRISPR / Cas14a1 systems with U-rich tails 2
[0239] Experiments were conducted to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system with U-rich tail sequences using DYtarget2, DYtarget10, DYtarget13, and Intergenic-22 as target nucleic acids. The protospacer sequences and PAM sequences of the target nucleic acids are listed in Table 7. [Table 7] Target nucleic acid sequence used in Experimental Example 3 [Table 7]
[0240] 1) A plasmid vector containing a human codon-optimized Cas14a1 sequence was prepared according to Experimental Example 1-2. 2) Amplicons capable of expressing wild-type or (artificial) single-guide RNAs with U-rich tail sequences were prepared according to Experimental Examples 1 and 2. The DNA sequences encoding the single-guide RNAs and the primer sequences used to prepare the amplicons are shown in Tables 8 and 9. [Table 8] DNA sequence encoding sgRNA in Experimental Example 3 [Table 8]
[0241] In the above table, for the DNA sequence encoding the sgRNA for each target nucleic acid, the sgRNA consensus sequence, protospacer sequence, and U-rich tail sequence are linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 254. The amplicon sequence capable of expressing an (artificial) single guide RNA is an amplicon sequence in which a U6 promoter is operably linked to the DNA sequence encoding the sgRNA. [Table 9] Primer sequences used in Experimental Example 3 [Table 9]
[0242] 3) HEK-293T cells cultured according to Experimental Example 1-8 were transfected with a plasmid vector containing the Cas14a1 sequence and a single guide RNA amplicon using the method disclosed in Experimental Example 1-9. 4) The indel efficiency of the target nucleic acid in each example was analyzed by the method disclosed in Experimental Examples 1-12 and 1-13. The results of the indel efficiency analysis are shown in FIG.
[0243] As disclosed in Figure 2, the CRISPR / Cas14a1 system containing an artificial guide RNA with a U-rich tail sequence linked to its 3' end exhibited higher indel efficiency than the wild type.
[0244] Experimental Example 4 Comparison of indel efficiency of artificial CRISPR / Cas14a1 systems with U-rich tails 3
[0245] Experiments were conducted to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system with U-rich tail sequences using DYtarget2, DYtarget10, and Intergenic-22 as target nucleic acids. The protospacer sequences and PAM sequences of the target nucleic acids are shown in Table 10. [Table 10] Target nucleic acid sequence used in Experimental Example 4 [Table 10]
[0246] 1) A plasmid vector containing a human codon-optimized Cas14a1 sequence was prepared according to Experimental Example 1-2. 2) Plasmid vectors and amplicons capable of expressing tracrRNA and wild-type or U-rich tailed (artificial) crRNA were prepared according to Experimental Examples 1-2 and 1-4. The DNA sequences encoding tracrRNA and crRNA, as well as the primer sequences used to prepare the plasmid vectors and amplicons, are shown in Tables 11 to 13. [Table 11] DNA sequence encoding tracrRNA in Experimental Example 4 [Table 11]
[0247] In this regard, an amplicon sequence or plasmid vector sequence capable of expressing tracrRNA is an amplicon sequence or plasmid vector sequence in which a U6 promoter is operably linked to a DNA sequence encoding tracrRNA. [Table 12] DNA sequence encoding crRNA in Experimental Example 4 [Table 12]
[0248] In the above table, the DNA sequence encoding the crRNA of each target nucleic acid is composed of a crRNA consensus sequence, a protospacer sequence, and a U-rich tail sequence linked in the 5' to 3' direction, and the crRNA repeat sequence in the crRNA consensus sequence is the sequence of SEQ ID NO: 58. The amplicon sequence or plasmid sequence capable of expressing an (artificial) crRNA is an amplicon sequence or plasmid sequence in which a U6 promoter is operably linked to a DNA sequence encoding the crRNA. [Table 13] Primer sequences used in Experimental Example 4 [Table 13]
[0249] 3) HEK-293T cells cultured according to Experimental Examples 1-8 were transfected with a plasmid vector containing the Cas14a1 sequence, a plasmid vector or amplicon of tracrRNA, and a plasmid vector or amplicon of crRNA by the method disclosed in Experimental Examples 1-9. In this regard, cells without crRNA injection were used as a control (data for No crRNA in Figure 3). 4) The indel efficiency for the target nucleic acid of each example was analyzed by the method disclosed in Experimental Examples 1-12 to 1-14. The results of the indel efficiency analysis are shown in FIG.
[0250] As shown in Figure 3, the CRISPR / Cas14a1 system containing an artificial guide RNA with a U-rich tail sequence attached to its 3' end showed higher indel efficiency than the wild type.
[0251] Experimental Example 5 Comparison of indel efficiency of artificial CRISPR / Cas14a1 systems with U-rich tails 4
[0252] We conducted experiments to evaluate the indel efficiency of an artificial CRISPR / Cas14a1 system with a U-rich tail sequence using various target sequences as target nucleic acids.
[0253] The experiments were performed with the target nucleic acid disclosed in Experimental Example 3 or 4, and instead of transfecting HEK-293T cells with a plasmid vector containing the Cas14a1 sequence and a single guide RNA amplicon, the experiments followed the following method. 1) Prepare single guide RNAs having the sgRNA sequences disclosed in Experimental Examples 3 and 4 according to Experimental Examples 1 to 5. 2) Prepare Cas14a1 protein according to Experimental Examples 1-6. 3) Prepare RNP from the resulting products of 1) and 2) according to Experimental Examples 1-7. 4) HEK-293T cells prepared according to Experimental Example 1-8 are transfected with RNP according to Experimental Example 1-10. 5) The following indel efficiency analysis method is as described in Experimental Example 2 or 3.
[0254] Experimental Example 6 Comparison of indel efficiency of artificial CRISPR / Cas14a1 systems with U-rich tails 5
[0255] We will conduct experiments to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system with a U-rich tail sequence using various target genes as target nucleic acids.
[0256] For the target nucleic acid disclosed in Experimental Examples 3 or 4, the experiment is performed according to the following method, instead of transfecting HEK-293T cells with a plasmid vector containing the Cas14a1 sequence and a single guide RNA amplicon. 1) A viral vector having the sgRNA sequence and a sequence encoding the Cas14a1 protein disclosed in Experimental Examples 3 and 4 is prepared according to Experimental Examples 1 to 3. 4) HEK-293T cells prepared according to Experimental Example 1-8 are transfected with a viral vector according to Experimental Example 1-11. 5) The following indel efficiency analysis method is as described in Experimental Example 2 or 3. Experimental Example 7 Comparison of indel efficiency depending on the U-rich tail structure 1
[0257] Using DYtarget2 and DYtarget10 as target nucleic acids, experiments were conducted to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system depending on the structure of the U-rich tail sequence. The protospacer sequences of the target nucleic acids and each PAM sequence are disclosed in Table 14. [Table 14] Target nucleic acid sequence used in Experimental Example 7 [Table 14]
[0258] 1) A plasmid vector containing a human codon-optimized Cas14a1 sequence was prepared according to Experimental Example 1-2. 2) Amplicons capable of expressing wild-type or (artificial) single-guide RNAs with U-rich tail sequences were prepared according to Experimental Examples 1-4. The DNA sequences encoding the single-guide RNAs and the primer sequences used to prepare the amplicons are shown in Tables 15 and 16. [Table 15] DNA sequence encoding sgRNA in Experimental Example 7 [Table 15]
[0259] In the above table, for the DNA sequence encoding the sgRNA for each target nucleic acid, the sgRNA consensus sequence, protospacer sequence, and U-rich tail sequence are linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 254. The amplicon sequence capable of expressing an (artificial) single guide RNA is an amplicon sequence in which a U6 promoter is operably linked to the DNA sequence encoding the sgRNA. [Table 16] Primer sequences used in Experimental Example 7 [Table 16]
[0260] 3) HEK-293T cells cultured according to Experimental Example 1-8 were transfected with a plasmid vector containing the Cas14a1 sequence and a single guide RNA amplicon using the method disclosed in Experimental Example 1-9. 4) The indel efficiency for the target nucleic acid of each example was analyzed by the method disclosed in Experimental Examples 1-12 to 1-14. The results of the indel efficiency analysis are shown in Figure 4. In this regard, HEK-293T cells that had not been subjected to any treatment were used as a control (see WT data in Figure 4). Experimental Example 8 Comparison of indel efficiency depending on U-rich tail structure 2
[0261] Using DYTarget2 as the target nucleic acid, an experiment was conducted to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system depending on the structure of the U-rich tail sequence. The protospacer sequence and PAM sequence of the target nucleic acid are shown as "DYTarget2" in Table 14 of Experimental Example 7. 1) A plasmid vector containing a human codon-optimized Cas14a1 sequence was prepared according to Experimental Example 1-2. 2) Amplicons capable of expressing wild-type or (artificial) single-guide RNAs with U-rich tail sequences were prepared according to Experimental Examples 1-4. The DNA sequences encoding the single-guide RNAs and the primer sequences used to prepare the amplicons are shown in Tables 17 and 18. [Table 17] DNA sequence encoding sgRNA of Experimental Example 8 [Table 17]
[0262] In the above table, the amplicon sequences encoding the sgRNA for each target nucleic acid are composed of the sgRNA consensus sequence, protospacer sequence, and U-rich tail sequence linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 254. The amplicon sequences capable of expressing an (artificial) single guide RNA are amplicon sequences in which a U6 promoter is operably linked to a DNA sequence encoding the sgRNA. [Table 18] Primer sequences used in Experimental Example 8 [Table 18]
[0263] 1) HEK-293T cells cultured according to Experimental Example 1-8 were transfected with a plasmid vector containing the Cas14a1 sequence and a single guide RNA amplicon by the method disclosed in Experimental Example 1-9. 2) The indel efficiency for the target nucleic acid of each example was analyzed by the method disclosed in Experimental Examples 1-12 to 1-14. The results of the indel efficiency analysis are shown in FIG. Experimental Example 9 Comparison of indel efficiency depending on the U-rich tail structure 3
[0264] Using CSMD1 as the target nucleic acid, experiments were conducted to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system depending on the structure of the U-rich tail sequence. The protospacer sequence and PAM sequence of the target nucleic acid are shown in Table 19. [Table 19] Target nucleic acid sequence used in Experimental Example 9 [Table 19]
[0265] 1) Recombinant Cas12f1 protein was prepared according to Experimental Examples 1-6. 2) Wild-type or (artificial) single-guide RNAs with U-rich tail sequences were synthesized according to Experimental Examples 1-5. The sequences of the single-guide RNAs are shown in Table 20 below. [Table 20] sgRNA sequence of Experimental Example 9 [Table 20]
[0266] In the table above, the sgRNA sequence for each target nucleic acid is composed of an sgRNA consensus sequence, a spacer sequence, and a U-rich tail sequence linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 253. 3) RNPs were prepared using recombinant Cas14a1 and single guide RNA according to Experimental Examples 1-7, and the RNPs were transfected into HEK293T cells cultured according to Experimental Examples 1-8 using the method disclosed in Experimental Examples 1-7. 2) The indel efficiency for the target nucleic acid of each example was analyzed by the method disclosed in Experimental Examples 1-12 and 1-13. The results of the indel efficiency analysis are shown in FIG.
[0267] As a result of the experiment, it was confirmed that artificial guide RNAs with a U-rich tail sequence attached to the 3' end exhibited significantly higher indel efficiency than wild-type guide RNAs without a U-rich tail sequence.
[0268] In particular, it was confirmed that artificial guide RNAs having a sequence linked with 10 or more uridines exhibit superior indel efficiency compared to wild-type guide RNAs. Experimental Example 10: Comparison of indel efficiency depending on the U-rich tail structure 3
[0269] Using DYtarget2 as the target nucleic acid, we conducted an experiment to evaluate the indel efficiency of the artificial CRISPR / Cas14a1 system depending on the structure of a modified consecutive uridine sequence in the U-rich tail sequence. The protospacer sequence and PAM sequence of the target nucleic acid are shown as "DYTarget2" in Table 14 of Experimental Example 7. 1) A plasmid vector containing a human codon-optimized Cas14a1 sequence was prepared according to Experimental Example 1-2. 2) Amplicons capable of expressing wild-type or (artificial) single-guide RNAs with U-rich tail sequences were prepared according to Experimental Examples 1-4. The DNA sequences encoding the single-guide RNAs and the primer sequences used to prepare the amplicons are shown in Tables 21 and 22. [Table 21] DNA sequence encoding sgRNA of Experimental Example 10 [Table 21]
[0270] In the above table, the DNA sequence encoding the sgRNA for each target nucleic acid comprises a sgRNA consensus sequence, a protospacer sequence, and a U-rich tail sequence linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 254. The amplicon sequence capable of expressing an (artificial) single guide RNA is an amplicon sequence in which a U6 promoter is operably linked to a DNA sequence encoding the sgRNA. [Table 22] Primer sequences used in Experimental Example 10 [Table 22]
[0271] 3) HEK-293T cells cultured according to Experimental Example 1-8 were transfected with a plasmid vector containing a nucleic acid sequence encoding the Cas14a1 sequence and a single guide RNA amplicon by the method disclosed in Experimental Example 1-9. 4) The indel efficiency for the target nucleic acid of each example was analyzed by the method disclosed in Experimental Examples 1-12 to 1-14. The results of the indel efficiency analysis are shown in FIG.
[0272] As a result of the experiment, the artificial guide RNA (U a V) n U b In this regard, modified guide RNAs with consecutive uridine sequences (where V can be represented by cytidine (C) and one of guanosine (G) or adenosine (A)) have been shown to have significantly higher indel efficiency in eukaryotic cells than wild-type guide RNAs without U-rich tail sequences. Experimental Example 11 Comparison of indel efficiency depending on the length of the spacer sequence
[0273] To compare the indel efficiency depending on the length of the spacer sequence, experiments were performed using DYtarget2 and DYtarget10 as target nucleic acids, with the U-rich tail sequence fixed as U4AU6, and then varying the length of the spacer sequence. The protospacer sequences of the target nucleic acids and each PAM sequence are shown in Table 23. [Table 23] Target nucleic acid sequence used in Experimental Example 11 [Table 23]
[0274] 1) A plasmid vector containing a nucleic acid sequence encoding a human codon-optimized Cas14a1 sequence was prepared according to Experimental Example 1-2. 2) Amplicons capable of expressing artificial single guide RNAs with different spacer sequence lengths were prepared according to Experimental Examples 1-4. The DNA sequences encoding the single guide RNAs and the primer sequences used to prepare the amplicons are shown in Tables 24 and 25. [Table 24] DNA sequence encoding sgRNA used in Experimental Example 11 [Table 24]
[0275] In the above table, the DNA sequence encoding the sgRNA for each target nucleic acid comprises a sgRNA consensus sequence, a protospacer sequence, and a U-rich tail sequence linked in the 5' to 3' direction, and the scaffold sequence in the sgRNA consensus sequence is the sequence portion of SEQ ID NO: 254. The amplicon sequence capable of expressing an (artificial) single guide RNA is an amplicon sequence in which a U6 promoter is operably linked to a DNA sequence encoding the sgRNA. [Table 25] Primer sequences used in Experimental Example 11 [Table 25]
[0276] 3) HEK-293T cells cultured according to Experimental Example 1-8 were transfected with a plasmid vector containing the Cas14a1 sequence and a single guide RNA amplicon using the method disclosed in Experimental Example 1-9. 4) The indel efficiency for the target nucleic acid of each example was analyzed by the method disclosed in Experimental Examples 1-12 to 1-14. The results of the indel efficiency analysis are shown in FIG.
[0277] Experimental Example 12: Confirmation of intracellular gene cleavage activity of Type V CRISPR / Cas system with U-rich tail
[0278] To confirm whether gene cleavage activity is enhanced when a U-rich tail sequence is introduced into a Type V CRISPR / Cas system other than the artificial CRISPR / Cas14a system used in the experimental example, the following experiment was performed on the CRISPR / SpCas12f1 system and CRISPR / AsCas12f1 system (Karvelis et al., Nucleic Acids Research, Vol. 48, No. 95017 (2020)), which are classified as CRISPR / Cas12f1 systems.
[0279] Experimental Example 12-1 Confirmation of intracellular expression of SpCas12f1 and AsCas12f1 proteins
[0280] To confirm whether the SpCas12f1 and AsCas12f1 proteins are appropriately expressed in eukaryotic cells after vectors encoding them are introduced into the cells, the expression of the SpCas12f1 and AsCas12f1 proteins was confirmed by the methods disclosed in Experimental Examples 1-13. In this regard, the amino acid sequences of the SpCas12f1 and AsCas12f1 proteins and human codon-optimized nucleic acid sequences encoding the amino acid sequences are shown in Table 26. [Table 26] Sequence information of SpCas12f1 and AsCas12f1 proteins. [Table 26]
[0281] The experimental results confirmed that the SpCas12f1 and AsCas12f1 proteins were properly expressed in HEK293T cells (Figure 9).
[0282] Experimental Example 12-2. Confirmation of intracellular gene cleavage activity of artificial CRISPR / SpCas12f1 and CRISPR / AsCas12f1 systems with U-rich tails
[0283] To confirm the gene cleavage activity of the artificial CRISPR / SpCas12f1 and CRISPR / AsCas12f1 systems with U-rich tail sequences, experiments were conducted using various gene sequences as target nucleic acids. The protospacer sequences and PAM sequences of the target nucleic acids are shown in Table 27. [Table 27] Target nucleic acid used in Experimental Example 12 [Table 27]
[0284] 1) According to Experimental Example 1-2, a plasmid vector containing human codon-optimized sequences encoding SpCas12f1 and AsCas12f1 proteins was prepared. 2) Plasmid vectors expressing tracrRNA and wild-type crRNA or (artificial) crRNA with a U-rich tail sequence were prepared according to Experimental Example 1-2. The DNA sequences encoding tracrRNA and crRNA and the primer sequences used to prepare the plasmid vectors are shown in Tables 28-30. [Table 28] [Table 28]
[0285] The CRISPR / SpCas12f1 system and the CRISPR / AsCas12f1 system have the same tracrRNA sequence. The plasmid vector sequence for expressing tracrRNA is a plasmid vector sequence in which a U6 promoter is operably linked to a DNA sequence encoding tracrRNA. [Table 29] DNA sequence encoding crRNA of Experimental Example 12 [Table 29]
[0286] In the table above, the DNA sequence encoding the crRNA for each target nucleic acid is composed of a crRNA consensus sequence, a protospacer sequence, and a U-rich tail sequence linked in the 5' to 3' direction. The plasmid vector sequence expressing the crRNA is a plasmid vector sequence in which a U6 promoter is operably linked to the DNA sequence encoding the crRNA. [Table 30] Primer sequences used in Experimental Example 12 [Table 30]
[0287] 3) HEK-293T cells cultured according to Experimental Examples 1-8 were transfected with plasmid vectors containing nucleic acid sequences encoding the SpCas12f1 protein, the AsCas12f1 protein, the tracrRNA plasmid vector, and the crRNA plasmid vector, using the method disclosed in Experimental Examples 1-9. 4) The indel efficiency of the target nucleic acid in each example was analyzed by the method disclosed in Experimental Examples 1-12 to 1-14. The results of the indel efficiency analysis are shown in Table 31. [Table 31] Indel efficiency analysis results for Experimental Example 12 [Table 31]
[0288] As a result of the experiment, no significant indel efficiency was observed in either the artificial CRISPR / SpCas12f1 system or the artificial CRISPR / AsCas12f1 system with the U-rich tail sequence introduced. These results confirm that the SpCas12f1 and AsCas12f1 proteins can be expressed in eukaryotic cells (Figure 9), but no significant indel efficiency was observed. Therefore, it was confirmed that the introduction of the U-rich tail into the CRISPR / SpCas12f1 and CRISPR / AsCas12f1 systems does not affect their gene cleavage activity. [Industrial Applicability]
[0289] The present disclosure provides a CRISPR / Cas12f1 system that can be used in gene editing techniques, and in particular, provides an artificial CRISPR / Cas12f1 system that has improved gene editing efficiency due to the introduction of a U-rich tail sequence.
[0290] When used for gene editing, the artificial CRISPR / Cas12f1 system having a U-rich tail sequence according to the present disclosure exhibits higher gene editing efficiency than the wild-type CRISPR / Cas12f1 system.
Claims
1. An artificial CRISPR RNA (crRNA) for a CRISPR / Cas12f1 system capable of editing a nucleic acid containing a target sequence, the artificial crRNA comprising: a spacer sequence complementary to the target sequence; a U-rich tail sequence linked to the 3' end of the spacer sequence; The repeat sequence of crRNA and Including, The U-rich tail sequence is (U a N) n U b wherein N is selected from adenosine (A), uridine (U), cytidine (C), and guanosine (G); a is an integer of 1 or more and 4 or less, n is an integer selected from 0, 1 and 2, and b is an integer of 1 or more and 10 or less, the U-rich tail sequence is not U, UU, or UUU, The repeat sequence of the crRNA, the spacer sequence, and the U-rich tail sequence are sequentially linked in the 5' to 3' direction of the artificial crRNA, and the repeat sequence is a sequence determined by the type of Cas12f1 protein of the CRISPR / Cas12f1 system and is a sequence linked to the 5' end of the spacer sequence.
2. DNA having a sequence encoding the artificial crRNA of claim 1.
3. the U-rich tail sequence is UUUUUU, UUUUAUUUU, UUUUAUUUUUU or UUUUGUU The artificial crRNA of claim 1, wherein the artificial crRNA is UUUU.
4. The artificial crRNA of claim 1, wherein the repeat sequence of the crRNA has the sequence of SEQ ID NO:
58.
5. An artificial guide RNA for a CRISPR / Cas12f1 system capable of editing a nucleic acid containing a target sequence, comprising: a tracrRNA sequence; and a crRNA sequence comprising a repeat sequence of CRISPR RNA and a spacer sequence complementary to the target sequence; a U-rich tail sequence linked to the 3' end of the crRNA sequence; and having a sequence comprising The U-rich tail sequence is (U a N) n U b wherein N is selected from adenosine (A), uridine (U), cytidine (C) and guanosine (G), a is an integer of 1 or more and 4 or less, n is an integer selected from 0, 1 and 2, and b is an integer of 1 or more and 10 or less, the U-rich tail sequence is not U, UU, or UUU, and the repeat sequence is a sequence determined by the type of Cas12f1 protein of the CRISPR / Cas12f1 system and is a sequence linked to the 5' end of the spacer sequence.
6. 6. The artificial guide RNA of claim 5, wherein the artificial guide RNA further comprises a linker sequence, and the tracrRNA sequence and the crRNA sequence are linked via the linker sequence.
7. The artificial guide RNA according to claim 6, wherein the linker sequence is 5'-gaaa-3'.
8. the U-rich tail sequence is UUUUUU, UUUUAUUUU, UUUUAUUUUUU or UUUUGUU The artificial guide RNA according to claim 6 or 7, which is UUUU.
9. A DNA having a sequence encoding the artificial guide RNA according to claim 5 or 6.
10. An artificial CRISPR / Cas12f1 complex capable of editing a nucleic acid containing a target sequence, the artificial CRISPR / Cas12f1 complex comprising: A Cas12f1 protein belonging to the Cas14 family; An artificial guide RNA, A scaffold sequence that interacts with the Cas12f1 protein; a spacer sequence complementary to the target sequence; U-rich tail sequence and and an artificial guide RNA comprising: Including, The U-rich tail sequence is (U a N) n U b is expressed as N is selected from the group consisting of adenosine (A), uridine (U), cytidine (C) and guanosine (G); An artificial CRISPR / Cas12f1 complex, wherein a is an integer between 1 and 4, n is an integer selected from 0, 1, and 2, and b is an integer between 1 and 10, inclusive, and the U-rich tail sequence is not U, UU, or UUU.
11. A vector for expressing an artificial CRISPR / Cas12f1 complex capable of editing a nucleic acid containing a target sequence, the vector comprising: a first sequence comprising a sequence encoding a Cas12f1 protein; a first promoter sequence operably linked to the first sequence; A second sequence comprising a sequence encoding an artificial guide RNA, the artificial guide RNA having a scaffold sequence that interacts with the Cas12f1 protein, a spacer sequence complementary to the target sequence, and a U-rich tail sequence linked to the 3' end of the spacer sequence, the U-rich tail sequence being (U a N) n U b wherein N is selected from the group consisting of adenosine (A), uridine (U), cytidine (C) and guanosine (G); a is an integer between 1 and 4, inclusive; n is an integer selected from 0, 1 and 2, and b is an integer between 1 and 10, inclusive; and the U-rich tail sequence is not U, UU, or UUU; a second promoter sequence operably linked to the second sequence; and Including, The vector is constructed to express the Cas12f1 protein and the artificial guide RNA, such that the Cas12f1 protein and the artificial guide RNA form the artificial CRISPR / Cas12f1 complex in the cell, and a nucleic acid containing the target sequence is edited by the artificial CRISPR / Cas12f1 complex.
12. The vector of claim 11 , wherein the second promoter sequence is a U6 promoter sequence.
13. The vector according to claim 11 or 12, wherein the vector is a viral vector.
14. The vector according to claim 11 or 12, wherein the vector is a plasmid vector.
15. 14. The vector of claim 13, wherein the vector is one or more selected from the group consisting of retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, and herpes simplex virus.
16. The vector of claim 11 , wherein the vector is a linear PCR amplicon.
17. A vector for expressing an artificial CRISPR / Cas12f1 complex capable of editing a nucleic acid comprising a first target sequence and a second target sequence, the vector comprising: a first sequence comprising a sequence encoding a Cas12f1 protein; a first promoter sequence operably linked to the first sequence; A second sequence comprising a sequence encoding a first artificial guide RNA, the first artificial guide RNA having a first scaffold sequence that interacts with the Cas12f1 protein, a first spacer sequence that is complementary to the first target sequence, and a first U-rich tail sequence, the first U-rich tail sequence being (U a N) n U b wherein N is selected from the group consisting of adenosine (A), uridine (U), cytidine (C) and guanosine (G); a is an integer between 1 and 4, inclusive; n is an integer selected from 0, 1 and 2, and b is an integer between 1 and 10, inclusive; and the U-rich tail sequence is not U, UU, or UUU; a second promoter sequence operably linked to the second sequence; a third sequence comprising a sequence encoding a second artificial guide RNA, the second artificial guide RNA having a second scaffold sequence that interacts with the Cas12f1 protein, a second spacer sequence that is complementary to the second target sequence, and a second U-rich tail sequence, the second U-rich tail sequence being (U a N) n U b wherein N is selected from adenosine (A), uridine (U), cytidine (C) and guanosine (G), a is an integer between 1 and 4, n is an integer selected from 0, 1 and 2, and b is an integer between 1 and 10, inclusive, and the U-rich tail sequence is not U, UU, or UUU; and a third promoter sequence operably linked to the third sequence; and A vector comprising:
18. The vector of claim 17 , wherein the second promoter sequence and the third promoter sequence are the same promoter sequence.
19. 18. The vector of claim 17, wherein the second promoter sequence is an H1 promoter sequence and the third promoter sequence is a U6 promoter sequence.
20. 1. An in vitro or ex vivo method for editing a nucleic acid comprising a target sequence in a cell, said method comprising: A method comprising delivering a Cas12f1 protein or a nucleic acid encoding the Cas12f1 protein and an artificial guide RNA or a nucleic acid encoding the artificial guide RNA into a cell, thereby forming a CRISPR / Cas12f1 complex in the cell; The nucleic acid containing the target sequence is edited by the CRISPR / Cas12f1 complex, and the artificial guide RNA comprises a scaffold sequence that interacts with the Cas12f1 protein, a spacer sequence complementary to the target sequence, and a U-rich tail sequence linked to the 3' end of the spacer sequence; The U-rich tail sequence is (U a N) n U b wherein N is selected from the group consisting of adenosine (A), uridine (U), cytidine (C) and guanosine (G), a is an integer between 1 and 4, n is an integer selected from 0, 1 and 2, and b is an integer between 1 and 10, inclusive, and the U-rich tail sequence is not U, UU, or UUU.
21. The method of claim 20, wherein the delivery involves introducing the Cas12f1 protein and the artificial guide RNA into the cell in the form of a CRISPR / Cas12f1 complex.
22. The method of claim 20, wherein the delivery comprises introducing into the cell a vector comprising the nucleic acid encoding the Cas12f1 protein and the nucleic acid encoding the artificial guide RNA.
23. 23. The method of claim 22, wherein the vector is one or more selected from the group consisting of retrovirus, lentivirus, adenovirus, adeno-associated virus (AAV), vaccinia virus, poxvirus, and herpes simplex virus.
24. 1. An in vitro or ex vivo method for editing a nucleic acid comprising a target sequence in a cell, said method comprising: contacting a CRISPR / Cas12f1 complex with said nucleic acid comprising said target sequence; The CRISPR / Cas12f1 complex comprises a Cas12f1 protein and an artificial guide RNA; The artificial guide RNA comprises a scaffold sequence that interacts with the Cas12f1 protein, a spacer sequence that is complementary to the target sequence, and a U-rich tail sequence; The U-rich tail sequence is (U a N) n U b wherein N is selected from the group consisting of adenosine (A), uridine (U), cytidine (C) and guanosine (G), a is an integer between 1 and 4, n is an integer selected from 0, 1 and 2, and b is an integer between 1 and 10, inclusive, and the U-rich tail sequence is not U, UU, or UUU.