Gene editing system based on sgRNA-donor DNA chimera
By covalently connecting sgRNA and donor DNA into chimera, an efficient CRISPR/Cas gene editing system was constructed, which solved the problems of low gene editing efficiency and high off-target rate in the prior art, and achieved efficient and low off-target gene editing effect.
Patent Information
- Application Number
- CN202311441854.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-01
- Publication Date
- 2025-05-06
AI Technical Summary
The existing CRISPR/Cas system has shortcomings in gene editing efficiency and off-target rate, especially when introducing donor DNA into cells, it is difficult to achieve efficient homologous recombination repair.
Long single-stranded RNA-DNA chimera are used to covalently connect sgRNA and donor DNA to form sgRNA-donor DNA chimera, which is used to build an efficient CRISPR/Cas gene editing system. The system guides the Cas9 enzyme to cleave target DNA through sgRNA and edits using donor DNA-mediated HDR.
It significantly improves the efficiency and accuracy of gene editing, reduces off-target rate, and makes high-precision gene editing of multiple types (replacement, insertion, and deletion).
Smart Images

Figure CN119932114A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of chemical biology and molecular biology. Specifically, the present disclosure relates to a gene editing system based on sgRNA-donor DNA chimera. Background Art
[0002] The clustered regularly interspaced shortpalindromic repeats (CRISPR) / Cas system is a simple and efficient gene editing tool that has been widely used for precise editing of multiple target genes in cells and in vivo. In a typical CRISPR / Cas9 system, three components are required, including CRISPR RNA (crRNA), trans-activating CRISPR RNA (tracrRNA) and CRISPR-Cas9 nuclease (referred to as Cas9 enzyme, Cas protein or Cas9). 1 . TracrRNA and Cas9 enzyme can produce a tight interaction to provide the necessary structural support, while crRNA and tracrRNA form a double-stranded structure and recognize the target genomic DNA. After the CRISPR / Cas system recognizes and cuts the genomic DNA, it can introduce exogenous DNA fragments (i.e., donor DNA) into the target area through the homologous recombination repair (HDR) process. 2 . In the traditional CRISPR / Cas system, since crRNA and tracrRNA need to be prepared separately, they must be assembled together to work; however, the stability of the assembly will affect the editing efficiency. In addition, introducing donor DNA into cells also faces additional challenges. Since donor DNA cannot specifically target the cell nucleus, its random distribution in the cell will reduce the efficiency of HDR and have the risk of causing non-homologous recombination. 3 .
[0003] To address these challenges, scientists have adopted a covalent ligation strategy to connect the different components of the CRISPR / Cas9 system. The most successful example is the formation of a single guide RNA (sgRNA) by connecting crRNA and tracrRNA, thereby improving the stability of the system and the efficiency of gene editing. 1This method simplifies the operation and improves editing efficiency by eliminating the assembly step and reducing disassembly during transport. It is particularly noteworthy that sgRNA can be easily obtained through a single transcription process, which greatly improves its convenience and applicability. On this basis, people have also connected donor DNA with the CRISPR / Cas9 enzyme system. For example, by using porcine circovirus 2 (PCV2) 4 or click chemistry 5 Connecting CRISPR-Cas nuclease and donor DNA can increase the editing efficiency to 8%. However, these strategies require modification of CRISPR / Cas nuclease, which is difficult to purify and has low yield. In addition, there are also reports on methods of connecting crRNA and donor DNA through click chemistry reactions. 6 In this method, crRNA and tracrRNA must be prepared separately, which affects the stability and efficiency of the system and has no significant improvement on the efficiency of HDR. Therefore, although the above methods all improve the final HDR efficiency, they still have their own limitations.
[0004] However, the prior art does not covalently link crRNA, tracrRNA and donor DNA, and its feasibility and impact on editing efficiency cannot be predicted.
[0005] References
[0006] 1.Jinek,M.;Chylinski,K.;Fonfara,I.;Hauer,M.;Doudna,JA;Charpentier,E.,A Programmable Dual-RNA-Guided DNAEndonuclease in Adaptive BacterialImmunity.Science 2012,337(6096),816-821.
[0007] 2. Mali, P.; Yang, L.; Esvelt, KM; Aach, J.; Guell, M.; DiCarlo, JE; Norville, JE; Church, GM, RNA-Guided Human Genome Engineering via Cas9. Science 2013, 339(6121), 823-826.
[0008] 3. Wang, J.Y.; Doudna, J.A., CRISPR technology: A decade of genome editing is only the beginning. Science 379(6629), eadd8643.
[0009] 4. Aird, E.J.; Lovendahl, K.N.; St. Martin, A.; Harris, R.S.; Gordon, W.R., Increasing Cas9-mediated homology-directed repair efficiency through covalent tethering of DNA repair template. Communications Biology 2018, 1(1), 54.
[0010] 5. Ling, X.; Xie, B.; Gao, X.; Chang, L.; Zheng, W.; Chen, H.; Huang, Y.; Tan, L.; Li, M.; Liu, T., Improving the efficiency of precise genome editing with site-specific Cas9-oligonucleotide conjugates. Science Advances 6(15), eaaz0051.
[0011] 6. Lee, K.; Mackley, V.A.; Rao, A.; Chong, A.T.; Dewitt, M.A.; Corn, J.E.; Murthy, N., Synthetically modified guide RNA and donor DNA are a versatile platform for CRISPR-Cas9 engineering. eLife 2017, 6, e25312. Summary of the Invention
[0012] Problem that the invention aims to solve
[0013] In response to the problems of low editing efficiency and high off-target rate in the current HDR-based CRISPR / Cas system, a new strategy of integrating sgRNA and donor DNA using long single-stranded RNA-DNA chimeras was proposed and developed to establish a new CRISPR / Cas gene editing system with high efficiency and low off-target performance.
[0014] Solutions for solving problems
[0015] [1]. A gene editing system comprising:
[0016] (i) a nucleic acid programmable DNA binding protein; and
[0017] (ii) sgRNA-donor DNA chimera,
[0018] The sgRNA-donor DNA chimera is a single-stranded structure, comprising a connected sgRNA and a donor DNA.
[0019] [2] The gene editing system according to [1], wherein the donor DNA comprises a 5' homology arm, an editing sequence and a 3' homology arm;
[0020] Preferably, the length of the 5' homology arm and the 3' homology arm are both 10-100 nucleotides, preferably 20-80 nucleotides.
[0021] [3] The gene editing system according to [1] or [2], wherein the sgRNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera; or, the sgRNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera;
[0022] Preferably, the sgRNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera.
[0023] [4] According to the gene editing system described in [3], in the sgRNA-donor DNA chimera, a linker sequence is also included between the sgRNA and the donor DNA.
[0024] [5] The gene editing system according to [4], wherein the linker sequence has a length of 1 to 200 nucleotides;
[0025] Preferably, the linker sequence comprises a length of 1 to 20 nucleotides;
[0026] More preferably, the linker sequence consists of thymidine deoxyribonucleotides.
[0027] [6] According to any one of [1] to [5], the sgRNA comprises a guide sequence and a scaffold sequence.
[0028] [7] The gene editing system according to any one of [1] to [6], wherein the nucleic acid programmable DNA binding protein has nuclease activity;
[0029] Preferably, the nucleic acid programmable DNA binding protein comprises a Cas protein;
[0030] More preferably, the nucleic acid programmable DNA binding protein comprises a Cas9 protein.
[0031] [8]. An isolated polynucleotide, wherein the polynucleotide encodes a gene editing system as described in any one of [1] to [7].
[0032] [9]. A vector comprising the polynucleotide described in [8].
[0033]
[10] A cell comprising one or more of the following:
[0034] (a) The gene editing system as described in any one of [1] to [7];
[0035] (b) the polynucleotide described in [8]; and
[0036] (c) The vector as described in [9].
[0037]
[11] . A pharmaceutical composition comprising one or more of the following (A) to (D):
[0038] (A) The gene editing system as described in any one of [1] to [7];
[0039] (B) the polynucleotide described in [8];
[0040] (C) the vector described in [9]; and
[0041] (D) the cell as described in
[10] ;
[0042] And, optionally, a pharmaceutically acceptable carrier.
[0043]
[12] . A kit comprising one or more of the following:
[0044] (A) The gene editing system as described in any one of [1] to [7];
[0045] (B) the polynucleotide described in [8];
[0046] (C) the vector described in [9]; and
[0047] (D) Cells as described in
[10] .
[0048]
[13] A method for editing a target gene, the method comprising the step of contacting the target gene with the gene editing system described in any one of [1] to [7], the polynucleotide described in [8], the vector described in [9], the cell described in
[10] , the pharmaceutical composition described in
[11] , or the kit described in
[12] ;
[0049] Preferably, the editing includes at least one of replacement, insertion and deletion.
[0050]
[14] Use of the gene editing system described in any one of [1] to [7], the polynucleotide described in [8], the vector described in [9] or the cell described in
[10] in the preparation of a reagent for editing a target gene;
[0051] Preferably, the editing includes at least one of replacement, insertion and deletion.
[0052]
[15] . A method for preventing and / or treating a disease caused by a disease-related gene, comprising the step of administering to a subject a therapeutically effective amount of the gene editing system described in any one of [1] to [7], the polynucleotide described in [8], the vector described in [9], the cell described in
[10] , the pharmaceutical composition described in
[11] , or the kit described in
[12] .
[0053]
[16] . Use of the gene editing system described in any one of [1] to [7], the polynucleotide described in [8], the vector described in [9], the cell described in
[10] , the pharmaceutical composition described in
[11] or the kit described in
[12] in the preparation of a drug for preventing and / or treating diseases caused by disease-related genes.
[0054] Effects of the Invention
[0055] In the gene editing system based on sgRNA-donor DNA chimera provided in the present disclosure, crRNA, tracrRNA and donor DNA are covalently linked to achieve stronger stability and higher editing efficiency. In the present disclosure, sgRNA is linked to donor DNA to achieve multi-type (replacement, insertion, deletion) and high-precision gene editing.
[0056] More specifically, the present disclosure provides a new gene editing tool based on long single-stranded sgRNA-donor DNA chimeras. Design and prepare sgRNA-donor DNA chimeras containing sgRNA targeting any specific genomic DNA and donor DNA sequences targeting specific base editing targets, guide CRISPR / Cas9 through its sgRNA part and stimulate its nuclease activity, and use HDR mediated by the donor DNA part to replace the target gene sequence with the sequence information on it, so as to achieve gene editing such as target base replacement and fragment insertion. At the same time, there are homology arm sequences on the donor DNA, which also have the ability to recognize the target sequence, and together with the gRNA, a dual targeting system is realized to achieve the effect of cascade amplification, significantly reducing the off-target rate of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 A schematic diagram showing efficient base substitution editing based on sgRNA-donor DNA chimeras;
[0058] Figure 2 A schematic diagram showing a highly efficient fragment insertion method based on sgRNA-donor DNA chimeras;
[0059] Figure 3A and Figure 3B The HDR efficiency determined by flow cytometry and high-throughput sequencing of the eGFP gene to eBFP gene based on the above-mentioned base substitution editing method are shown;
[0060] Figure 4 A data graph showing that sgRNA-donor DNA improves HDR efficiency in a concentration-dependent manner;
[0061] Figure 5 The HDR efficiency determined by flow cytometry based on the above-mentioned fragment insertion method in which the Y-FAST gene is inserted into the N-terminus of the ACTB gene is shown. DETAILED DESCRIPTION
[0062] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below. The word "exemplary" used herein means "used as an example, embodiment or illustrative". Any embodiment described herein as "exemplary" is not necessarily to be interpreted as being superior or better than other embodiments.
[0063] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In other examples, methods, means, equipment and steps well known to those skilled in the art are not described in detail in order to highlight the main purpose of the present disclosure.
[0064] Unless otherwise stated, the units used in this specification are all international standard units, and the numerical values and numerical ranges appearing in this disclosure should be understood to include the inevitable systematic errors in industrial production.
[0065] In this specification, the word "may" includes both performing a certain process and not performing a certain process.
[0066] In this specification, the references to "some specific / preferred embodiments", "other specific / preferred embodiments", "embodiments", etc., mean that the specific elements (e.g., features, structures, properties and / or characteristics) described in connection with the embodiments are included in at least one embodiment described herein, and may or may not exist in other embodiments. In addition, it should be understood that the elements may be combined in various embodiments in any suitable manner.
[0067] In this specification, the numerical range expressed using "a numerical value A to a numerical value B" means a range including the endpoints numerical values A and B.
[0068] In the present specification, "optional" or "optionally" means that the event or situation described below may or may not occur, and the description includes cases where the event occurs and cases where it does not occur.
[0069] As used in the present disclosure, the terms "nucleic acid" and "nucleic acid molecule" refer to compounds comprising a core base and an acidic portion, such as a polymer of a nucleoside, a nucleotide or a nucleotide. Typically, a polymeric nucleic acid, such as a nucleic acid molecule comprising three or more nucleotides is a linear molecule in which adjacent nucleotides are interconnected by a phosphodiester bond. In some embodiments, "nucleic acid" refers to a single nucleic acid residue (such as a nucleotide and / or a nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" are used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" includes RNA and single-stranded and / or double-stranded DNA. Nucleic acid can be naturally occurring, such as in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or their fragments, or synthetic DNA, RNA, DNA / RNA hybrids or nucleotides or nucleosides that comprise non-naturally occurring. In addition, term " nucleic acid ", " DNA ", " RNA " and / or similar terms comprise nucleic acid analogs, such as analogs with other skeletons except the phosphodiester backbone. Nucleic acid can be purified from natural sources, use recombinant expression systems to produce and optionally purify, chemosynthesis etc. In suitable cases, such as in the case of chemosynthesized molecules, nucleic acid can comprise nucleoside analogs, such as analogs or sugar and backbone modifications with chemically modified bases.
[0070] As used in this disclosure, the terms "polypeptide," "peptide," and "protein" are used interchangeably herein and are amino acid polymers of any length. The polymer may be linear or branched, it may contain modified amino acids, and it may be interrupted by non-amino acids. The term also includes amino acid polymers that have been modified (e.g., disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation, such as conjugation with a labeling component).
[0071] As used in the present disclosure, CRISPR-Cas system or CRISPR / Cas system refers to a class of bacterial systems for defending against exogenous nucleic acids. CRISPR-Cas systems are widely found in eubacteria and archaeal organisms. CRISPR-Cas systems include type I, type II, and type III subtypes. Wild-type type II CRISPR-Cas systems utilize RNA-mediated nucleases (e.g., Cas9 proteins) in combination with guide and activation RNAs (e.g., single guide RNAs or sgRNAs) to recognize and cleave exogenous nucleic acids, i.e., exogenous nucleic acids comprising natural or modified nucleotides.
[0072] As used in the present disclosure, the term "Cas protein" refers to a clustered regularly spaced short palindromic repeats associated protein or nuclease. The Cas protein can be a wild-type Cas protein or a Cas protein variant. The Cas9 protein is an example of a Cas protein belonging to a type II CRISPR / Cas system (e.g., Rath et al., Biochimie 117: 119, 2015). Naturally occurring Cas proteins require crRNA and tracrRNA for site-specific DNA recognition and cutting. The crRNA associates with the tracrRNA through a partially complementary region to guide the Cas protein to a region homologous to the crRNA in the target DNA called a "protospacer adjacent motif" (PAM). The naturally occurring Cas protein cuts DNA at a site specified by a guide sequence contained in the crRNA transcript to generate a flat end at a double-strand break.
[0073] As used in the present disclosure, the term "Cas protein variant" refers to a Cas protein having at least one amino acid substitution (e.g., one, two, three, four, five, six, seven, eight, nine, ten or more amino acid substitutions) relative to a wild-type Cas protein sequence, and / or a truncated form or fragment of a wild-type Cas protein. In some embodiments, the Cas protein variant has at least 75% sequence identity (e.g., at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94% 95%, 96%, 97%, 98%, 99% or 100% sequence identity) with a wild-type Cas protein sequence. In some embodiments, the Cas protein variant is a fragment of a wild-type Cas protein and has at least one amino acid substitution relative to a wild-type Cas protein sequence. The Cas protein variant may be a Cas9 protein variant. In some embodiments, the Cas protein variant has nuclease activity.
[0074] As used in the present disclosure, the term "guide RNA" or "gRNA" refers to a DNA targeting RNA that can guide an RNA-guided nuclease (e.g., Cas protein) to a target gene by hybridizing with the target gene. In some embodiments, the guide RNA can be a single guide RNA (sgRNA), which includes a guide sequence that targets the RNA-guided nuclease to the target gene (i.e., the crRNA equivalent of a single guide RNA) and a scaffold sequence that interacts with the RNA-guided nuclease (i.e., the tracrRNA equivalent of a single guide RNA). In other embodiments, the guide RNA may include two components, namely, a guide sequence that targets the RNA-guided nuclease to the target gene (i.e., the crRNA equivalent of a single guide RNA) and a scaffold sequence that interacts with the RNA-guided nuclease (i.e., the tracrRNA equivalent of a single guide RNA). A portion of the guide sequence can be hybridized with a portion of the scaffold sequence to form a two-component guide RNA.
[0075] As used in the present disclosure, the term "single guide RNA" or "sgRNA" refers to a DNA targeting RNA that comprises a guide sequence that targets a Cas protein to a target DNA (i.e., the crRNA equivalent portion of a single guide RNA) and a scaffold sequence that interacts with the Cas protein (i.e., the tracrRNA equivalent portion in a single guide RNA).
[0076] As used in the present disclosure, the term "crRNA" includes a repeat sequence (repeat) and a spacer sequence (spacer), and CRISPR transcribes to form a long chain of pre-CRISPR RNA (pre-crRNA), and pre-crRNA is processed to obtain a short crRNA containing a repeat region sequence and a spacer region sequence. In some CRISPR / Cas systems, crRNA is obtained by the action of Cas protein on pre-crRNA. In other CRISPR / Cas systems, crRNA is obtained by the joint action of Cas protein and tracrRNA (trans-activating crRNA) on pre-crRNA.
[0077] As used in the present disclosure, "guide sequence of crRNA" refers to a sequence in crRNA that hybridizes with a target sequence of a target nucleic acid (target gene), which is correspondingly formed by a spacer sequence of crRNA.
[0078] As used in the present disclosure, the term "target sequence" refers to a nucleotide sequence in a target nucleic acid (target gene) that is complementary or at least partially complementary to crRNA. After the Cas protein, crRNA and target sequence form a ternary complex, the Cas protein exerts specific cutting activity on the target chain and / or non-target chain in the target nucleic acid (target gene).
[0079] As used in the present disclosure, the term "target strand" refers to a nucleotide strand in a target nucleic acid (target gene) that hybridizes with crRNA; the term "non-target strand" refers to a nucleotide strand in a target nucleic acid (target gene) that does not hybridize with crRNA.
[0080] As used in the present disclosure, the terms "nucleic acid programmable nucleotide binding domain", "nucleic acid programmable DNA-binding protein (napDNAbp)" refer to a protein that binds to a nucleic acid (e.g., DNA or RNA), such as a guide polynucleotide (e.g., gRNA), which guides the napDNAbp to a specific nucleic acid sequence, for example, by hybridizing with a target sequence of a target gene. For example, the Cas9 protein can be bound to a guide RNA that guides the Cas9 protein to a specific DNA sequence that is complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as a nuclease-active Cas9, a Cas9 nickase (nCas9), or a nuclease-free Cas9 (dCas9). Examples of nucleic acid programmable DNA-binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, C2c1, C2c2, C2c3, and Argonaute protein (AGO). However, it should be understood that nucleic acid programmable DNA-binding proteins also include nucleic acid programmable proteins that bind RNA. For example, the napDNAbp can be bound to a nucleic acid that guides the napDNAbp to the RNA.Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically described in the present disclosure.
[0081] As used in the present disclosure, the term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., a protein comprising an active, inactive or partially active DNA cleavage domain of Cas9, and / or a binding domain of gRNA Cas9). Cas9 nucleases are sometimes also referred to as CRISPR-associated nuclease 9. As previously described, CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster contains numerous short and conserved repeat regions and spacers. The CRISPR cluster is transcribed and processed into pre-crRNA. In the type II CRISPR / cas9 system, transcoded small RNA (tracrRNA), endogenous ribonuclease 3 (RNase III), and Cas9 protein are required for the correct processing of pre-crRNA. TracrRNA serves as a guide for ribonuclease 3 to assist in the processing of pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleases a linear or circular dsDNA target complementary to the spacer sequence. The target strand that is not complementary to crRNA is cut by endonucleolytic cutting. In nature, DNA binding and cutting usually require proteins and two RNAs. However, a single guide RNA (sgRNA) can be engineered to integrate various aspects of crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif in the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.Cas9 nuclease sequence and structure are well known to those skilled in the art (see, e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti et al., Proc. Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607 (2011); and "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M. et al., Science 337:816-821 (2012)). Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun and Charpentier, “The tracrRNA and Cas9 families of type IICRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.
[0082] As used in the present disclosure, Cas9 nickase can cut a chain of double-stranded DNA.Cas9 nickase can be generated by introducing an inactivating mutation into the HNH subdomain or the RuvC subdomain.For example, an inactivating mutation (D10A) can be introduced into the RuvC domain of Streptococcus pyogenes Cas9, while the HNH domain remains active, i.e., the residue at position 840 remains as histidine.Such Cas9 variants can generate single-stranded DNA breaks (nicks) at specific positions based on the target sequence determined by gRNA.Those skilled in the art can identify the catalytic residues in the RuvC and HNH domains of any known Cas9 protein and introduce inactivating mutations to generate corresponding dCas9 or nCas9.
[0083] Similarly, for other Cas proteins, those skilled in the art can obtain the corresponding Cas proteins without nuclease activity and the nicking enzyme that cuts one strand of double-stranded DNA in the same manner.
[0084] <Gene editing system based on sgRNA-donor DNA chimera>
[0085] In some aspects of the present disclosure, a gene editing system is provided, comprising:
[0086] (i) Nucleic acid programmable DNA binding proteins (napDNAbp) ;and
[0087] (ii) sgRNA-donor DNA chimera (ssRDC),
[0088] The sgRNA-donor DNA chimera is a single-stranded structure, comprising a connected sgRNA and a donor DNA.
[0089] Nucleic acid programmable DNA binding protein (napDNAbp)
[0090] In the present disclosure, there is no particular limitation on nucleic acid programmable DNA binding proteins, and in some embodiments, nucleic acid programmable DNA binding proteins have nuclease activity. For example, nucleic acid programmable DNA binding proteins can edit target genes by cutting target genes. Then, the cut target gene can be repaired by homologous recombination with donor DNA in a nearby sgRNA-donor DNA chimera. For example, Cas protein (Cas nuclease) can directly cut one or two chains at a certain position of the target gene. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, homologs of the above, variants of the above, mutants of the above, and derivatives of the above. There are three main types of Cas proteins (type I, type II and type III) and 10 subtypes (including 5 type I, 3 type II and 2 type III proteins) (for example, see Hochstrasser and Doudna, Trends Biochem Sci, 2015: 40 (1): 58-66). Type II Cas proteins include Cas1, Cas2, Csn2, Cas9 and Cfp1. These Cas nucleases are known to those skilled in the art. For example, the amino acid sequence of the wild-type Cas9 polypeptide of Streptococcus pyogenes is listed, for example, in NBCI Ref. Seq. No. NP_269215, and the amino acid sequence of the wild-type Cas9 polypeptide of Streptococcus thermophilus is listed, for example, in NBCI Ref. Seq. No. WP_011681470.
[0091] In some preferred embodiments, the napDNAbp is Cas9 or a variant thereof.
[0092] sgRNA-donor DNA chimeras
[0093] In the present disclosure, sgRNA and donor DNA are fused on a single strand to form an sgRNA-donor DNA chimera, thereby optimizing and improving existing HDR-based gene editing tools, and guiding multi-type (replacement, insertion, deletion) and high-precision gene editing in cells.
[0094] In some embodiments, the sgRNA-donor DNA chimera is a single-stranded structure. In some specific embodiments, the sgRNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera. In other specific embodiments, the sgRNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera.
[0095] In some preferred embodiments, the sgRNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera. The present invention finds that when the sgRNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera, the editing efficiency (e.g., cleavage activity) of the gene editing system is significantly better than when the sgRNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera.
[0096] In some optional embodiments, in the sgRNA-donor DNA chimera, a linker sequence is also included between the sgRNA and the donor DNA.
[0097] In some alternative embodiments, the linker sequence can be a stretch of nucleotides and has a length of 1 to 20 nucleotides (nt). In some embodiments, the linker sequence is a stretch of nucleotides and has a length of 1 to 10 nucleotides (nt).
[0098] In some exemplary embodiments, the linker sequence is 1 nucleotide (nt) in length, specifically thymidine deoxyribonucleotide (T).
[0099] (sgRNA)
[0100] The Cas protein can be guided to its target DNA by a single guide RNA (sgRNA). sgRNA is a form of naturally occurring two-segment guide RNA (crRNA and tracrRNA) that is engineered into a single continuous sequence. The sgRNA can include a guide sequence that targets the Cas protein to the target DNA (e.g., a crRNA equivalent of the sgRNA), and a scaffold sequence that interacts with the Cas protein (e.g., a tracrRNA equivalent of the sgRNA). sgRNA can be selected using software. As a non-limiting example, considerations for selecting sgRNA can include, for example, the PAM sequence of the Cas protein to be used, and strategies for minimizing off-target editing. Such as The tools of CRISPR design tools can provide sequences for preparing sgRNAs for evaluating the editing efficiency of target genes and / or evaluating the cutting of off-target sites. In some specific embodiments, the length of sgRNA is about 100 nt.
[0101] Boot Sequence
[0102] The guide sequence in the sgRNA may be complementary to a specific sequence in the target DNA. The 3' end of the target DNA sequence may be followed by a PAM sequence. About 20 nucleotides upstream of the PAM sequence is the target DNA. Typically, the Cas9 protein or its variants cut about three nucleotides upstream of the PAM sequence. The guide sequence in the sgRNA can be complementary to either strand of the target DNA.
[0103] In some embodiments, the guide sequence of the sgRNA is at the 5' end of the sgRNA that can use RNA-DNA complementary base pairing to guide the Cas protein to the target DNA site, and can include about 10 to about 2000 nucleotides, for example, about 10 to about 100 nucleotides, about 10 to about 500 nucleotides, about 10 to about 1000 nucleotides, about 10 to about 1500 nucleotides, about 10 to about 2000 nucleotides, about 50 to about 100 nucleotides, about 50 to about 500 nucleotides, about 50 to about 10 00 nucleotides, about 50 to about 1500 nucleotides, about 50 to about 2000 nucleotides, about 100 to about 500 nucleotides, about 100 to about 1000 nucleotides, about 100 to about 1500 nucleotides, about 100 to about 2000 nucleotides, about 500 to about 1000 nucleotides, about 500 to about 1500 nucleotides, about 500 to about 2000 nucleotides, about 1000 to about 1500 nucleotides, about 1000 to about 2000 nucleotides, or about 1500 to about 2000 nucleotides. In some embodiments, the guide sequence of the sgRNA comprises about 100 nucleotides at the 5' end of the sgRNA that can use RNA-DNA complementary base pairing to guide the Cas protein to the target DNA site. In some embodiments, the guide sequence comprises 20 nucleotides at the 5' end of the sgRNA that can use RNA-DNA complementary base pairing to guide the Cas protein to the target DNA site. In other embodiments, the guide sequence comprises less than 20 (e.g., 19, 18, 17, 16, 15 or less) nucleotides complementary to the target DNA site. In some cases, the guide sequence in the sgRNA comprises at least one nucleotide mismatch in the complementary region of the target DNA site. In some cases, the guide sequence comprises about 1 to about 10 nucleotide mismatches (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotide mismatches) in the complementary region of the target DNA site.
[0104] Scaffold sequence
[0105] The scaffold sequence in sgRNA can be used as a protein binding sequence that interacts with Cas protein or its variants. In some embodiments, the scaffold sequence in sgRNA can include two complementary nucleotide fragments that hybridize to form a double-stranded RNA duplex (dsRNA duplex). The scaffold sequence can have a structure such as a lower stem, a protrusion, an upper stem, a nexus and / or a hairpin. In some embodiments, the scaffold sequence in sgRNA can be about 90 nucleotides to about 120 nucleotides, for example, about 90 nucleotides to about 115 nucleotides, about 90 nucleotides to about 110 nucleotides, about 90 nucleotides to about 105 nucleotides, about 90 nucleotides to about 100 nucleotides, about 90 nucleotides to about 95 nucleotides, about 95 nucleotides to about 120 nucleotides, about 100 nucleotides to about 120 nucleotides, about 105 nucleotides to about 120 nucleotides, about 110 nucleotides to about 120 nucleotides, or about 115 nucleotides to about 120 nucleotides.
[0106] In some specific embodiments, the scaffold sequence of the sgRNA comprises the nucleotide sequence shown in SEQ ID NO:4, or a nucleotide sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with SEQ ID NO:4.
[0107] (Donor DNA)
[0108] In the present disclosure, the donor DNA may include a 5' homology arm (or upstream homology arm, left homology arm), an editing sequence (e.g., an exogenous nucleotide sequence and / or a sequence encoding a heterologous protein or a fragment thereof), and a 3' homology arm (or downstream homology arm, right homology arm).
[0109] Under the guidance of sgRNA, Cas protein generates site-specific double-strand breaks (DSBs) or single-strand breaks (SSBs) in double-stranded DNA (dsDNA) target genes in some cases (for example, when Cas protein is a nickase variant), which is repaired by HDR. The gene editing system of the present disclosure is introduced into the cell (for example, by methods such as electrotransfection, liposome transfection or microinjection), and the donor DNA is integrated into the target gene. In some embodiments, an exogenous nucleotide sequence is introduced into the cell, and the editing includes inserting the exogenous nucleotide sequence into the target gene. In other embodiments, the editing includes excision of the target gene. In some embodiments, the gene editing system provided by the present disclosure can be performed in vivo, in vitro or in vitro for gene editing. In some embodiments, the gene editing system provided by the present disclosure enables the use of a lower concentration or amount of the donor DNA described herein relative to the concentration or amount of the corresponding donor DNA to achieve the same or higher gene editing (e.g., knock-in) efficiency.
[0110] In the present disclosure, "homologous arms" have the same meaning as commonly understood by those skilled in the art, and refer to flanking sequences on both sides of the editing sequence (target sequence) on the donor DNA that are completely consistent with the genomic sequence and are used to identify and undergo recombination.
[0111] In some embodiments of the present disclosure, the length of the 5' homology arm and the 3' homology arm are both 10-100 nucleotides, preferably 20 nt-80 nt. The present disclosure finds that under this length, the gene editing system of the present invention has excellent gene editing efficiency.
[0112] In some embodiments, the length and sequence of the edited sequence are not particularly limited, and can be determined based on the length of the edited sequence to be inserted, replaced, or deleted.
[0113] In some embodiments, when a target gene is deleted, the gene editing system provided by the present invention can delete 1-500 nucleotides in the target gene.
[0114] In some embodiments, the sgRNA-donor DNA chimera can be obtained by solid phase synthesis or short chain splicing.
[0115] In some specific embodiments, the sgRNA-donor DNA chimera has the following structure:
[0116] [sgRNA]-[optional linker sequence]-[donor DNA].
[0117] In some more specific embodiments, the sgRNA-donor DNA chimera has the following structure:
[0118] [Guide sequence]-[Scaffold sequence]-[Optional linker sequence]-[5' homology arm]-[Editing sequence]-[3' homology arm].
[0119] <Polynucleotide>
[0120] The present disclosure provides an isolated polynucleotide encoding the above-mentioned gene editing system.
[0121] In some specific embodiments, the polynucleotide comprises one or more of the following:
[0122] i) a nucleotide sequence encoding a nucleic acid programmable DNA binding protein in the above-mentioned gene editing system; and,
[0123] ii) A nucleotide sequence encoding the sgRNA-donor DNA chimera in the above-mentioned gene editing system.
[0124] In some embodiments, the nucleotide sequence of the nucleic acid programmable DNA binding protein is codon optimized. This type of optimization may require mutation of the nucleotide sequence encoding the nucleic acid programmable DNA binding protein to mimic the codon preference of the intended host organism or cell while encoding the same protein.
[0125] <Carrier>
[0126] The present disclosure provides a vector comprising the above-mentioned polynucleotide.
[0127] In some specific embodiments, the carrier comprises one or more of the following:
[0128] (i) a polynucleotide encoding a nucleic acid programmable DNA binding protein in the above-mentioned gene editing system; and,
[0129] (ii) A polynucleotide encoding the sgRNA-donor DNA chimera in the above-mentioned gene editing system.
[0130] In some embodiments, the above (i) to (ii) may be in the same vector. In other embodiments, the above (i) to (ii) may be in different vectors.
[0131] In some embodiments, the vector is an expression vector, more specifically a recombinant expression vector. Suitable expression vectors include viral expression vectors (e.g., viral vectors based on the following viruses, vaccinia virus, polio virus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviral vectors (e.g., murine leukemia virus, spleen necrosis virus, and vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukemia virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus), etc.
[0132] Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, and the like may be used in the expression vector.
[0133] Methods for introducing nucleic acid into host cells are known in the art, and any convenient method can be used to introduce nucleic acid (e.g., expression construct) into cells. Suitable methods include, for example, viral infection, transfection, liposome transfection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc.
[0134] <Cell>
[0135] The present disclosure provides a cell comprising one or more of the following:
[0136] (a) the gene editing system described above;
[0137] (b) the polynucleotide described above; and,
[0138] (c) The above-mentioned vector.
[0139] The cell can be any of a variety of cells, including, for example, in vitro cells, in vivo cells, ex vivo cells, primary cells, cancer cells, animal cells, plant cells, algae cells, fungal cells, and the like.
[0140] In some embodiments, the cell is a receptor for the gene editing system provided by the present disclosure, which may also be referred to as a "host cell" or a "target cell". A host cell or a target cell may be a receptor for the gene editing system provided by the present disclosure. A host cell or a target cell may be a receptor for the gene editing system provided by the present disclosure. A host cell or a target cell may be a receptor for a single component in the gene editing system provided by the present disclosure.
[0141] In some specific embodiments, non-limiting examples of cells include: prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of unicellular eukaryotic organisms, protozoan cells, cells from plants, algal cells, fungal cells, animal cells, cells from invertebrates, cells from vertebrates, cells from mammals (e.g., ungulates; rodents; non-human primates; humans; cats; dogs, etc.), etc. In some cases, the cell is a cell that is not derived from a natural organism (e.g., the cell can be a synthetic cell; also known as an artificial cell).
[0142] <Pharmaceutical Composition and Kit>
[0143] The present disclosure provides a pharmaceutical composition comprising one or more of the following:
[0144] (A) The gene editing system described above;
[0145] (B) the above-mentioned polynucleotide;
[0146] (C) the above-mentioned vector; and,
[0147] (D) The above cells.
[0148] In some optional embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier.
[0149] In some embodiments, the pharmaceutically acceptable carrier may be a delivery carrier, such as a lipid, a cationic lipid or other polymers having drug delivery function.
[0150] The present disclosure provides a reagent or kit comprising one or more of the following:
[0151] (A) The gene editing system described above;
[0152] (B) the above-mentioned polynucleotide;
[0153] (C) the above-mentioned vector; and,
[0154] (D) The above cells.
[0155] In some specific embodiments, the kit can be a disease treatment kit. In other specific embodiments, the kit can be a kit for target gene editing.
[0156] <Methods and uses of editing target genes>
[0157] The present disclosure provides a method for editing a target gene, the method comprising contacting the target gene with a gene editing system provided by the present disclosure, a polynucleotide provided by the present disclosure, a vector provided by the present disclosure, a cell provided by the present disclosure, a pharmaceutical composition provided by the present disclosure, or a kit provided by the present disclosure. In some embodiments, the contact results in editing of the target gene by the gene editing system. The present disclosure also provides the use of the gene editing system provided by the present disclosure, the polynucleotide provided by the present disclosure, the vector provided by the present disclosure, and the cell provided by the present disclosure in the preparation of a reagent for editing a target gene.
[0158] In some exemplary embodiments, a CRISPR / Cas9 gene editing system based on RNA-DNA chimeras is constructed: equimolar amounts of Cas9 protein and sgRNA-donor DNA chimeras are fully mixed in an incubation buffer and incubated at 25°C for 15 minutes. The incubation mixture is electroporated into a corresponding number of cells by transfection. After transfection, the cells are cultured in a 24-well plate in a 5% CO2, 37°C incubator for 3-5 days.
[0159] In some specific embodiments, the editing is a change in a nucleotide in the target gene. In some specific embodiments, the nucleotide change can be a single nucleotide substitution (e.g., transition or transversion change), deletion or insertion.
[0160] In some specific embodiments, the nucleotide change can be (1) a G to T substitution, (2) a G to A substitution, (3) a G to C substitution, (4) a T to G substitution, (5) a T to A substitution, (6) a T to C substitution, (7) a C to G substitution, (8) a C to T substitution, (9) a C to A substitution, (10) an A to T substitution, (11) an A to G substitution, or (12) an A to C substitution.
[0161] In some specific embodiments, the nucleotide change can be a conversion of (1) a G:C base pair to a T:A base pair,
[0162] (2) G:C base pair to A:T base pair, (3) G:C base pair to C:G base pair, (4) T:A base pair to G:C base pair, (5) T:A base pair to A:T base pair, (6) T:A base pair to C:G base pair, (7) C:G base pair to G:C base pair, (8) C:G base pair to T:A base pair, (9) C:G base pair to A:T base pair, (10) A:T base pair to T:A base pair, (11) A:T base pair to G:C base pair, or (12) A:T base pair to C:G base pair.
[0163] In some specific embodiments, the nucleotide change can be an insertion. In some cases, the length of the insertion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.
[0164] In some specific embodiments, the nucleotide change can be a deletion (deletion). In some other cases, the length of the deletion is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, or at least 500 nucleotides.
[0165] In some specific embodiments, the contacting occurs in vitro or in vivo. In some specific embodiments, the contacting occurs inside a cell or outside a cell.
[0166] In some specific embodiments, the cell is a eukaryotic cell or a prokaryotic cell.
[0167] In some more specific embodiments, the cell is selected from the group consisting of: plant cells, fungal cells, mammalian cells, reptile cells, insect cells, avian cells, fish cells, parasite cells, arthropod cells, invertebrate cells, vertebrate cells, rodent cells, mouse cells, rat cells, primate cells, non-human primate cells and human cells.
[0168] <Methods and uses for treating diseases>
[0169] The present disclosure provides a method for preventing and / or treating diseases caused by disease-related genes, which comprises the step of administering to a subject a therapeutically effective amount of the gene editing system, polynucleotide, vector, cell, pharmaceutical composition or kit described in the present disclosure.
[0170] In the present invention, by editing the disease-related genes through the gene editing system, polynucleotide, vector, cell, pharmaceutical composition or kit of the present invention, it is possible to achieve upregulation, downregulation, inactivation, activation, mutation correction or introduction of disease-related sites of disease-related genes, thereby achieving the prevention and / or treatment of diseases and / or the creation of disease-related models. For example, the target gene described in the present invention may be located in the protein coding region of the disease-related gene, or, for example, may be located in a gene expression regulatory region such as a promoter region or an enhancer region, so that the editing of the disease-related gene function or the editing of the disease-related gene expression can be achieved. Therefore, the editing of disease-related genes described herein includes editing of the disease-related gene itself (e.g., protein coding region), and also includes editing of its expression regulatory region (e.g., promoter, enhancer, intron, etc.).
[0171] "Disease-related" gene refers to any gene that produces a transcription or translation product at an abnormal level or in an abnormal form in cells derived from tissues affected by the disease, compared to tissues or cells of non-disease controls. In the case where the altered expression is related to the appearance and / or progression of the disease, it can be a gene expressed at an abnormally high level; it can be a gene expressed at an abnormally low level. Disease-related genes also refer to genes with one or more mutations or genetic variations that are directly responsible or unbalanced with one or more genes responsible for the etiology of the disease. The mutation or genetic variation is, for example, a single nucleotide variation (SNV). The transcribed or translated product can be known or unknown, and can be at normal or abnormal levels.
[0172] Therefore, the present invention also provides a method for treating a disease in a subject in need thereof, comprising delivering an effective amount of the gene editing system, polynucleotide, vector, cell, pharmaceutical composition or kit of the present invention to the subject to edit a gene associated with the disease.
[0173] The present disclosure also provides use of the gene editing system, polynucleotide, vector, cell, pharmaceutical composition or kit described in the present disclosure in the preparation of a drug for preventing and / or treating diseases caused by disease-related genes.
[0174] The present disclosure also provides the gene editing system, polynucleotide, vector, cell, pharmaceutical composition or kit described in the present disclosure, which is used to prevent and / or treat diseases caused by disease-related genes.
[0175] Example
[0176] The embodiments of the present disclosure will be described in detail below in conjunction with the examples, but those skilled in the art will appreciate that the following examples are only used to illustrate the present disclosure and should not be considered to limit the scope of the present disclosure. Where specific conditions are not specified in the examples, they are carried out under conventional conditions or conditions recommended by the manufacturer. Where the manufacturers of the reagents or instruments used are not specified, they are all conventional products that can be obtained commercially.
[0177] The experimental techniques and experimental methods used in this example are all conventional technical methods unless otherwise specified. For example, the experimental methods in the following examples that do not specify specific conditions are usually carried out under conventional conditions such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or under conditions recommended by the manufacturer. The materials, reagents, etc. used in the examples can be obtained through regular commercial channels unless otherwise specified.
[0178] Experimental methods
[0179] Characterization of gene editing efficiency in transfected cells
[0180] Construct a sequencing library and use high-throughput sequencing technology to analyze the gene sequence of the target site. The specific implementation steps are as follows:
[0181] A. The amplification primers containing Illumina forward and reverse adapters are different for different editing sequences, and the adapters are the same. The specific structure is as follows:
[0182] Amplification primers for forward adapter: forward adapter + forward specific binding sequence;
[0183] Amplification primer for reverse adapter: reverse adapter + reverse specific binding sequence.
[0184] The sequence of the forward adapter is as follows (SEQ ID NO: 1):
[0185] ACACTCTTTCCCTACACGACCGCTTCCGATCT
[0186] The sequence of the reverse adapter is as follows (SEQ ID NO: 2):
[0187] TGGAGTTCAGACGTGTGCTCTTCCGATCT
[0188] The forward specific binding sequence and the reverse specific binding sequence are sequences that specifically bind to different sequence sites, and their sequences are (N) n Wherein, N is A, T, G or C; n=15-30, preferably 18-24.
[0189] For the first round of PCR (PCR 1) to amplify the genomic region of interest. PCR 1 reaction (25 μl) was performed with 0.5 μM each forward and reverse primer, 1 μL genomic DNA extract and 12.5 μL PCR premix. The PCR reaction was performed as follows: 98°C for 2 minutes followed by 30 cycles [98°C for 10 seconds, 60°C for 10 seconds, 72°C for 30 seconds], followed by a final extension of 72°C for 2 minutes.
[0190] B. The unique Illumina barcode (index) primer (the primer is provided by the kit and has a series of different sequences. The primer product number is 12412ES02, and the sequence information is detailed in the instructions (Yi Sheng Company, Hieff 384CDIPrimer for ) pairs were added to each sample. PCR 2 reactions were performed by mixing 0.5 μM of each unique forward and reverse Illumina barcode primer pair, 1 μl of unpurified PCR 1 reaction mixture, and 12.5 μL of PCR master mix. Reactions were performed as follows: 98°C for 2 minutes, followed by 12 cycles of [98°C for 10 seconds, 60°C for 10 seconds, and 72°C for 30 seconds], followed by a final extension at 72°C for 2 minutes.
[0191] C. Purify the PCR 2 product (pooled from the common amplicons) by 1.5% agarose gel electrophoresis using the QIAquick Gel Extraction Kit (Qiagen) and elute with 10 μL of water.
[0192] D. Take 5 μL and sequence it on Illumina MiSeq.
[0193] E. The sequencing data was analyzed using Crispresso2 software to obtain specific homologous recombination repair (HDR) efficiency, non-homologous end joining (NHEJ) efficiency, etc.
[0194] Alternatively, cells containing fluorescent expression can be used to measure gene editing efficiency using flow cytometry.
[0195] Example 1: Implementation of the gene editing system based on three base substitutions of this method:
[0196] Step 1: synthesize sgRNA-donor DNA chimera by short chain splicing.
[0197] The specific sequence of the sgRNA-donor DNA used is as follows (SEQ ID NO: 3):
[0198]
[0199] Among them, the bold sequence (double underline part + dotted underline part) is the sgRNA part, the italic is the connection part (i.e., the connection sequence), the single underline sequence at the front (5' end) is the left homology arm, the single underline sequence at the back (3' end) is the right homology arm, the three bases (GTG) in the middle of the homology arm are the replaced three bases, the double underline is the sequence for guiding the recognition of the target gene (i.e., the guide sequence), and the dotted underline is the skeleton (i.e., the scaffold sequence, the specific sequence is: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU; SEQ ID NO: 4). The sgRNA in the sgRNA-donor DNA chimera targets the eGFP gene, which will convert tyrosine at position 66 of the eGFP gene into histidine (Y66H), thereby converting the eGFP gene from expressing green fluorescent protein (eGFP) to expressing blue fluorescent protein (eBFP). The cells used are HEK293T cells (HEK293T-eGFP) in which the eGFP gene is inserted by lentivirus, which can be purchased directly.
[0200] Step 2: Incubate equimolar amounts of the above-mentioned sgRNA-donor DNA chimera and Cas9 protein (300 nM) in electroporation buffer at room temperature for 15 min (electrotransfer RNP, i.e., sgRNA-donor DNA chimera / Cas9 protein complex, wherein the Cas9 protein is a commercially purchased protein, NEB, catalog number M0646T).
[0201] Step 3: Mix the above incubation mixture with 5×10 4 After mixing the cells, use an electroporator according to the instructions of the Neon NxT electroporation system to transfect the electroporation mixture into the cells, culture them in a 24-well plate containing 1 mL of cell culture medium for 3-5 days, and replace the culture medium every 24 hours.
[0202] Step 4: After the above cultured cells were placed in PBS buffer, they were analyzed using a BD LSRFortessa multidimensional high-definition flow cytometer, with an excitation wavelength of 405 nm (eBFP signal) and an excitation wavelength of 488 nm (eGFP signal). The analysis results are as follows: Figure 3A As shown. FlowJo software was used to divide the flow cytometer analysis results, and the signal points with 405 excitations were counted and divided by the total signal points to obtain the ratio of eGFP to eBFP, that is, the HDR efficiency obtained by the flow cytometer. In the figure, NC is the blank control, no ssRDC or ssODN was added, only sgRNA was added for treatment, C3-20HA is the donor sequence code used, that is, the left and right homologous arms of the donor DNA are 20nt long and connected to the 3' end of the sgRNA, ssRDC refers to the sgRNA-donor DNA chimera used, ssODN is a single-stranded nucleic acid donor DNA, which is not connected to the sgRNA, and its sequence is consistent with the donor DNA of ssRDC. The quantitative data are as follows:
[0203]
[0204] Step 5: Extract genomic DNA from the cultured cells, construct a high-throughput sequencing library for the target sequence using the method described in the “Characterization of gene editing efficiency in transfected cells” section of the experimental method, and perform high-throughput sequencing using an Illumina MiSeq instrument. The sequencing results are as follows: Figure 3B As shown. ssRDC refers to the sgRNA-donor DNA chimera used, and ssODN is a single-stranded nucleic acid donor DNA that is not connected to sgRNA and has the same sequence as the donor DNA of ssRDC. The sequencing results were analyzed using CRISPResso2 software, the basic principle of which is to determine the target site sequence information of 2G data, analyze the number of correctly edited reads, and the ratio of the number of reads with insertions and deletions to the total number of reads to obtain the results. The quantitative data are as follows:
[0205]
[0206] The results showed that after flow cytometry characterization, it was found that the eGFP gene edited by the ssRDC / CRISPR system could successfully express BFP protein like the control group, and its editing efficiency was much higher than that of the control group. High-throughput sequencing also showed that the accurate editing efficiency of the ssRDC / CRISPR system was much higher than that of the control group, and it could significantly reduce the proportion of insertions and deletions.
[0207] Example 2: Determination of the concentration dependence of sgRNA-donor DNA in improving HDR efficiency
[0208] Step 1: synthesize sgRNA-donor DNA chimera by short chain splicing.
[0209] The sgRNA-donor DNA used is the same as the sgRNA-donor DNA used in Example 1. As mentioned above, the sgRNA in the sgRNA-donor DNA chimera targets the eGFP gene, which converts tyrosine 66 of the eGFP gene into histidine (Y66H), thereby converting the eGFP gene from expressing green fluorescent protein to expressing blue fluorescent protein (eBFP). The cells used are HEK293T cells (HEK293T-eGFP) in which the eGFP gene is inserted by lentivirus, which can be directly purchased.
[0210] Step 2: Add equimolar amounts of the above-mentioned sgRNA-donor DNA chimera and Cas9 protein in electroporation buffer at concentrations of 100 nM, 300 nM, 500 nM, 1000 nM, 1500 nM, 2000 nM, and 3000 nM, respectively; incubate at room temperature for 15 min (electrotransfer RNP, i.e., sgRNA-donor DNA chimera / Cas9 protein complex, Cas9 protein is a commercially purchased protein, NEB, catalog number M0646T).
[0211] Step 3: The above incubation mixture was mixed with 5×10 4 After mixing the cells, use an electroporator according to the instructions of the Neon NxT electroporation system to transfect the electroporation mixture into the cells, culture them in a 24-well plate containing 1 mL of cell culture medium for 3-5 days, and replace the culture medium every 24 hours.
[0212] Step 4: After the above cultured cells were placed in PBS buffer, they were analyzed using a BD LSRFortessa multidimensional high-definition flow cytometer, with an excitation wavelength of 405 nm (eBFP signal) and an excitation wavelength of 488 nm (eGFP signal). The analysis results are as follows: Figure 4As shown. The flow cytometer analysis results were divided into zones using FlowJo software, and the signal points with 405 excitations were counted and divided by the total signal points to obtain the ratio of eGFP to eBFP, which is the HDR efficiency obtained by the flow cytometer. In the figure, the data correspond to the substance concentration of the added ssRDC / Cas9, and the quantitative data are as follows:
[0213]
[0214] The results showed that the ssRDC / CRISPR system is significantly less dependent on high-concentration donors than the ssODN / CRISPR system. When the system concentration is only 500nM, the efficiency of the ssRDC / CRISPR system can reach more than 20%, and the efficiency of ssODN / CRISPR still tends to increase when the concentration reaches 3000nM, which can significantly reduce the amount of donors used in the editing system, thereby greatly reducing the cytotoxicity it brings.
[0215] Example 3: Implementation of the 390nt fragment insertion system based on this method:
[0216] Step 1: synthesize sgRNA-donor DNA chimera by short chain splicing.
[0217] The specific sequence of the sgRNA-donor DNA used is as follows (SEQ ID NO: 5):
[0218]
[0219] The bold sequence (double underlined part + dotted underlined part) is the sgRNA part, the italic part is the connection part (i.e., the connection sequence), the single underlined sequence at the front (5' end) is the left homology arm, the single underlined sequence at the back (3' end) is the right homology arm, and the 390 bases inserted into the N-terminus of the ACTB gene are in the middle of the homology arm. The double underlined sequence is the sequence that guides the recognition of the target gene (i.e., the guide sequence), and the dotted underline is the skeleton (i.e., the scaffold sequence); this ssRDC targets the ACTB gene in the human genome, and inserts a Y-FAST fluorescent protein at the N-terminus of the protein expressed by ACTB. This experiment uses HEK293T cells, which can be purchased directly.
[0220] Step 2: Incubate equimolar amounts of the above-mentioned sgRNA-donor DNA chimera and Cas9 protein (300 nM) in electroporation buffer at room temperature for 15 min (electrotransfer RNP, i.e., sgRNA-donor DNA chimera / Cas9 protein complex, wherein the Cas9 protein is a commercially purchased protein, NEB, catalog number M0646T).
[0221] Step 3: Mix the above incubation mixture with 5×10 4 After mixing the cells, use an electroporator according to the instructions of the Neon NxT electroporation system to transfect the electroporation mixture into the cells, culture them in a 24-well plate containing 1 mL of cell culture medium for 3-5 days, and replace the culture medium every 24 hours.
[0222] Step 4: Take the above cultured cells and place them in PBS buffer, and use BD LSRFortessa multidimensional high-definition flow cytometer to analyze them, select 488nm excitation wavelength (Y-FAST fluorescent protein signal), and the analysis results are as follows: Figure 5 As shown. The flow cytometer analysis results were zoned using FlowJo software, and the signal points with 405 excitations were counted and divided by the total signal points to obtain the proportion of successful and correct insertion of the Y-FAST fluorescent protein gene, that is, the HDR efficiency obtained by the flow cytometer. In the figure, NC is the blank control, in which no ssRDC or ssODN is added, and only sgRNA is added for treatment; ssRDC refers to the sgRNA-donor DNA chimera used, and ssODN is a single-stranded nucleic acid donor DNA, whose sequence is consistent with the donor DNA of ssRDC. Figure 5 In the figure, the left side is the flow cytometry analysis chart after ssODN treatment, the middle is the flow cytometry analysis chart after ssRDC treatment, and the right side is the quantitative analysis picture, and the quantitative data are as follows:
[0223] NG ssODN ssRDC 0.013 0.59 11.5 0 0.75 11.5 0 0.47 12.4
[0224] The results showed that the ssRDC / CRISPR system successfully inserted a 400nt Y-FAST gene sequence into the cell genome and expressed the corresponding Y-FAST protein, which means that the system can be used not only for editing single bases or shorter fragments, but also for inserting fragments of at least 400nt in length, and the insertion efficiency is much higher than that of the control group and is not limited by the length of the editing sequence.
[0225] It should be noted that, although the technical solutions of the present disclosure are introduced with specific examples, those skilled in the art will appreciate that the present disclosure should not be limited thereto.
[0226] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A gene editing system, comprising: (i) a nucleic acid programmable DNA binding protein; and (ii) sgRNA-donor DNA chimera, in, The sgRNA-donor DNA chimera is a single-stranded structure comprising a connected sgRNA and a donor DNA.
2. The gene editing system according to claim 1, characterized in that: The donor DNA includes a 5' homology arm, an editing sequence and a 3' homology arm; Preferably, the length of the 5' homology arm and the 3' homology arm are both 10-100 nucleotides, preferably 20-80 nucleotides.
3. The gene editing system according to claim 1 or 2, characterized in that: The sgRNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera; or, the sgRNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera; Preferably, the sgRNA is located at the 5' end of the single-stranded sgRNA-donor DNA chimera, and the donor DNA is located at the 3' end of the single-stranded sgRNA-donor DNA chimera.
4. The gene editing system according to claim 3, characterized in that In the sgRNA-donor DNA chimera, a linker sequence is further included between the sgRNA and the donor DNA.
5. The gene editing system according to claim 4, characterized in that: The linker sequence comprises a length of 1 to 200 nucleotides; Preferably, the linker sequence comprises a length of 1 to 20 nucleotides; More preferably, the linker sequence consists of thymidine deoxyribonucleotides.
6. The gene editing system according to any one of claims 1 to 5, characterized in that: The sgRNA comprises a guide sequence and a scaffold sequence.
7. The gene editing system according to any one of claims 1 to 6, characterized in that: The nucleic acid programmable DNA binding protein has nuclease activity; Preferably, the nucleic acid programmable DNA binding protein comprises a Cas protein; More preferably, the nucleic acid programmable DNA binding protein comprises a Cas9 protein.
8. An isolated polynucleotide, wherein The polynucleotide encodes the gene editing system as described in any one of claims 1 to 7. A vector comprising the polynucleotide according to claim 8.
10. A cell comprising one or more of the following: (a) The gene editing system according to any one of claims 1 to 7; (b) the polynucleotide of claim 8; and, (c) The vector according to claim 9.
11. A pharmaceutical composition comprising one or more of the following (A) to (D): (A) The gene editing system according to any one of claims 1 to 7; (B) the polynucleotide according to claim 8; (C) the vector according to claim 9; and (D) the cell according to claim 10; And, optionally, a pharmaceutically acceptable carrier.
12. A kit comprising one or more of the following: (A) The gene editing system according to any one of claims 1 to 7; (B) the polynucleotide according to claim 8; (C) the vector according to claim 9; and (D) The cell according to claim 10.
13. A method for editing a target gene, the method comprising the step of contacting the target gene with the gene editing system according to any one of claims 1 to 7, the polynucleotide according to claim 8, the vector according to claim 9, the cell according to claim 10, the pharmaceutical composition according to claim 11, or the kit according to claim 12; Preferably, the editing includes at least one of replacement, insertion and deletion.
14. Use of the gene editing system according to any one of claims 1 to 7, the polynucleotide according to claim 8, the vector according to claim 9 or the cell according to claim 10 in preparing a reagent for editing a target gene; Preferably, the editing includes at least one of replacement, insertion and deletion.
15. A method for preventing and / or treating diseases caused by disease-related genes, comprising the step of administering to a subject a therapeutically effective amount of the gene editing system as described in any one of claims 1 to 7, the polynucleotide as described in claim 8, the vector as described in claim 9, the cell as described in claim 10, the pharmaceutical composition as described in claim 11, or the kit as described in claim 12.
16. Use of the gene editing system according to any one of claims 1 to 7, the polynucleotide according to claim 8, the vector according to claim 9, the cell according to claim 10, the pharmaceutical composition according to claim 11 or the kit according to claim 12 in the preparation of drugs for preventing and / or treating diseases caused by disease-related genes.
Citation Information
Cited By
Chimera reverse transcription primer and application thereof in precise gene editing
CN120683098A