Use of a gene editing system in the preparation of a gene editing product
By using the Cas9 nucleases iSpymacCas9 and sgRNA in the gene editing system, the sequence dependence of preferred insertion or removal of wild-type Cas9 nucleases during gene editing is reversed, and the limitations of gene editing accuracy and flexibility in the prior art are solved, achieving a wider sgRNA sequence-guided precise base insertion.
Patent Information
- Application Number
- CN202410236230.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-03-01
AI Technical Summary
The existing insertion or deletion preferences during gene editing are affected by the -4th base type of PAM, limiting the accuracy and flexibility of gene editing.
A gene editing system containing the Cas9 nuclease iSpymacCas9 and sgRNA was used. This system reversed the insertion or deletion preferences affected by the base type of position-4 of PAM compared to the gene editing system using the wild-type Cas9 nuclease.
By changing the characteristics of Cas9 nuclease, iSpymacCas9 prefers to produce base deletion when position -4 of PAM is A/T, while base insertion is more likely to produce base insertion when position -4 of PAM is C/G, which improves the accuracy and flexibility of gene editing.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and particularly to the use of a gene editing system in the preparation of gene editing products. Background Art
[0002] For a long time, genetic manipulation of genomes has been a hot topic of concern (1-6). Traditional gene editing was based on transposon and transgenic technologies (7), but due to their large limitations, gene editing tools including ZFNs and TALENs were successively developed (8). As the first-generation gene editing system, ZFNs rely on the fusion of zinc finger modules that recognize specific sequences. By fusing zinc finger modules that recognize different sequences, cleavage of specific target sequences can be achieved (9). However, since a single zinc finger module recognizes three bases and is difficult to construct, there are certain limitations. Therefore, to achieve more extensive gene editing, the second-generation gene editing system TALENs was developed (10). TALENs are essentially similar to ZFNs and also rely on protein recognition of target sequences. Tandem arrangement of TALE motifs can achieve recognition of target sequences. Since a single motif is responsible for the recognition of a single base pair, the design is more flexible, but it is also more difficult to assemble than ZFNs because it needs to contain more motifs (11, 12).
[0003] In recent years, to achieve more efficient gene editing, the third-generation gene editing system CRISPR / Cas9 [Clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR-associated nuclease 9 (Cas9)] has been developed (11, 13-16). This system is derived from the immune defense systems of bacteria and archaea. After the initial invasion of foreign nucleic acid fragments, bacteria and archaea will insert the characteristic sequences of the virus into the CRISPR system. When the virus invades again, the system can respond quickly, recognize and cleave the target sequence to achieve immunity. Since this system relies on RNA to recognize target sequences, its construction is simpler than traditional gene editing systems, greatly improving the efficiency of gene editing.
[0004] After the Cas9 nuclease generates DSB (Double-strand break) by cleaving the target sequence under the mediation of sgRNA (single guide RNA), it is mainly repaired in cells by cellular repair mechanisms such as NHEJ (non-homologous end joining) and HR (homologous recombination) (17). Although NHEJ is highly efficient and not restricted by the cell cycle, it has not been used for precise gene editing because it generates small insertions or deletions (indels) at the broken ends (18) (19). In 2018, Shou Jia et al. found that among the repair results of gene editing products mediated by CRISPR / Cas9, small insertions are more precise than deletions. In particular, single-base insertions usually show the recurrence of the base at position -4 upstream of the PAM site (20). Similarly, in 2019, Shi Xin et al. also reported that there is an insertion preference for two-base insertions, which are more inclined to insert bases at positions -5 and -4 upstream of the PAM. For this insertion preference, they speculated that it originated from the sticky-end cleavage activity of Cas9 itself. When the RuvC domain of Cas9 cleaves at positions -4 or even -5 and -6 upstream of the PAM on the non-target strand to generate a 5'-protrusion of single or multiple bases, predictable base insertions can be generated under the action of polymerase (21).
[0005] The human genetic variation database records thousands of diseases caused by frameshift mutations. Although the cleavage activity of Cas9 itself can be used to guide base insertions to repair gene function inactivation caused by frameshift mutations to a certain extent, due to the sequence dependence of this cleavage activity, the scope of application is limited. Summary of the Invention
[0006] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide the use of a gene editing system in the preparation of gene editing products to solve the problems in the prior art.
[0007] To achieve the above object and other related objects, the present invention provides the use of a gene editing system in the preparation of gene editing products. The gene editing system comprises Cas9 nuclease iSpymacCas9 and sgRNA. Compared with the gene editing system using wild-type Cas9 nuclease, the insertion or deletion preference affected by the base type at position -4 of the PAM is reversed.
[0008] Preferably, the Cas9 nuclease iSpymacCas9 comprises a polypeptide with an amino acid sequence as shown in SEQ ID No.1; and / or, the sgRNA comprises a backbone sequence with a nucleotide sequence as shown in SEQ ID No.2.
[0009] The present invention also provides a composition for gene editing, which composition comprises the aforementioned Cas9 nuclease iSpymacCas9 and sgRNA; the sgRNA comprises a backbone sequence with a nucleotide sequence as shown in SEQ ID No.2, and compared with the gene editing system using wild-type Cas9 nuclease, the insertion or deletion preference affected by the -4th base type of PAM is reversed.
[0010] The present invention also provides a gene editing method for editing genes of isolated cells, which method is a method for reversing the editing type preference of wild-type Cas9 affected by the -4th base of PAM, and the gene editing method is: introducing the aforementioned composition into the isolated cells to achieve gene editing of the isolated cells.
[0011] As described above, the use of a gene editing system of the present invention in the preparation of gene editing products has the following beneficial effects: Compared with wild-type Cas9 nuclease, the sequence-dependent characteristics of the base insertion preference of Cas9 nuclease (iSpymacCas9) in the present invention are different. Although both wild-type Cas9 and the Cas9 nuclease (iSpymacCas9) involved in the present invention can cleave at the target sequence to generate a 5' overhang and can generate precise base insertions under the mediation of a polymerase, the frequencies of base insertions mediated by the two are different. Further, when the -4th base of PAM or the 17th base of sgRNA is A / T, compared with the gene editing mediated by wild-type Cas9, the probability of base deletion mediated by Cas9 nuclease iSpymacCas9 is greater; and when the -4th base of PAM or the 17th base of sgRNA is C / G, compared with the gene editing mediated by wild-type Cas9, the probability of base insertion mediated by iSpymacCas9 is greater. In other words, although the cleavage activities of the two are similar, both can generate 5' overhang ends and direct precise base insertions, but the sgRNA sequences on which base insertions depend are different for the two, resulting in nearly opposite results, that is, different sequence-dependent preferences, and more precise base insertions guided by a wider range of sgRNA sequences can be achieved. Therefore, it can be expected that using these two nucleases can provide more possibilities and available tools for precise base editing mediated by base insertions. Brief Description of the Drawings
[0012] Figure 1 It shows a schematic diagram of the domains of wild-type Cas9 nuclease, SpymacCas9 nuclease and iSpymacCas9 nuclease in the present invention.
[0013] Figure 2 It shows a schematic diagram of the results that when performing single-site editing at the CTCF locus in the present invention, using an optimized sgRNA backbone sequence has higher editing activity.
[0014] Figure 3 Shown is a schematic diagram of the results of expressing Cas9 nuclease by wild-type Cas9 nuclease and iSpymacCas9 nuclease in the present invention.
[0015] Figure 4 Shown is the relative editing result after editing at the RNF112 locus in the present invention.
[0016] Figure 5 Shown is the relative editing result after editing at the TUBGLP5 locus in the present invention.
[0017] Figure 6 Shown is the relative editing result after editing at the NAXD locus in the present invention.
[0018] Figure 7 Shown is the relative editing result after editing at the CNNM4 locus in the present invention.
[0019] Figure 8 Shown is the relative editing result after editing at the RBBP8 locus in the present invention.
[0020] Figure 9 Shown is the relative editing result after editing at the EXD1 locus in the present invention.
[0021] Figure 10 Shown is the relative editing result after editing at the RP11-568G13.3 locus in the present invention.
[0022] Figure 11 Shown is the relative editing result after editing at the CA2M2 locus in the present invention.
[0023] Figure 12 Shown is a schematic diagram of the results of the difference in the base insertion frequency of two Cas9 nucleases after editing at the EXD1 locus in the present invention.
[0024] Figure 13 Shown is a schematic diagram of the results of the difference in the base insertion frequency of two Cas9 nucleases after editing at the RP11-568G13.3 locus in the present invention.
[0025] Figure 14 Shown is a schematic diagram of the results of the difference in the base insertion frequency of two Cas9 nucleases after editing at the CA2M2 locus in the present invention.
[0026] Figure 15 Shown is a schematic diagram of the results of the motif analysis of the sgRNA sequence where wild-type Cas9 in the present invention is more inclined to produce base insertions.
[0027] Figure 16Schematic diagram showing the analysis results of the sgRNA sequence motif for which iSpymacCas9 in the present invention is more inclined to generate base insertions. Detailed implementation manners
[0028] The present invention provides the use of a gene editing system in the preparation of gene editing products, wherein the gene editing system comprises the Cas9 nuclease iSpymacCas9 and sgRNA, and for the gene editing products, compared with the gene editing system using the wild-type Cas9 nuclease, the insertion or deletion preference affected by the base type at the -4th position of the PAM is reversed.
[0029] In some detailed implementation manners, the reversal of the insertion or deletion preference affected by the base type at the -4th position of the PAM is as follows: when the -4th position of the PAM is A / T, the gene editing type mediated by wild-type Cas9 is changed from a high base insertion probability to a high base deletion probability; and / or when the -4th position of the PAM is C / G, the gene editing type mediated by wild-type Cas9 is changed from a high base deletion probability to a high base insertion probability. In some detailed implementation manners, the high base insertion probability or high base deletion probability is an event with an occurrence ratio of at least greater than 60%.
[0030] In some detailed implementation manners, the Cas9 nuclease iSpymacCas9 comprises a polypeptide with an amino acid sequence as shown in SEQ ID No.1; and / or the sgRNA comprises a backbone sequence with a nucleotide sequence as shown in SEQ ID No.2.
[0031] In some detailed implementation manners, the gene editing products are products for gene editing at the -4th position of the protospacer adjacent motif (PAM) or the 17th position of the sgRNA sequence. Among them, the -4th position is represented as: taking the 5'-end to 3'-end of the nucleic acid fragment as the positive direction, moving 4 nucleotides along the negative direction from the 5'-end of the PAM is the -4th position; the 17th position is represented as: taking the 5'-end to 3'-end of the nucleic acid fragment as the positive direction, moving 17 nucleotides along the positive direction from the 5'-end of the sgRNA is the 17th position.
[0032] Furthermore, the gene editing products are nucleotide insertion products for which the base at the -4th position of the PAM or the 17th position of the sgRNA sequence is C / G; and / or the gene editing products are nucleotide deletion products for which the base at the -4th position of the PAM or the 17th position of the sgRNA sequence is A / T.
[0033] The present invention also provides a composition for gene editing, which composition comprises the aforementioned Cas9 nuclease iSpymacCas9 and sgRNA. The sgRNA comprises a backbone sequence with a nucleotide sequence as shown in SEQ ID No. 2. Compared with a gene editing system using a wild-type Cas9 nuclease, the insertion or deletion preference affected by the type of the -4th base of the PAM is reversed.
[0034] In some specific embodiments, the composition further comprises a pharmaceutically acceptable carrier and medium. The pharmaceutically acceptable carrier and medium may include, for example, sterile water or physiological saline, stabilizers, excipients, antioxidants (such as ascorbic acid), buffers (such as phosphoric acid, citric acid, and other organic acids), preservatives, surfactants (such as PEG, Tween, etc.), chelating agents (such as EDTA, etc.), adhesives, etc. Moreover, it may also contain other low-molecular-weight polypeptides; proteins such as serum albumin, gelatin, or immunoglobulins; amino acids such as glycine, glutamine, asparagine, arginine, and lysine; saccharides or carbohydrates such as polysaccharides and monosaccharides; sugar alcohols such as mannitol or sorbitol. When preparing an aqueous solution for injection, such as physiological saline, an isotonic solution containing glucose or other auxiliary drugs, such as D-sorbitol, D-mannose, D-mannitol, sodium chloride, appropriate solubilizers such as alcohols (such as ethanol), polyols (such as propylene glycol, PEG, etc.), non-ionic surfactants (such as Tween 80, HCO-50), etc. may be used. In some embodiments, the composition comprises a buffer for stabilizing nucleic acids.
[0035] The present invention also provides a gene editing method for disease diagnosis or treatment purposes, or non-disease diagnosis or treatment purposes, or for editing the genes of isolated cells, which method is a method in which, compared with a gene editing system using a wild-type Cas9 nuclease, the insertion or deletion preference affected by the type of the -4th base of the PAM is reversed. The gene editing method is: introducing the aforementioned composition into isolated cells to achieve gene editing of the isolated cells. Among them, the gene editing method for in vitro or non-disease diagnosis or treatment purposes can be a gene editing method for industrial production purposes or for scientific research purposes.
[0036] In some specific embodiments, the gene editing method is a method for nucleotide insertion when the -4th base of the PAM or the 17th base of the sgRNA sequence is C / G; and / or, the gene editing method is a method for nucleotide deletion when the -4th base of the PAM or the 17th base of the sgRNA sequence is A / T.
[0037] In some specific embodiments, the isolated cells are selected from prokaryotic cells or eukaryotic cells, and the eukaryotic cells do not include fertilized egg cells. Specifically, the prokaryotic cells may be selected from one or more of Escherichia coli, Salmonella typhimurium, Klebsiella pneumoniae, Bacteroides ovatus, Campylobacter jejuni, Staphylococcus saprophyticus, Enterococcus faecalis, Bacteroides thetaiotaomicron, Bacteroides vulgatus, Bacteroides uniformis, Lactobacillus casei, Bacteroides fragilis, Acinetobacter lwoffii, Fusobacterium nucleatum, Bacteroides johnsonii, Bacteroides arabinosus, Lactobacillus rhamnosus, Bacteroides massiliensis, Parabacteroides distasonis, Clostridium scindens, Bifidobacterium breve, or Listeria monocytogenes; the eukaryotic cells may be selected from one or more of yeast cells, insect cells of Drosophila S2 or Sf9, CHO cells, COS cells, 293 cells, Bowes melanoma cells, 293T cells, T lymphocytes, B lymphocytes, N2A cells, Hela cells, white blood cells, or pluripotent stem cells.
[0038] In some specific embodiments, the introduction methods may be selected from one or more of DEAE-dextran-mediated transfection, liposome-mediated transfection, virus or phage infection, transfection, conjugation, protoplast fusion, polyethyleneimine-mediated transfection, electroporation, gene gun, calcium phosphate precipitation, microinjection, and nanoparticle-mediated nucleic acid delivery.
[0039] In the present invention, the terms "nucleic acid fragment" and "polynucleotide" are used interchangeably, and they refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides or ribonucleotides or their analogs. Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. Examples of polynucleotides include, but are not limited to, the following: genes or gene fragments (including probes, primers, ESTs or SAGE tags), exons, introns, messenger RNA, transfer RNA, ribosomal RNA, ribozymes, cDNA, dsRNA, siRNA, miRNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides also include modified nucleotides, such as methylated nucleotides and nucleotide analogs. If there are modifications on the polynucleotide, the modifications can be conferred before or after the assembly of the polynucleotide. The nucleotide sequence can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, for example, by being labeled with a labeled component through conjugation. The term refers to both double-stranded and single-stranded polynucleotide molecules. Unless otherwise specified or required, any embodiment of the polynucleotides disclosed in the present invention includes its double-stranded form and either of the two complementary single-stranded forms that are known or predicted to form a double-stranded form.
[0040] In the present invention, the terms "peptide", "polypeptide" and "protein" are used interchangeably and denote a polymeric form of amino acids of any length, which may include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones.
[0041] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be applied based on different viewpoints and various modifications or changes can be made without departing from the spirit of the present invention.
[0042] Before further describing the specific embodiments of the present invention, it should be understood that the protection scope of the present invention is not limited to the specific embodiments described below; it should also be understood that the terms used in the embodiments of the present invention are for describing specific embodiments and not for limiting the protection scope of the present invention; in the specification and claims of the present invention, unless otherwise clearly indicated in the text, the singular forms "a", "an" and "the" include plural forms.
[0043] When the embodiments give a numerical range, it should be understood that unless otherwise specified in the present invention, any value between the two endpoints of each numerical range and either of the two endpoints can be selected. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art of this technology. In addition to the specific methods, devices and materials used in the embodiments, according to the knowledge of the prior art by those skilled in the art of this technology and the description of the present invention, any methods, devices and materials similar or equivalent to those described in the embodiments of the present invention can also be used to implement the present invention.
[0044] Unless otherwise specified, the experimental methods, detection methods, and preparation methods disclosed in the present invention all adopt conventional techniques in the fields of molecular biology, biochemistry, chromatin structure and analysis, analytical chemistry, cell culture, recombinant DNA technology, and related fields in the art. These techniques have been well described in the existing literature. For details, see Sambrook et al., MOLECULAR CLONING: A LABORATORY MANUAL, Second edition, Cold Spring Harbor Laboratory Press, 1989 and Third edition, 2001; Ausubel et al., CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, New York, 1987 and periodic updates; the series METHODS IN ENZYMOLOGY, Academic Press, San Diego; Wolffe, CHROMATIN STRUCTURE AND FUNCTION, Third edition, Academic Press, San Diego, 1998; METHODS IN ENZYMOLOGY, Vol. 304, Chromatin (P.M. Wassarman and A.P. Wolffe, eds.), Academic Press, San Diego, 1999; and METHODS IN MOLECULAR BIOLOGY, Vol. 119, Chromatin Protocols (P.B. Becker, ed.) Humana Press, Totowa, 1999, etc.
[0045] The sequence information used in the present invention is as follows:
[0046] SEQ ID No.1
[0047]
[0048] SEQ ID No.2
[0049] Gtttcagagctatgctggaaacagcatagcaagttgaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgc (where base t and base u can be replaced with each other)
[0050] SEQ ID No.3
[0051]
[0052] Example: The nuclease iSpymacCas9 has a new sgRNA sequence feature for precise base insertion.
[0053] 1. Construction of the SpymacCas9 hybrid protein
[0054] (1) Gene synthesis
[0055] Use the codon optimization software provided by the company to optimize the codons of the PAM recognition domain from Streptococcus macacae. After synthesizing the corresponding nucleotide sequence by the company, the synthesized sequence was ligated to the pUC19 vector through cloning technology. The ligation product was verified to have the correct sequence by first-generation sequencing.
[0056] (2) Design of primer pair PID for amplification
[0057] Use Q5 (NEB) to amplify PID with the plasmid as the template. The reaction system is as follows:
[0058]
[0059]
[0060] Reaction conditions:
[0061]
[0062] Purify the amplification product using a PCR purification kit (Roche). The purification steps strictly follow the instructions.
[0063] (2) Construction of the SpymacCas9 nuclease expression plasmid
[0064] Based on the existing wild-type SpyCas9 expression vector in the laboratory, replace PID with that from Streptococcus macacae to obtain the SpymacCas9 nuclease expression plasmid. Select appropriate enzymes to linearize the vector. The reaction system is as follows:
[0065]
[0066] Purify the enzyme digestion product using gel extraction and purification according to the instructions of the gel extraction kit (Axygen).
[0067] Use homologous recombination to load the purified amplification product onto the linearized vector. The reaction system is as follows:
[0068]
[0069] After reacting at 50 °C for 30 min, transform the ligation product. The transformation steps are as follows:
[0070] After transforming the ligation product with Stbl3 competent cells, culture it overnight on an LB plate containing ampicillin antibiotic (Amp, 100 mg / L) at a culture temperature of 37 °C. After colonies grow, pick a single colony and transfer it to an LB (Amp, 100 mg / L) liquid medium for overnight culture. Then extract the plasmid according to the instructions of the plasmid miniprep kit (Axygen) and perform first-generation sequencing verification.
[0071] 2. Construction of the iSpymacCas9 mutant hybrid protein
[0072] On the basis of SpymacCas9, mutate arginine 221 and asparagine 394 to lysine to obtain the iSpymacCas9 mutant hybrid protein. The specific steps are as follows:
[0073] (1) Use the NEB mutagenesis kit (Q5 Site-Directed Mutagenesis Kit, #E0554S) to construct the Cas9 mutant. First, purchase primers for PCR amplification. The primer sequences are as follows:
[0074] ispy-enzyF: AGGGGTCGGCAATTGAACCGGT (SEQ ID NO.4);
[0075] isp-R221F: CTTCCGGGATTTGGACAGCCTA (SEQ ID NO.5);
[0076] isp-R221F: TCCAAATCCCGGAAGCTCGAAAACCTCA (SEQ ID NO.6);
[0077] isp-N394R: CTTAAGCTTTACCAGCAGCTCCTC (SEQ ID NO.7);
[0078] isp-N394F: GCTGGTAAAGCTTAAGAGAGAAGATCTGTTG (SEQ ID NO.8);
[0079] ispy-enzyR: CACTCTGCTTGTCTCGGATCC (SEQ ID NO.9);
[0080] The reaction system is as follows:
[0081]
[0082] Reaction conditions:
[0083]
[0084] (2) Treatment with KLD (Kinase, Ligase & DpnI), the reaction system is as follows:
[0085]
[0086] Reaction conditions: room temperature for 10 minutes
[0087] (3) All of the reaction products in (2) were used for the transformation of competent bacteria Stbl3 (50 μl), cultured overnight on an LB plate containing ampicillin antibiotic (Amp, 100 mg / L) at 37°C. Single colonies were picked and transferred to a liquid medium, and the plasmid was extracted and then subjected to first-generation sequencing. Plasmids containing the nucleic acid fragment shown in SEQ ID No. 3 were selected. If the sequencing was successful, the successfully sequenced plasmid could be transformed again for mid-prep. The specific steps are as follows:
[0088] The successfully sequenced plasmid was re-transformed with Stbl3 competent cells, cultured overnight on an LB plate containing Amp (100 mg / L), and in the morning, single colonies were picked and cultured in 2 ml of LB (Amp, 100 mg / L) liquid medium for 8 hours and then transferred to 200 ml of LB (Amp, 100 mg / L) liquid medium and cultured overnight; the cells were collected at 4,000 g for 15 min at 4°C, and the plasmid was extracted according to the instructions of the plasmid mid-prep kit (Qiagen).
[0089] 3. Designing sgRNA sequences and their corresponding target sequences using bioinformatics means
[0090] To comprehensively study the cleavage and repair mechanisms of Cas9 nuclease (iSpymacCas9), a large-capacity sgRNA library was designed and constructed. The constructed sgRNA library contains 37,266 sgRNAs and their corresponding target sequences, covering almost all coding genes, with a balanced and uniform base distribution. The designed library was handed over to Custom Array Company for chip synthesis using an electrochemical method.
[0091] 4. Obtaining an expression vector with an optimized sgRNA backbone sequence using site-directed mutagenesis
[0092] The NEB mutagenesis kit (Q5 Site-Directed Mutagenesis Kit, #E0554S) was used to construct an expression vector with an optimized sgRNA backbone.
[0093] (1) First, primers were purchased for PCR amplification, and the primer sequences are as follows:
[0094] gRNA-sca-F1: AGAGATCCAGTTTGGTTAATTAAGGTACCGAGGGCCT (SEQ ID NO.10);
[0095] gRNA-sca-R1: GCTGTTTCCAGCATAGCTCTGAAACAGAGACGTACAAAAAAGAGCAAGA (SEQ ID NO.11);
[0096] gRNA-sca-F2: TATGCTGGAAACAGCATAGCAAGTTGAAATAAGGCTAGTCC (SEQ ID NO.12); gRNA-sca-R2: GAGGCTGATCAGCGGGTTTAAACGGGCCCTGC (SEQ ID NO.13);
[0097] The reaction conditions are as follows:
[0098]
[0099] Reaction conditions:
[0100]
[0101] (2) Treatment with KLD (Kinase, Ligase & DpnI), the reaction is as follows:
[0102]
[0103] Reaction conditions: 10 minutes at room temperature.
[0104] 5. Construction of the first-step sgRNA library
[0105] The method for preparing the sgRNA library refers to Nature Protocol (22), including Gibson Assembly and Golden Gate. After the synthesized sequences are amplified by PCR and loaded onto the U6 promoter of the vector using Gibson Assembly, the first-step sgRNA library is obtained.
[0106] (1) Determination of the minimum number of amplification cycles using qPCR
[0107] To avoid random mutations introduced by over-amplification and resulting library imbalance, qPCR is used to determine the minimum number of amplification cycles. The specific steps are as follows:
[0108]
[0109] Refer to the instruction manual for the amplification conditions, change the annealing temperature to 67 °C, and extend the extension time to 1 min. The optimal number of amplification cycles is half of the number of cycles at the plateau phase. Finally, 20 cycles are selected as the optimal condition.
[0110] (2) Amplify the synthetic library using the determined number of cycles
[0111] Prepare the following reaction system:
[0112]
[0113]
[0114] The reaction conditions are as follows:
[0115]
[0116] Purify the amplified product using a PCR purification kit (Roche), and strictly follow the instructions for the steps.
[0117] (3) Digest the 52961-cas-del-scaffold-del vector
[0118] Linearize the 52961-cas-del-scaffold-del vector and use it for the assembly of synthetic fragments.
[0119] Prepare the following reaction system:
[0120]
[0121] After digesting at 55 °C for 9 hours, perform gel extraction and purification (Qiagen) on the digested product, and follow the instructions for the purification steps.
[0122] (4) Homologous recombination
[0123] Assemble the amplified synthetic library sequence onto the vector using homologous recombination, and the reaction conditions are as follows:
[0124]
[0125] React at 50 °C for 2 h.
[0126] (5) Electroporation
[0127] Transform the homologous recombination product into electrocompetent cells (EnduraTM Electroporationcompetent cells) using electroporation, and strictly follow the electroporation conditions of the Bio-Rad electroporator. Culture overnight at 32 °C on an LB plate containing ampicillin antibiotic (Amp, 100 mg / L).
[0128] (6) Collect the bacterial cells for large-scale plasmid extraction
[0129] After collecting the bacterial cells with a spreading rod, centrifuge at 4,000 g for 15 min at 4 °C to collect the bacteria, and extract the plasmid according to the instructions of the plasmid large-scale extraction kit (Qiagen).
[0130] 6. Insert the sgRNA backbone sequence into the constructed first-step sgRNA library to form the second-step sgRNA library
[0131] Due to the limitation of the synthesis length, the synthetic sequence of the chip does not contain the sgRNA backbone sequence. The sgRNA backbone sequence needs to be inserted into a specific position of the sgRNA sequence to achieve the normal expression of sgRNA. After amplifying and digesting the vector containing the optimized sgRNA backbone sequence constructed in 4, the backbone sequence can be inserted through Golden Gate ligation.
[0132] The specific steps are as follows:
[0133] (1) Linearization of the first-step sgRNA library
[0134] Prepare the following reaction system:
[0135]
[0136] After reacting at 37 °C for 4 hours, inactivate the enzyme by reacting at 65 °C for 15 minutes.
[0137] (2) Dephosphorylation of the vector
[0138] To avoid self-ligation of the vector, the digested vector needs to be dephosphorylated. Prepare the following reaction system:
[0139] Product from the last step 50 μl
[0140] shrimp Alkaline phosphatase (rSAP) 2.89 μl
[0141] React the above system at 37 °C for 30 min and then inactivate the enzyme by reacting at 65 °C for 5 min. Then, recover the linearized vector using the Axygen gel extraction kit, and the steps of gel extraction follow the instructions.
[0142] (3) PCR amplification of the optimized sgRNA backbone sequence
[0143] Prepare the following reaction system:
[0144]
[0145]
[0146] (4) Digest the amplified backbone sequence with restriction enzymes
[0147] Prepare the following reaction system:
[0148]
[0149] After digesting at 55 °C for 2 h 20 min, purify the digested product using a PCR purification kit (Qiagen).
[0150] (5) Ligation of the backbone sequence
[0151] Prepare the following reaction system:
[0152]
[0153] After ligation overnight at 16 °C, purify the ligation product using 1.8x beads (Vazyme).
[0154] (6) Transform the ligation product into electrocompetent cells (EnduraTM Electroporationcompetent cells) by electroporation
[0155] The transformation steps strictly follow the electroporation conditions of the Bio-Rad electroporator. Then, culture overnight at 32 °C on an LB plate containing ampicillin antibiotic (Amp, 100 mg / L).
[0156] (7) Collect the bacterial cells for large-scale plasmid extraction
[0157] After collecting the bacterial cells with a spreader, centrifuge at 4,000 g for 15 min at 4 °C to collect the bacteria, and extract the plasmid according to the instructions of the plasmid large-scale extraction kit (Qiagen).
[0158] 7. Lentivirus packaging, concentration and titer determination
[0159] In this study, to introduce Cas9 and the sgRNA library into target cells and achieve stable and long-term expression, the obtained plasmid needs to be packaged into lentivirus. Stable integration can be achieved by infecting the target cells with the packaged virus particles. For safety reasons, the components for producing the virus are divided into three plasmids, namely the packaging plasmid, the envelope plasmid and the plasmid containing the fragment to be integrated. Co-transfecting them into the packaging cell line HEK293T can produce the virus for cell infection. The specific steps are as follows (23):
[0160] (1) Lentivirus packaging
[0161] All the lentiviruses involved in this study were prepared by similar methods. Here, only SpyCas9 is described in detail as an example. For virus packaging, HEK293T cells with a lower passage number and good growth vitality were used and packaged using a three-plasmid system. One day in advance, HEK293T cells were inoculated into a 10-cm cell culture dish at a density of 3×10^5 / ml to ensure that the cell confluence was about 70% at the time of transfection the next day. After diluting lipofectamine 3000 with OPTI-MEM, it was placed at room temperature for 5 minutes; the plasmids to be packaged, SpyCas9, psPAX2, and pMD2.G, were also diluted in OPTI-MEM at a ratio of 10:7.5:5, and then P3000 was added and mixed well. The DNA was added to the diluted lipofectamine 3000, incubated at room temperature for 15 minutes, and then dropped onto the cells, and the cells were cultured in an incubator at 37°C with 5% CO2. After 12 - 18 hours, the medium was changed to a harvest medium containing 30% fetal bovine serum. The supernatants containing virus particles were harvested at 24 hours, 36 hours, and 48 hours of culture respectively, and the crude virus solutions collected at different time points were mixed together. The virus solution can be temporarily stored at 4°C when collected.
[0162] (2) Concentration of Lentivirus
[0163] To remove toxins as much as possible, ultracentrifugation was selected here to concentrate the virus solution. At the same time, to ensure the viability of the virus, buffering was required during centrifugation. Here, a sucrose cushion formulation is provided as follows:
[0164] Component Final Concentration Sucrose 20%(w / v) NaCl 100 mM HEPES (pH 7.4) 2 mM EDTA 1 mM
[0165] The crude virus solution collected in (1) can be stored at 4°C for one week after filtering through a 0.45-μm filter membrane. Here, the crude solution was further concentrated, and the specific concentration steps are as follows:
[0166] The crude virus solutions collected at different time points were centrifuged at 1,000 g for 10 min at 4°C to remove cell debris. After transferring the supernatant to a new centrifuge tube, it was filtered through a 0.45-μm filter membrane again to remove all cell debris as much as possible. The virus solution at this time can be stored at 4°C for no more than one week. Then, the filtered virus solution can be loaded into a centrifuge tube that has been disinfected by wiping with alcohol and ultraviolet irradiation. To protect the virus activity before adding the virus solution, 6 ml of 20% sucrose solution should be carefully added to the bottom of the centrifuge tube. The liquid in the tube is generally added to a stop 1 - 3 cm from the tube mouth. If there is still virus solution, repeat the above steps, but pay attention to careful balancing (<0.1 g).
[0167] After selecting the appropriate rotor and adapter, centrifuge at 25,000 r.p.m at 4 °C for 2 h. After centrifugation, carefully remove the centrifuge tube, discard the supernatant, and blot the remaining liquid with absorbent paper. Then, an appropriate amount of PBS can be added to dissolve overnight. Aliquot and store at -80 °C.
[0168] (3) Lentivirus titer determination
[0169] To more accurately determine the titer of the lentivirus, an infection method is adopted here. Viruses with infectivity are screened through antibiotics to evaluate the titer. One day in advance, HEK293T cells in good growth state are seeded in a six-well plate. When the cell confluence is about 70%, change the medium to DMEM complete medium containing polybrene at a final concentration of 8 mg / ml. Infect the cells with virus solutions of different dilution factors. At 24 hours after infection, change the medium to a resistant medium containing antibiotics for screening. After all the uninfected cells die, count the remaining cells and calculate the titer.
[0170] 8. Construction of monoclonal cell lines
[0171] Since the virus integrates randomly, even cells with successful integration may show silencing of expression and low protein levels. It is necessary to monoclonalize the population of infected cells. After plating the single-cell suspension obtained by the dilution method, identify the single cell clusters at the gene and protein levels. The specific steps are as follows:
[0172] One day in advance, HEK293T cells in good growth state are seeded in a six-well plate. When the cell confluence is about 70%, change the medium to DMEM complete medium containing polybrene at a final concentration of 8 μg / ml. According to the calculated titer, add an appropriate amount of virus concentrate for infection. At 24 hours after infection, change to normal medium, and then continue to culture for 12 - 24 hours and add a resistant medium containing blasticidin for screening. After the screening is completed, count the cells, dilute the cells to a concentration of 7.5 cells / ml, and use a multichannel pipette to seed the diluted cells in a 96-well plate, 200 μl per well. After culturing for 6 days, observe the cell growth under a microscope and change the medium for the wells with only one cell cluster. After culturing for a few more days, aspirate some cells for identification.
[0173] PCR and western blot are techniques well-known to researchers in this field, and the specific steps will not be elaborated. The correctly identified cell lines can be used for subsequent experiments.
[0174] 9. Infecting monoclonal cell lines with the sgRNA library
[0175] The sgRNA library was packaged, purified, concentrated, and titer-determined using the method similar to that mentioned in 7 above. The corresponding monoclonal cell lines were infected using the similar infection method mentioned in 8. The sgRNA library used in this study was large in quantity and high in complexity. To facilitate downstream analysis, it was necessary to ensure that at most one virus particle could enter a single cell. Therefore, a low MOI (multiplicity of infection) value was used to infect the cell line stably expressing Cas9. At the same time, to ensure the high coverage of the sgRNA library and minimize the loss of the library characteristics, multiple sets of repeated experiments were required.
[0176] 10. Enrichment of the editing results after infection with the sgRNA library
[0177] After a few days of antibiotic screening of the infected cells, the cells could be collected, and the genomic DNA was extracted using the Promega kit. The specific steps refer to the instruction manual.
[0178] Since the sgRNA library ends contain consensus sequences, the editing status of the target sequences was amplified using consensus primers and then combined with high-throughput sequencing technology to analyze the cleavage mechanism and repair results. After amplifying the extracted genomic DNA using the high-fidelity enzyme Q5 (NEB), the first-round PCR product was used as a template, and a second-round amplification was performed using primers adapted to the adapter sequences of the Illumina platform. After recovering the amplified products using the gel extraction kit (Qiagen), the constructed high-throughput sequencing library was quantified using Qubit. After mixing the samples, they were sent to a sequencing company for sequencing using the Novaseq platform. The primer sequences involved are as follows:
[0179] Illumina-R
[0180] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCTTTCAAGACCTAGCTAGCGAATT(SEQ IDNO.14);
[0181] Illumina-b9-R
[0182] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTTCGAACGTGCTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.15);
[0183] Illumina-b10-R、
[0184] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGTAACGTCACCTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.16);
[0185] Illumina-b11-R
[0186] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTAGTGTCACCTACTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.17);
[0187] Illumina-b12-R
[0188] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTTAGTACGCTTGACCTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.18);
[0189] Illumina-b13-R
[0190] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCTAGTCTATGCTAGCTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.19);
[0191] Illumina-b14-R
[0192] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGCTAGTAGCGTACGTCTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.20);
[0193] Illumina-b15-R
[0194] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTAGCTAGTAACGTCTGACTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.21);
[0195] Illumina-b16-R
[0196] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTTAGCTAGTTCATGTGACCTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.22);
[0197] Illumina-b17-R
[0198] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCTAGCTAGTGCCAGACTACTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.23);
[0199] Illumina-b18-R
[0200] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTGCTAGCTAGTCTAGAGGTTCTTTCAAGACCTAGCTAGCGAATT(SEQ ID NO.24);
[0201] Illumina-F
[0202] TTTCCCTACACGACGCTCTTCCGATCTTCTTGTGGAAAGGACGAAACACC(SEQ ID NO.25);
[0203] Illumina-b9-F
[0204] TTTCCCTACACGACGCTCTTCCGATCTTAAGTAGAGTCTTGTGGAAAGGACGAAACACC(SEQ IDNO.26);
[0205] Illumina-b10-F
[0206] TTTCCCTACACGACGCTCTTCCGATCTATCATGCTTATCTTGTGGAAAGGACGAAACACC(SEQ IDNO.27);
[0207] Illumina-b11-F
[0208] TTTCCCTACACGACGCTCTTCCGATCTGATGCACATCTTCTTGTGGAAAGGACGAAACACC(SEQ IDNO.28);
[0209] Illumina-b12-F
[0210] TTTCCCTACACGACGCTCTTCCGATCTCGATTGCTCGACTCTTGTGGAAAGGACGAAACACC(SEQ ID NO.29);
[0211] Illumina-b13-F
[0212] TTTCCCTACACGACGCTCTTCCGATCTTCGATAGCAATTCTCTTGTGGAAAGGACGAAACACC(SEQ ID NO.30);
[0213] Illumina-b14-F
[0214] TTTCCCTACACGACGCTCTTCCGATCTATCGATAGTTGCTTTCTTGTGGAAAGGACGAAACACC(SEQ ID NO.31);
[0215] Illumina-b15-F
[0216] TTTCCCTACACGACGCTCTTCCGATCTGATCGATCCAGTTAGTCTTGTGGAAAGGACGAAACACC(SEQ ID NO.32);
[0217] Illumina-b16-F
[0218] TTTCCCTACACGACGCTCTTCCGATCTCGATCGATTTGAGCCTTCTTGTGGAAAGGACGAAACACC(SEQ ID NO.33);
[0219] Illumina-b17-F
[0220] TTTCCCTACACGACGCTCTTCCGATCTACGATCGATACACGATCTCTTGTGGAAAGGACGAAACACC(SEQ ID NO.34);
[0221] Illumina-b18-F
[0222] TTTCCCTACACGACGCTCTTCCGATCTTACGATCGATGGTCCAGATCTTGTGGAAAGGACGAAACACC(SEQ ID NO.35);
[0223] Illumina-index-F1
[0224] AATGATACGGCGACCACCGAGATCTACACACGCAGTCACACTCTTTCCCTACACGACGCTCTTC(SEQ ID NO.36);
[0225] Illumina-index-F2
[0226] AATGATACGGCGACCACCGAGATCTACACAGTCGACTACACTCTTTCCCTACACGACGCTCTTC(SEQ ID NO.37);
[0227] Illumina-index-F3
[0228] AATGATACGGCGACCACCGAGATCTACACATGACTGCACACTCTTTCCCTACACGACGCTCTTC(SEQ ID NO.38);
[0229] Illumina-index-F4
[0230] AATGATACGGCGACCACCGAGATCTACACTACGACATACACTCTTTCCCTACACGACGCTCTTC(SEQ ID NO.39);
[0231] Illumina-index-R
[0232] CAAGCAGAAGACGGCATACGAGATACCTACGTGTGACTGGAGTTCAGACGTGT(SEQ ID NO.40);
[0233] Illumina-index-R2
[0234] CAAGCAGAAGACGGCATACGAGATAATCGACTGTGACTGGAGTTCAGACGTGT(SEQ ID NO.41);
[0235] Illumina-index-R3
[0236] CAAGCAGAAGACGGCATACGAGATCTACGATTGTGACTGGAGTTCAGACGTGT(SEQ ID NO.42)
[0237] 10. Analyzing the editing results using bioinformatics
[0238] To achieve the goal of inferring the Cas9 cleavage spectrum and repair mode and predicting the cleavage results of any target sequence by analyzing the Cas9 gene editing results, bioinformatics means are used to analyze the data. After quality control screening of the sequencing data, using the synthetic sequence as the reference sequence, and with the personalized analysis program already written, the Bowtie alignment of the sequencing results is carried out to obtain the gene editing results. To infer the Cas9 cleavage spectrum, different gene editing results are subjected to feature extraction. After optimizing the algorithm, the data can be summarized, and different Cas9 cleavage models can be refined and summarized. The purpose is to use this model to predict the editing results corresponding to any target sequence.
[0239] The above experimental results are as follows:
[0240] Figure 1 Shown are three Cas9 nucleases involved in the present invention. The present invention mainly compares the activities between wild-type Cas9 and iSpymacCas9. As shown in the figure, wild-type Cas9 and SpymacCas9 only differ in PID, while iSpymacCas9 is obtained by mutating the 221st arginine and the 394th asparagine to lysine on the basis of SpymacCas9.
[0241] Figure 2 The results show that the optimized sgRNA backbone sequence is also applicable to the iSpymacCas9 nuclease. After introducing a single sgRNA at the CTCF site for single-site editing, first-generation sequencing can be used to evaluate the editing efficiency according to the double-peak ratio at the editing site. It can be seen that the sequence within the box, that is, the sequence at the editing site, has changed, indicating that an editing event has occurred. Among them, the optimized backbone sequence, that is, the sgRNA backbone sequence involved in the present invention, shows the highest editing efficiency (CTCF-N), followed by the sgRNA backbone sequence corresponding to wild-type Cas9 (CTCF-O), and the sgRNA backbone sequence corresponding to Streptococcus macacaeCas9 has the worst efficiency and hardly any editing occurs at this site (CTCF-S).
[0242] Figure 3 Shows the Cas9 expression intensity of monoclonal cell lines stably expressing two Cas9 nucleases involved in the present invention. There are differences in Cas9 expression intensity among different cell lines.
[0243] Referring to the method of the above-mentioned embodiment, after editing the corresponding target sequence with sgRNA and Cas9 nuclease, the editing results can be detected by high-throughput sequencing technology, and then the proportion of various editing types when different Cas9 nucleases cut the target sequence under the mediation of each sgRNA can be calculated.
[0244] After capturing the editing results of wild-type SpyCas9 nuclease (abbreviated as Cas9WT, WT) and iSpymacCas9 nuclease under the mediation of sgRNA using the above-mentioned high-throughput sequencing primers, significant differences were found. As Figure 4 shown in the editing results at the RNF112 locus, the corresponding sgRNA sequence is shown below. The 17th position of this sgRNA sequence is T. It can be seen that in the editing results after wild-type Cas9 cleavage, the ratio of the proportion of base insertion repair type to the proportion of deletion is approximately 1:3, while for iSpymacCas9, the proportion of base insertion repair type is much smaller than that of base deletion. This indicates that iSpymacCas9 is more likely to produce base deletions when the 17th position of sgRNA is T.
[0245] Similarly, as Figure 5 shown in the analysis of the editing results at TUBGCP5, wild-type Cas9 is more likely to produce base insertions at this site, which is similar to previous research results. Previous research results have shown that wild-type Cas9 is more likely to produce base insertions when the 17th position of sgRNA is T (24)(25). It can be seen that the ratio of base insertion to base deletion at this site reaches 3:1. However, when iSpymacCas9 is used for editing, almost no base insertions are found. On the contrary, base deletions dominate. This also indicates that iSpymacCas9 is more likely to produce base deletions when the 17th position of sgRNA is T, and at the same time it reverses the ratio of the wild-type Cas9 repair type.
[0246] As Figure 6 shown, when analyzing the editing results at another different locus NAXD, the ratio of insertion to deletion in the editing of wild-type Cas9 at this site is approximately 1:5. However, when iSpymacCas9 is used for editing, almost all editing results become deletions. This also indicates that iSpymacCas9 is more likely to produce base deletions when the 17th position of sgRNA is T.
[0247] Continuing to analyze the gene editing results of two Cas9 nucleases mediated by this type of sgRNA with the 17th position of sgRNA being A. First is the CNNM4 locus ( Figure 7) It can be seen that at this site, the repair result of base insertion is slightly higher than that of base deletion when cut by wild-type Cas9. However, when switched to iSpymacCas9 cleavage, the repair type of base deletion dominates, and this preference is reversed. This preference of wild-type Cas9 has been reported in the literature (24)(25), indicating that base insertion is more likely to occur when the 17th position of the sgRNA is A. But the results here show that base deletion is more likely to occur when the 17th position of iSpymacCas9 is A.
[0248] Analyze the editing results of another site, RBBP8 ( Figure 8 ). It can be seen that the ratio of base insertion to base deletion in the editing results of wild-type Cas9 is approximately 1:3. However, in the editing results mediated by iSpymacCas9, this gap is further amplified, and there is almost no base insertion. This is similar to the results at the CNNM4 site.
[0249] Similar situations were also found after analyzing the editing results of the EXD1 site ( Figure 9 ). At this site, the ratio of base insertion to base deletion generated by wild-type Cas9 is basically the same, but the base deletion generated by iSpymacCas9 at this site is more than base insertion.
[0250] It can be seen that when the 17th position of the sgRNA is A / T, iSpymacCas9 is more likely to generate base deletions compared to wild-type Cas9. Does this phenomenon stem from the fact that iSpymacCas9 itself is more likely to generate base deletions?
[0251] Furthermore, the editing results mediated by another type of sgRNA were analyzed. The 17th position of this type of sgRNA is G. First, analyze the editing results of the sgRNA at RP11-569G13.3 ( Figure 10 ). It can be seen that at this site, wild-type Cas9 is more likely to generate base deletions, which also conforms to the laws reported in the literature (24)(25). However, after analyzing the editing results of iSpymacCas9, it is found that it is more likely to generate base insertions compared to wild-type Cas9. This suggests that iSpymacCas9 is more likely to generate base insertions when the 17th position of the sgRNA is G compared to wild-type Cas9, rather than the previous speculation that iSpymacCas9 itself is more likely to generate base deletions.
[0252] Furthermore, the editing situation mediated by the sgRNA at another site, CA2M2, was analyzed ( Figure 11) It can be seen that for wild-type Cas9, base deletions are more likely to occur at this site, and the ratio of base deletions to base insertions can reach 4:1. However, for iSpymacCas9, base insertions are more likely to occur at this site, and the ratio of base deletions to base insertions is 3:7. This once again confirms that when the 17th position of the sgRNA is G, iSpymacCas9 is more likely to generate base insertions.
[0253] To further analyze this type of repair of base insertion, base insertions were further divided into 1 - 10 bp for analysis. First, as Figure 12 shown are the editing results mediated by the sgRNA sequence at the NAXD site. The 17th position of the corresponding sgRNA sequence at this site is T. It can be seen that at this site, wild-type Cas9 indeed has a higher frequency of base insertions in each length segment than iSpymacCas9 as previously described, once again verifying that for this type of sgRNA, wild-type Cas9 is more likely to generate base insertions. Taking the RBBP8 site as an example ( Figure 13 ), the 17th position of the corresponding sgRNA sequence at this site is A. It can be seen that wild-type Cas9 is more likely to generate base insertions than iSpymacCas9, and the overall frequency of base insertions is higher than that of iSpymacCas9. Then, the base insertion situation mediated by the sgRNA with the 17th position of the sgRNA sequence being G was analyzed. The corresponding site is CA2M2 ( Figure 14 ). At this site, the frequency of base insertions generated by wild-type Cas9 is much lower than that of iSpymacCas9, which once again coincides with the previous data.
[0254] To analyze the sequence-dependent characteristics of the base insertion preferences of these two nucleases, the sgRNA motifs corresponding to these two Cas9 nucleases that are more likely to generate base insertions were analyzed. As Figure 15 shown are the motif characteristics corresponding to wild-type Cas9. It can be seen that there is no obvious preference for the selected sgRNA sequences ( Figure 15 left), ensuring the reliability of the analysis results. By analyzing the sgRNAs that are more likely to generate base insertions, it was found that when the 17th position of the sgRNA sequence is A / T, base insertions are likely to occur, and conversely, when it is C / G, base deletions are likely to occur. It should be noted that the "more likely" here does not refer to the relative ratio between base insertions and base deletions, but rather the relative situation between sgRNA sequences. This is similar to the characteristics reported in the literature.
[0255] Then, the motif characteristics corresponding to iSpymacCas9 were analyzed ( Figure 16 ). Similarly, there is no obvious preference for the selected sgRNA sequences themselves ( Figure 16On the left), a similar method was used to analyze the sgRNA sequences that are more likely to produce base insertions. It can be seen that for iSpymacCas9, different from wild-type Cas9, when the 17th position of the sgRNA sequence is C / G, base insertions are more likely to occur, and when it is A / T, base deletions are more likely to occur. The meaning of "more likely" here is consistent with the above.
[0256] In summary, compared with the wild-type Cas9 nuclease, the Cas9 nuclease (iSpymacCas9) involved in the present invention has different sequence characteristics dependent on base insertions or base deletions. Although both the wild-type Cas9 and the Cas9 nuclease (iSpymacCas9) involved in the present invention can cleave at the target sequence to generate 5'-overhangs and produce precise base insertions under the mediation of polymerase, the sequence characteristic preferences for mediating base insertions are different between the two. More specifically, when the 17th position of the sgRNA is A / T, more base insertions are mediated by wild-type Cas9, but when the base type at this position is C / G, more base insertions are mediated by iSpymacCas9. In other words, although the cleavage activities of the two are similar, manifested as both being able to generate 5'-overhang ends and direct precise base insertions, the sgRNA sequence characteristics on which they are more likely to produce base insertions are different, with almost opposite results, that is, different sequence-dependent preferences, and can achieve precise base insertions guided by a wider range of sgRNA sequences. Therefore, it can be expected that these two nucleases can provide more possibilities and available tools for base insertion-mediated precise base editing.
[0257] The references of this application are as follows:
[0258] 1. B, M. (1950) The origin and behavior of mutable loci in maize. Proc Natl Acad Sci U S A., 36, 344 - 355. 2. B, M. (1984) The Significance of Responses of the Genome to Challenge. Science, 226, 792 - 801.
[0259] 3. Brinster RL, C. H., Trumbauer M, Senear AW, Warren R, Palmiter RD. (1981) Somatic expression of herpes thymidine kinase in mice following injection of a fusion gene into eggs. Cell, 27, 223 - 231.
[0260] 4. Harbers K, J.D., Jaenisch R. (1981) Microinjection of cloned retroviral genomes into mouse zygotes: integration and expression in the animal. Nature, 293, 540 - 542.
[0261] 5. Gordon JW, S.G., Plotkin DJ, Barbosa JA, Ruddle FH. (1980) Genetic transformation of mouse embryos by microinjection of purified DNA. Proc Natl Acad Sci U S A., 77, 7380 - 7384.
[0262] 6. Palmiter RD, B.R., Hammer RE, Trumbauer ME, Rosenfeld MG, Birnberg NC, Evans RM. (1982) Dramatic growth of mice that develop from eggs microinjected with metallothionein - growth hormone fusion genes. Nature, 300, 611 - 615.
[0263] 7. MR, C. (2005) Gene targeting in mice: functional analysis of the mammalian genome for the twenty - first century. Nat Rev Genet, 6, 507 - 512.
[0264] 8. Bibikova, M., Beumer, K., Trautman, J.K. and Carroll, D. (2003) Enhancing Gene Targeting with Designed Zinc Finger Nucleases. Science, 300, 764 - 764.
[0265] 9. Petolino, J.F., Worden, A., Curlee, K., Connell, J., Strange Moynahan, T.L., Larsen, C. and Russell, S. (2010) Zinc finger nuclease-mediated transgenic deletion. Plant Molecular Biology, 73, 617 - 628.
[0266] 10. Cermak, T., Doyle, E.L., Christian, M., Wang, L., Zhang, Y., Schmidt, C., Baller, J.A., Somia, N.V., Bogdanove, A.J. and Voytas, D.F. (2011) Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic Acids Research, 39, e82 - e82.
[0267] 11. Carroll, D. (2014) Genome Engineering with Targetable Nucleases. Annual Review of Biochemistry, 83, 409 - 439.
[0268] 12. Ain, Q.U., Chung, J.Y. and Kim, Y.-H. (2015) Current and future delivery systems for engineered nucleases: ZFN, TALEN and RGEN. Journal of Controlled Release, 205, 120 - 127.
[0269] 13. Jinek M, C.K., Fonfara I, Hauer M, Doudna JA, Charpentier E. (2012) A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science, 337, 816 - 821.
[0270] 14. Cong L, R.F., Cox D, Lin S, Barretto R, Habib N, Hsu PD, Wu X, Jiang W, Marraffini LA, Zhang F. (2013) Multiplex genome engineering using CRISPR / Cas systems. Science, 339, 819 - 823.
[0271] 15. Mali P, Y.L., Esvelt KM, Aach J, Guell M, DiCarlo JE, Norville JE, Church GM. (2013) RNA-guided human genome engineering via Cas9. Science, 339, 823 - 826.
[0272] 16. Jiang, F. and Doudna, J.A. (2017) CRISPR–Cas9 Structures and Mechanisms. Annual Review of Biophysics, 46, 505 - 529.
[0273] 17. Symington, L.S. and Gautier, J. (2011) Double-Strand Break End Resection and Repair Pathway Choice. Annual Review of Genetics, 45, 247 - 271.
[0274] 18. Pannunzio, N.R., Watanabe, G. and Lieber, M.R. (2018) Nonhomologous DNA end-joining for repair of DNA double-strand breaks. Journal of Biological Chemistry, 293, 10512 - 10523.
[0275] 19. Ali, A., Xiao, W., Babar, M.E. and Bi, Y. (2022) Double-Stranded Break Repair in Mammalian Cells and Precise Genome Editing. Genes, 13.
[0276] 20. Shou, J., Li, J., Liu, Y. and Wu, Q. (2018) Precise and Predictable CRISPR Chromosomal Rearrangements Reveal Principles of Cas9-Mediated Nucleotide Insertion. Molecular Cell, 71, 498-509.e494.
[0277] 21. Shi, X., Shou, J., Mehryar, M.M., Li, J., Wang, L., Zhang, M., Huang, H., Sun, X. and Wu, Q. (2019) Cas9 has no exonuclease activity resulting in staggered cleavage with overhangs and predictable di- and tri-nucleotide CRISPR insertions without template donor. Cell Discovery, 5.
[0278] 22. Joung, J., Konermann, S., Gootenberg, J.S., Abudayyeh, O.O., Platt, R.J., Brigham, M.D., Sanjana, N.E. and Zhang, F. (2017) Genome-scale CRISPR-Cas9 knockout and transcriptional activation screening. Nature Protocols, 12, 828-863.
[0279] 23. Tiscornia, G., Singer, O. and Verma, I.M. (2006) Production and purification of lentiviral vectors. Nature Protocols, 1, 241-245.
[0280] 24. Shen, M.W., Arbab, M., Hsu, J.Y., Worstell, D., Culbertson, S.J., Krabbe, O., Cassa, C.A., Liu, D.R.,
[0281] Gifford, D.K. and Sherwood, R.I. (2018) Predictable and precise template-free CRISPR editing of pathogenic variants. Nature, 563, 646-651.
[0282] 25. Chen, W., McKenna, A., Schreiber, J., Haeussler, M., Yin, Y., Agarwal, V., Noble, W.S. and Shendure, J. (2019) Massively parallel profiling and predictive modeling of the outcomes of CRISPR / Cas9-mediated double-strand break repair. Nucleic Acids Research, 47, 7989-8003.
[0283] The above embodiments are for illustrative purposes of the disclosed embodiments of the present invention and should not be construed as limitations on the present invention. In addition, various modifications listed herein and changes in the methods of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the present invention. Although the present invention has been specifically described in conjunction with various specific preferred embodiments of the present invention, it should be understood that the present invention should not be limited to these specific embodiments. In fact, all such modifications that are obvious to those skilled in the art as described above to obtain the invention should be included within the scope of the present invention.
Claims
1. Use of a gene editing system in preparing a gene editing product, the gene editing system comprising a Cas9 nuclease iSpymacCas9 and sgRNA, wherein the insertion or deletion preference of the gene editing product affected by the type of the PAM-4 base is reversed compared to a gene editing system using a wild-type Cas9 nuclease; The Cas9 nuclease iSpymacCas9 is a polypeptide with an amino acid sequence as shown in SEQ ID No. 1; the sgRNA comprises a backbone sequence with a nucleotide sequence as shown in SEQ ID No. 2; The gene editing product is a nucleotide insertion product with a C / G nucleotide at the -4th base of PAM; Alternatively, the gene editing product is a nucleotide deletion product targeting the -4th base of PAM which is A / T.
2. The use according to claim 1, characterized in that Compared with a gene editing system using a wild-type Cas9 nuclease, the gene editing product has an improved nucleotide insertion efficiency when the -4 base of PAM is C / G, and an improved nucleotide deletion efficiency when the -4 base of PAM is A / T.
3. A composition for gene editing, characterized in that: The composition comprises the Cas9 nuclease iSpymacCas9 and sgRNA for use according to any one of claims 1 to 2; the sgRNA comprises a nucleotide sequence such as a backbone sequence shown in SEQ ID No. 2, and the composition is compared with a gene editing system using a wild-type Cas9 nuclease, and the insertion or deletion preference affected by the -4 base type of the PAM is reversed, wherein when the -4 base of the PAM is C / G, it is a nucleotide insertion, and when the -4 base of the PAM is A / T, it is a nucleotide deletion; The Cas9 nuclease iSpymacCas9 is a polypeptide whose amino acid sequence is shown in SEQ ID No.
1.
4. The composition according to claim 3, characterized in that The composition further comprises a pharmaceutically acceptable carrier and medium; the pharmaceutically acceptable carrier and medium are selected from one or more of sterile water, physiological saline, stabilizers, excipients, antioxidants, buffers, preservatives, chelating agents or adhesives.
5. A gene editing method for editing isolated cell genes, characterized in that: Compared with the gene editing system using wild-type Cas9 nuclease, the insertion or deletion preference affected by the -4th base type of PAM is reversed. The gene editing method is: introducing the composition as described in any one of claims 3-4 into isolated cells to achieve gene editing of the isolated cells.
6. The gene editing method according to claim 5, characterized in that The gene editing method is a method for inserting nucleotides when the -4th position of PAM or the 17th base of the sgRNA sequence is C / G; and / or, the gene editing method is a method for deleting nucleotides when the -4th position of PAM or the 17th base of the sgRNA sequence is A / T.
7. The gene editing method according to claim 5, characterized in that The separated cells are selected from prokaryotic cells or eukaryotic cells, and the eukaryotic cells do not include fertilized egg cells; the prokaryotic cells are selected from one or more of Escherichia coli, Salmonella typhimurium, Klebsiella pneumoniae, Bacteroides ovale, Campylobacter jejuni, Staphylococcus saprophyticus, Enterococcus faecalis, Bacteroides thetaiotaomicron, Bacteroides vulgaris, Bacteroides monothecae, Lactobacillus casei, Bacteroides fragilis, Acinetobacter lwoffii, Fusobacterium nucleatum, Bacteroides johnsii, Bacteroides thaliana, Lactobacillus rhamnosus, Bacteroides massiliense, Bacteroides faecalis, Fusobacterium mortis, Bifidobacterium breve or Listeria; and / or, the eukaryotic cells are selected from one or more of yeast cells, insect cells of Drosophila S2 or Sf9, CHO cells, COS cells, 293 cells, Bowes melanoma cells, 293T cells, T lymphocytes, B lymphocytes, N2A cells, Hela cells, leukocytes or pluripotent stem cells.
8. The gene editing method according to claim 5, characterized in that The introduction method is selected from one or more of DEAE-dextran mediated transfection, liposome mediated transfection, virus or phage infection, conjugation, protoplast fusion, polyethyleneimine mediated transfection, electroporation, gene gun, calcium phosphate precipitation, microinjection, and nanoparticle mediated nucleic acid delivery.
Citation Information
Patent Citations
Cas9 nuclease K918A and application thereof
CN106957831A
Application of crispr / cas system in gene editing
CN111971389A