CRISPR / Cas9 nucleic acid, recombinant vector, recombinant engineering bacterium, mRNA and preparation method, recombinant protein, composition, expression and gene editing method and product

By optimizing the nucleic acid sequence of CRISPR/Cas9 and adjusting the GC content and mRNA secondary structure, the gene editing efficiency and specificity of Cas9 were improved, and the transcription efficiency and stability issues of hSpCas9 in human codon optimization were resolved.

CN120989073APending Publication Date: 2025-11-21YUNZHOU BIOSCIENCES (GUANGZHOU) INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410636900.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, codon optimization of hSpCas9 does not take into account local GC content and mRNA secondary structure, which affects its transcription efficiency and stability, thus limiting the editing efficiency and specificity of CRISPR/Cas9.

Method used

Provide the nucleic acid sequence of CRISPR/Cas9, including specific nucleotide sequences or sequences with similar or high identity to them, and improve the editing efficiency and specificity of Cas9 by adjusting the nucleotide combination to optimize GC content and mRNA secondary structure.

Benefits of technology

By optimizing the nucleotide sequence, the gene editing efficiency and specificity of Cas9 were improved, and the transcription efficiency and stability issues of hSpCas9 in human codon optimization were resolved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120989073A_ABST
    Figure CN120989073A_ABST
Patent Text Reader

Abstract

The invention relates to the field of bioengineering, in particular to CRISPR / Cas9 nucleic acid, a recombinant vector, a recombinant engineering bacterium, mRNA and a preparation method thereof, a recombinant protein, a composition, an expression and gene editing method and a product, and the CRISPR / Cas9 nucleic acid provided by the embodiment of the invention comprises a nucleotide sequence as shown in SEQ ID NO: 1, SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4 or SEQ ID NO: 5. According to the present invention, by optimizing the codon ratio, the GC content, the sequence repeatability, the RNA secondary structure, the RNA free energy and the like in the DNA sequence, the DNA sequence capable of highly expressing the Cas9 protein version is obtained; the expression quantity of the CRISPR / Cas9 protein provided in the technical scheme is increased by 4-8 times compared with that of a wild type or other contrast, and a similar editing effect can be achieved with a lower dosage, so that the application value of the Cas9 protein in gene editing can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioengineering, and more particularly to CRISPR / Cas9 nucleic acids, recombinant vectors, recombinant engineered bacteria, mRNA and their preparation methods, recombinant proteins, compositions, expression and gene editing methods and products. Background Technology

[0002] Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) were initially identified as a special type of DNA sequence found in the genomes of many bacteria and archaea. These sequences are an adaptive immune system developed by microorganisms (especially prokaryotes) during the evolution of life to combat the invasion of viruses and foreign DNA.

[0003] CRISPR-associated nuclease 9 (Cas9) is an effector protein of the CRISPR system, derived from bacteria and archaea, and plays a central role, especially in the Type I CRISPR system. It is an RNA-guided nuclease, meaning that Cas9, guided by specific sgRNAs (single guide RNAs), can recognize and cleave double-stranded DNA, enabling precise genome editing.

[0004] In genome editing applications, a major optimization direction for CRISPR / Cas9 is to improve editing efficiency and enhance editing specificity. For example, codon optimization of the Cas9 sequence can enhance protein expression, thereby improving editing efficiency. Codon optimization refers to adjusting synonymous codons in a gene, eliminating rare codons, and optimizing related parameters such as mRNA secondary structure and motif to improve translation efficiency. Different species and cells may choose different synonymous codons when encoding the same amino acid; this is known as codon bias. During codon optimization, it is necessary to analyze the frequency of synonymous codon usage in the target gene to understand the expression status of the original gene and its optimization potential.

[0005] Existing technologies have addressed human codon optimization of SpCas9, resulting in hSpCas9. However, during hSpCas9's human codon optimization, the original codons were replaced with the most frequent codons among human synonyms. This leads to a locally high GC content (70%-80%) in hSpCas9, which may affect transcription efficiency and consequently gene expression. Simultaneously, current versions of hSpCas9 also exhibit partially low GC content (below 30%), which may affect stability and, consequently, gene expression.

[0006] In summary, existing hSpCas9 technologies simply replace the codons with the optimal human codons without considering the secondary structure of hSpCas9, such as local GC content and mRNA secondary structure. This hinders further optimization of CRISPR / Cas9, namely, improving editing efficiency and enhancing editing specificity. Summary of the Invention

[0007] In view of this, this disclosure provides CRISPR / Cas9 nucleic acids, recombinant vectors, recombinant engineered bacteria, mRNA and preparation methods, recombinant proteins, compositions, expression and gene editing methods and products, aiming to solve or improve the problems in the prior art and enhance the application value of Cas9 in gene editing.

[0008] This disclosure specifically provides the following technical solutions:

[0009] The first aspect of this disclosure provides the nucleic acids of CRISPR / Cas9, including:

[0010] (1) A nucleotide sequence as shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, or SEQ ID NO:5; or

[0011] (2) A nucleotide sequence that is functionally identical or similar to the nucleotide sequence shown in (1) obtained by one or more base substitutions, insertions, deletions, or inversions; or

[0012] (3) A nucleotide sequence that has at least 80% identity with the nucleotide sequence shown in (1) or (2).

[0013]

[0014]

[0015]

[0016] GGACAGCCGGATGAACACCAAGTACGACGAGAACGACAAGCTGAT

[0017] CCGGGAAGTGAAAGTGATCACCCTGAAGAGCAAGCTGGTGAGCGA

[0018] CTTCCGGAAGGACTTCCAGTTCTACAAAGTGCGCGAGATCAACAAC

[0019] TACCACCACGCCCACGACGCCTACCTGAACGCCGTCGTGGGAACCG

[0020] CCCTGATCAAAAAGTACCCCAAGCTGGAAAGCGAGTTCGTGTACG

[0021] GCGACTACAAGGTGTACGACGTGCGGAAGATGATCGCCAAGAGCG

[0022] AGCAGGAAATCGGCAAGGCCACCGCCAAGTACTTCTTCTACAGCA

[0023] ACATCATGAACTTCTTCAAGACCGAGATCACCCTGGCCAACGGCGA

[0024] GATCCGGAAGCGGCCCCTGATCGAGACAAACGGCGAAACCGGGGA

[0025] GATCGTGTGGGACAAGGGCCGGGACTTCGCCACCGTGCGGAAAGT

[0026] GCTGAGCATGCCCCAAGTGAACATCGTGAAAAAGACCGAGGTGCA

[0027] GACAGGCGGCTTCAGCAAAGAGAGCATCCTGCCCAAGAGGAACAG

[0028] CGACAAGCTGATCGCCAGAAAGAAGGACTGGGACCCCAAGAAGTA

[0029] CGGCGGCTTCGACAGCCCCACCGTGGCCTACAGCGTGCTGGTGGTG

[0030] GCCAAAGTGGAAAAGGGCAAGAGCAAGAAACTGAAGAGCGTGAA

[0031] AGAGCTGCTGGGGATCACCATCATGGAAAGAAGCAGCTTCGAGAA

[0032] GAACCCCATCGACTTCCTGGAAGCCAAGGGCTACAAAGAAGTGAA

[0033] AAAGGACCTGATCATCAAGCTGCCCAAGTACAGCCTGTTCGAGCTG

[0034] GAAAACGGCCGGAAGAGAATGCTGGCCAGCGCCGGCGAACTGCAG

[0035] AAGGGAAACGAACTGGCCCTGCCCAGCAAATACGTGAACTTCCTGT

[0036] ACCTGGCCAGCCACTACGAGAAGCTGAAGGGCAGCCCCGAGGACA

[0037] ACGAGCAGAAACAGCTGTTCGTGGAACAGCACAAGCACTACCTGG

[0038] ACGAGATCATCGAGCAGATCAGCGAGTTCAGCAAGAGAGTGATCC

[0039] TGGCCGACGCCAACCTGGACAAAGTGCTGAGCGCCTACAACAAGC

[0040] ACCGGGACAAGCCCATCAGAGAGCAGGCCGAGAACATCATCCACC

[0041] TGTTCACCCTGACCAACCTGGGAGCCCCCGCCGCCTTCAAGTACTT

[0042] CGACACCACCATCGACCGGAAGAGGTACACCAGCACCAAAGAGGT

[0043] GCTGGACGCCACCCTGATCCACCAGAGCATCACCGGCCTGTACGAG

[0044] ACACGGATCGACCTGAGCCAGCTGGGAGGCGAC。(VB3)

[0045]

[0046]

[0047] The second aspect of this disclosure provides a recombinant vector for CRISPR / Cas9, including the nucleic acid and other acceptable elements described in the first aspect of the technical solution.

[0048] In some implementations, the aforementioned acceptable element specifically includes the pmRVac skeleton.

[0049] This disclosure provides, in a third aspect, recombinant engineered bacteria for CRISPR / Cas9, including: a recombinant vector as described in the second aspect above; and a host that is transformed and / or transfected with the recombinant vector.

[0050] In some implementations, the aforementioned host may be Escherichia coli.

[0051] This disclosure provides a fourth aspect of a method for preparing mRNA, the method comprising:

[0052] The recombinant vector described in the second aspect above is processed through steps including linearization, in vitro transcription, and purification to obtain the mRNA.

[0053] The fifth aspect of this disclosure provides mRNA obtained by the aforementioned preparation method.

[0054] The sixth aspect of this disclosure provides a CRISPR / Cas9 recombinant protein, which is composed of mRNA translated as described in the fifth aspect above, and is capable of binding to gRNA.

[0055] The seventh aspect of this disclosure provides a CRISPR / Cas9 composition comprising a CRISPR / Cas9 recombinant protein and gRNA as described in the sixth aspect above.

[0056] This disclosure, in its eighth aspect, provides a method for CRISPR / Cas9 expression, the method comprising the following steps:

[0057] The recombinant vector described in the second aspect above was prepared into mRNA;

[0058] The mRNA was transfected into host cells to induce protein expression.

[0059] In some implementations, the aforementioned host cell may specifically be a 293T cell.

[0060] The ninth aspect of this disclosure provides a method for CRISPR / Cas9 gene editing, the method comprising the following steps:

[0061] The recombinant vector described in the second aspect above was prepared into mRNA;

[0062] The mRNA and gRNA are transfected into host cells to achieve gene editing.

[0063] The tenth aspect of this disclosure provides a gene editing product comprising: nucleic acid as described above, recombinant vector as described above, recombinant engineered bacteria as described above, mRNA as described above, recombinant protein as described above, and / or a composition as described above, as well as acceptable adjuvants, vectors, or devices.

[0064] The aforementioned technical solution, by optimizing factors such as codon ratio, GC content, sequence repeatability, RNA secondary structure, and RNA free energy in the DNA sequence, yielded a DNA sequence capable of highly expressing the Cas9 protein. The CRISPR / Cas9 protein expression level provided by this technical solution is 4-8 times higher than that of the wild type or other comparative types, exhibiting higher gene editing efficiency and thus significantly enhancing the application value of the Cas9 protein in gene editing. Attached Figure Description

[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0066] Figure 1 This is a schematic diagram of DNA sequence analysis for the comparative example (WT) of this disclosure;

[0067] Figure 2 This is a schematic diagram of DNA sequence analysis in Embodiment 1 (VB1) of this disclosure;

[0068] Figure 3 This is a schematic diagram of DNA sequence analysis in Embodiment 2 (VB2) of this disclosure;

[0069] Figure 4 This is a schematic diagram of DNA sequence analysis in Embodiment 3 (VB3) of this disclosure;

[0070] Figure 5 This is a schematic diagram of DNA sequence analysis in Embodiment 4 (VB4) of this disclosure;

[0071] Figure 6 This is a schematic diagram of DNA sequence analysis in Embodiment 5 (VB5) of this disclosure;

[0072] Figure 7 This is a schematic diagram of an in vitro transcription vector for DNA sequences provided in an embodiment of this disclosure;

[0073] Figure 8 The Cas9 protein immunoblot images are for the comparative examples and embodiments of this disclosure.

[0074] Figure 9 This is a graph showing the relative expression levels of Cas9 protein in the comparative examples and embodiments of this disclosure.

[0075] Figure 10 This is a comparative example and embodiment of the present disclosure for functional verification of the Cas9 protein in cells;

[0076] Figure 11 Visual data analysis charts for functional verification comparison of this disclosure. Detailed Implementation

[0077] This disclosure relates to CRISPR / Cas9 nucleic acids, recombinant vectors, recombinant engineered bacteria, mRNA and preparation methods, recombinant proteins, compositions, expression and gene editing methods and products.

[0078] It should be understood that the expression “one or more of…” individually includes each of the objects described after the expression, as well as various different combinations of two or more of the described objects, unless otherwise understood from the context and usage. The expression “and / or” combined with three or more described objects should be understood to have the same meaning, unless otherwise understood from the context.

[0079] The terms “including,” “having,” or “containing,” including the use of their grammatical synonyms, should generally be understood as open-ended and non-restrictive, for example, not excluding other unstated elements or steps, unless otherwise specifically stated or understood from the context.

[0080] It should be understood that the order of the steps or the order in which certain actions are performed is not important as long as the invention remains operational. Furthermore, two or more steps or actions can be performed simultaneously.

[0081] The use of any and all instances or exemplary language such as “e.g.” or “including” in this document is merely intended to better illustrate the invention and is not intended to limit the scope of the invention unless the claims are made. No language in this specification should be construed as indicating that any unclaimed element is essential to the practice of the invention.

[0082] Furthermore, the numerical ranges and parameters used to define the present invention are approximate values, and the relevant values ​​in the specific embodiments have been presented as precisely as possible. However, any value inevitably contains standard deviations due to individual test methods. Therefore, unless explicitly stated otherwise, it should be understood that all ranges, quantities, values, and percentages used in this disclosure are modified with the word "approximately." Here, "approximately" generally means an actual value within plus or minus 10%, 5%, 1%, or 0.5% of a particular value or range.

[0083] The partial sequences involved in the embodiments of this disclosure are as follows:

[0084] gRNA sequence (5'→3')

[0085] gEMX1: GUCACCUCCAAUGACUAGGG. (SEQ ID NO:39)

[0086] PCR amplification primer sequence (5'→3')

[0087] EXM1:

[0088] F:AAAACCACCCTTCTCTCTGGC. (SEQ ID NO:40)

[0089] R:GGAGATTGGAGACACGGAGAG. (SEQ ID NO:41)

[0090] In view of the deficiencies in the prior art, the present disclosure adopts the codon optimization ratio developed by the applicant. By optimizing the codon ratio, GC content, sequence repeatability, RNA secondary structure, RNA free energy, etc. in the DNA sequence, a DNA sequence capable of highly expressing the Cas9 protein version is obtained.

[0091] It should be understood that the aforementioned codon refers to a triplet of nucleotide residues on mRNA (or DNA) that encodes a specific amino acid. The most common codon, for example, is ATG, which stands for methionine.

[0092] The aforementioned codon optimization specifically refers to improving translation efficiency by adjusting synonymous codons in genes, eliminating rare codons, and optimizing related parameters such as mRNA secondary structure and motif. Different species and cells may choose different synonymous codons when encoding the same amino acid; this is known as codon bias. Therefore, to address the shortcomings of existing technologies, codon optimization requires analyzing the frequency of synonymous codon usage in the target gene to understand the expression status of the original gene and its optimization potential.

[0093] The overall route design for synthesizing and optimizing gene sequences mentioned above specifically includes: (1) designing candidate sequences; (2) optimizing codon sequences using the codon probability in the codon usage table; (3) eliminating other unfavorable factors, such as extreme GC content, repetitive sequences, unfavorable mRNA structures, etc.; and (4) adding or deleting restriction enzyme sites as needed.

[0094] The present disclosure is further illustrated below with reference to embodiments:

[0095] Examples 1 to 5 based on artificial codon optimization

[0096] The specific methods for artificial codon optimization are as follows:

[0097] (1) Analysis of amino acid sequence

[0098] The invention only involves one amino acid sequence (SpCas9):

[0099] DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIG

[0100] ALLFDSGETAEATRLKRTARRRRYTRRKNRICYLQEIFSNEMAKVDDSF

[0101] FHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDST

[0102] DKADLRIYLALAMHIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQ

[0103] LFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIAL

[0104] SLGLTPNFKSNFDLAEDAKLQLSKDTYDDDNLLAQIGDQYADLFL

[0105] AAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVR

[0106] QQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEEL

[0107] LVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNRE

[0108] KIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGAS

[0109] AQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGM

[0110] RKPAFLSGEQKKAIVDLLFKTTNRKVTVKQLKEDYFKKIECFDSVEISGV

[0111] EDRFNASLGTYHDLLKIIKDKDFLDNEENEDELEDIVLTLTLFEDREMIE

[0112] ERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTIL

[0113] DFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLA

[0114] GSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQK

[0115] NSRERMKRIEEGIGELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMY

[0116] VDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVP

[0117] SEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIK

[0118] RQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLSKLVSDF

[0119] RKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYKPLESEFVYGDY

[0120] KVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFKTEITLANGEIRKRPL

[0121] IETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESI

[0122] LPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKK

[0123] LKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFEL

[0124] ENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNE

[0125] QKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPI

[0126] REQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD. (As shown in SEQ ID NO:7)

[0127] (2) DNA sequence conversion

[0128] Based on the codon preference table, the amino acid sequence was converted to the most commonly used codon sequence for the desired species. The frequency of human codon usage is shown in the table below:

[0129]

[0130]

[0131] For example, in the species "Homo sapiens", all "alanine" is replaced with "GCC". This yields the first version of the DNA sequence, hSpCas9-WT.

[0132] The hSpCas9-WT sequence:

[0133] GACAAGAAGTACAGCATCGGCCTGGACATCGGCACCAACTCTGTG

[0134] GGCTGGGCCGTGATCACCGACGAGTACAAGGTGCCCAGCAAGAAA

[0135] TTCAAGGTGCTGGGCAACACCGACCGGCACAGCATCAAGAAGAAC

[0136] CTGATCGGAGCCCTGCTGTTCGACAGCGGCGAAACAGCCGAGGCC

[0137] ACCCGGCTGAAGAGAACCGCCAGAAGAAGATACACCAGACGGAAG

[0138] AACCGGATCTGCTATCTGCAAGAGATCTTCAGCAACGAGATGGCCA

[0139] AGGTGGACGACAGCTTTCTTCCACAGACTGGAAGAGTCCTTCCTGGT

[0140] GGAAGAGGATAAGAAGCACGAGCGGCACCCCATCTTCGGCAACAT

[0141] CGTGGACGAGGTGGCCTACCACGAGAAGTACCCCACCATCTACCAC

[0142] CTGAGAAAGAAACTGGTGGACAGCACCGACAAGGCCGACCTGCGG

[0143] CTGATCTATCTGGCCCTGGCCCACATGATCAAGTTCCGGGGCCACT

[0144] TCCTGATCGAGGGCGACCTGAACCCCGACAACAGCGACGTGGACA

[0145] AGCTGTTCATCCAGCTGGTGCAGACCTACAACCAGCTGTTCGAGGA

[0146] AAACCCCATCAACGCCAGCGGCGTGGACGCCAAGGCCATCCTGTCT

[0147] GCCAGACTGAGCAAGAGCAGACGGCTGGAAAATCTGATCGCCCAG

[0148] CTGCCCGGCGAGAAGAAGAATGGCCTGTTCGGAAACCTGATTGCCC

[0149] TGAGCCTGGGCCTGACCCCCAACTTCAAGAGCAACTTCGACCTGGC

[0150] CGAGGATGCCAAACTGCAGCTGAGCAAGGACACCTACGACGACGA

[0151] CCTGGACAACCTGCTGGCCCAGATCGGCGACCAGTACGCCGACCTG

[0152] TTTCTGGCCGCCAAGAACCTGTCCGACGCCATCCTGCTGAGCGACA

[0153] TCCTGAGAGTGAACACCGAGATCACCAAGGCCCCCCTGAGCGCCTC

[0154] TATGATCAAGAGATACGACGAGCACCACCAGGACCTGACCCTGCT

[0155] GAAAGCTCTCGTGCGGCAGCAGCTGCCTGAGAAGTACAAAGAGAT

[0156] TTTCTTCGACCAGAGCAAGAACGGCTACGCCGGCTACATTGACGGC

[0157] GGAGCCAGCCAGGAAGAGTTCTACAAGTTCATCAAGCCCATCCTGG

[0158] AAAAGATGGACGGCACCGAGGAACTGCTCGTGAAGCTGAACAGAG

[0159] AGGACCTGCTGCGGAAGCAGCGGACCTTCGACAACGGCAGCATCC

[0160] CCCACCAGATCCACCTGGGAGAGCTGCACGCCATTCTGCGGCGGCA

[0161] GGAAGATTTTTACCCATTCCTGAAGGACAACCGGGAAAAGATCGA

[0162] GAAGATCCTGACCTTCCGCATCCCCTACTACGTGGGCCCTCTGGCC

[0163] AGGGGAAACAGCAGATTCGCCTGGATGACCAGAAAGAGCGAGGAA

[0164] ACCATCACCCCCTGGAACTTCGAGGAAGTGGTGGACAAGGGCGCTT

[0165] CCGCCCAGAGCTTCATCGAGCGGATGACCAACTTCGATAAGAACCT

[0166] GCCCAACGAGAAGGTGCTGCCCAAGCACAGCCTGCTGTACGAGTA

[0167] CTTCACCGTGTATAACGAGCTGACCAAAGTGAAATACGTGACCGAG

[0168] GGAATGAGAAAGCCCGCCTTCCTGAGCGGCGAGCAGAAAAAGGCC

[0169] ATCGTGGACCTGCTGTTCAAGACCAACCGGAAAGTGACCGTGAAG

[0170] CAGCTGAAAGAGGACTACTTCAAGAAAATCGAGTGCTTCGACTCCG

[0171] TGGAAATCTCCGGCGTGGAAGATCGGTTCAACGCCTCCCTGGGCAC

[0172] ATACCACGATCTGCTGAAAATTATCAAGGACAAGGACTTCCTGGAC

[0173] AATGAGGAAAACGAGGACATTCTGGAAGATATCGTGCTGACCCTG

[0174] ACACTGTTTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACC

[0175] TATGCCCACCTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGG

[0176] CGGAGATACACCGGCTGGGGCAGGCTGAGCCGGAAGCTGATCAAC

[0177] GGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGATTTCCTGA

[0178] AGTCCGACGGCTTCGCCAACAGAAACTTCATGCAGCTGATCCACGA

[0179] CGACAGCCTGACCTTTAAAGAGGACATCCAGAAAGCCCAGGTGTC

[0180] CGGCCAGGGCGATAGCCTGCACGAGCACATTGCCAATCTGGCCGG

[0181] CAGCCCCGCCATTAAGAAGGGCATCCTGCAGACAGTGAAGGTGGT

[0182] GGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGAGAACAT

[0183] CGTGATCGAAATGGCCAGAGAGAACCAGACCACCCCAGAAGGGGACA

[0184] GAAGACAGCCGCGAGAGAGAATGAAGCGGATCGAAGAGGGCATCA

[0185] AAGAGCTGGGCAGCCAGATCCTGAAAGAACACCCCGTGGAAAACCA

[0186] CCCAGCTGCAGAACGAGAAGCTGTACCTGTACTACCTGCAGAATGG

[0187] GCGGGATATGTACGTGGACCAGGAACTGGACATCAACCGGCTGTC

[0188] CGACTACGATGTGGACCATATCGTGCCTCAGAGCTTTCTGAAGGAC

[0189] GACTCCATCGACAACAAGGTGCTGACCAGAAGCGACAAGAACCGG

[0190] GGCAAGAGCGACAACGTGCCCTCCGAAGGAGGTCGTGAAAGAGATG

[0191] AAGAACTGGCGGCAGCTGCTGAACGCCAAGCTGATTACCCAG

[0192] AGAAAGTTCGACAATCTGACCAAGGCCGAGAGAGGCGGCCTGAGC

[0193] GAACTGGATAAGGCCGGCTTCATCAAGAGACAGCTGGTGGGAAACC

[0194] CGGCAGATCACAAAGCACGTGGCACAGATCCTGGACTCCCGGATG

[0195] AACACTAAGTACGACGAGAATGACAAGCTGATCCGGGAAGTGAAA

[0196] GTGATCACCCTGAAGTCCAAGCTGGTGTCCGATTTCCGGAAGGATT

[0197] TCCAGTTTTACAAAGTGCGCGAGATCAACAACTACCACCACGCCCA

[0198] CGACGCCTACCTGAACGCCGTCGTGGGAACCGCCCTGATCAAAAA

[0199] GTACCCTAAGCTGGAAAGCGAGTTCGTGTACGGCGACTACAAGGT

[0200] GTACGACGTGCGGAAGATGATCGCCAAGAGCGAGCAGGAAATCGG

[0201] CAAGGCTACCGCCAAGTACTTCTTCTACAGCAACATCATGAACTTT

[0202] TTCAAGACCGAGATTACCCTGGCCAACGGCGAGATCCGGAAGCGG

[0203] CCTCTGATCGAGACAAACGGCGAAACCGGGGAGATCGTGTGGGAT

[0204] AAGGGCCGGGATTTTGCCACCGTGCGGAAAGTGCTGAGCATGCCCC

[0205] AAGTGAATATCGTGAAAAAGACCGAGGTGCAGACAGGCGGCTTCA

[0206] GCAAAGAGTCTATCCTGCCCAAGAGGAACAGCGATAAGCTGATCG

[0207] CCAGAAAGAAGGACTGGGACCCTAAGAAGTACGGCGGCTTCGACA

[0208] GCCCCACCGTGGCCTATTCTGTGCTGGTGGTGGCCAAAGTGGAAAA

[0209] GGGCAAGTCCAAGAAACTGAAGAGTGTGAAAGAGCTGCTGGGGAT

[0210] CACCATCATGGAAAGAAGCAGCTTCGAGAAGAATCCCATCGACTTT

[0211] CTGGAAGCCAAGGGCTACAAAGAAGTGAAAAAGGACCTGATCATC

[0212] AAGCTGCCTAAGTACTCCCTGTTCGAGCTGGAAAACGGCCGGAAG

[0213] AGAATGCTGGCCTCTGCCGGCGAACTGCAGAAGGGAAACGAACTG

[0214] GCCCTGCCCTCCAAATATGTGAACTTCCTGTACCTGGCCAGCCACT

[0215] ATGAGAAGCTGAAGGGCTCCCCCGAGGATAATGAGCAGAAACAGC

[0216] TGTTTGTGGAACAGCACAAGCACTACCTGGACGAGATCATCGAGCA

[0217] GATCAGCGAGTTCTCCAAGAGAGTGATCCTGGCCGACGCTAATCTG

[0218] GACAAAGTGCTGTCCGCCTACAACAAGCACCGGGATAAGCCCATC

[0219] AGAGAGCAGGCCGAGAATATCATCCACCTGTTTACCCTGACCAATC

[0220] TGGGAGCCCCTGCCGCCTTCAAGTACTTTGACACCACCATCGACCG

[0221] GAAGAGGTACACCAGCACCAAAGAGGTGCTGGACGCCACCCTGAT

[0222] CCACCAGAGCATCACCGGCCTGTACGAGACACGGATCGACCTGTCTCAGCTGGGAGGCGAC. (As shown in SEQ ID NO:6)

[0223] (3) DNA sequence analysis

[0224] DNA sequence analysis was performed on hSpCas9-WT, including GC content percentage analysis (https: / / www.vectorbuilder.cn / tool / gc-content-calculator.html), sequence repeat analysis (https: / / www.vectorbuilder.cn / tool / sequence-dot-plot.html), and RNA secondary structure and free energy analysis (http: / / rna.tbi.univie.ac.at / / cgi-bin / RNAWebSuite / RNAfold.cgi).

[0225] (4) DNA sequence adjustment

[0226] By combining parameters such as codon bias, GC content, sequence repetition, RNA secondary structure, and RNA free energy, an optimized version of the DNA sequence is obtained.

[0227] The optimization of each parameter is described below.

[0228] 1) Codon bias: After obtaining the first version of the DNA sequence based on the most frequently used codons, the first-most frequently used codons are replaced sequentially with the second, third, and fourth most frequently used codons. The general principle is that the percentages of the first, second, third, and fourth most frequently used codons in the total sequence of the protein should be controlled within the following ranges: 40–100%, 0–50%, 0–25%, and 0–15%, respectively.

[0229] 2) GC content percentage: higher than 0% to 15% of the total GC content of the species. The total GC content of the species "Homo sapiens" is 52.27%, so the total GC content percentage of the optimized DNA sequence should be "52.27% to 67.27%".

[0230] 3) Sequence repetition: Optimize the number of repetitive sequences and the sequence length as much as possible.

[0231] 4) RNA secondary structure and free energy: Reduce the number of consecutive base-pairing regions to lower the minimum free energy. If both cannot be achieved simultaneously, prioritize reducing the number of consecutive base-pairing regions.

[0232] Based on the above optimization method, a total of 5 versions of nucleic acid sequences were formed, including hSpCas9 from the existing technology as a comparative example, and they are named as follows:

[0233] SEQ ID NO Name Description 1 WT Comparative Example (SEQ ID NO: 6) 2 VB1 Example 1 (SEQ ID NO: 1) 3 VB2 Example 2 (SEQ ID NO: 2) 4 VB3 Example 3 (SEQ ID NO: 3) 5 VB4 Example 4 (SEQ ID NO: 4) 6 VB5 Example 5 (SEQ ID NO: 5)

[0234] Carrier construction

[0235] like Figure 7As shown, the optimized hSpCas9 DNA sequence was inserted into the pmRVac backbone for in vitro transcription using the Gibson method as the antigen coding region. The comparison ratio was consistent with the DNA sequences of other elements besides the hSpCas9 DNA sequence in each version of the examples, yielding a Gibson reaction buffer containing the pmRVac-hSpCas9 vector. The reaction buffer was then chemically transformed into VB UltraStable competent cells, incubated on ice for 30 min, in a 42°C hot water bath for 1 min, and on ice for 2 min. LB medium was added and incubated at 37°C for 1 h. Finally, the cells were transferred to LB plates containing Kanab antibiotics and incubated overnight at 37°C for 16 h.

[0236] Cloning identification

[0237] Multiple single clones were picked from the LB plates cultured overnight in the aforementioned vector construction steps, dissolved in sterile water to obtain bacterial culture, 2 μl of bacterial culture was taken for PCR amplification, the correct clones were identified by gel electrophoresis, and the remaining bacterial culture was placed in LB medium and cultured overnight at 37°C for 16 h.

[0238] All sequencing primer sequences are shown in the table below:

[0239]

[0240]

[0241] The primers used for sequencing each vector are shown in the table below:

[0242] Name Primer Used WT Seq F1, Seq F2, Seq F3, Seq F4, Seq F5, Seq F6, Seq F7, Seq F29, Seq F30 VB1 Seq F6, Seq F8, Seq F9, Seq F10, Seq F11, Seq F17, Seq F27, Seq F29, Seq F31 VB2 Seq F1, Seq F3, Seq F5, Seq F12, Seq F13, Seq F14, Seq F17, Seq F29, Seq F31 VB3 Seq F1, Seq F4, Seq F5, Seq F6, Seq F16, Seq F17, Seq F28, Seq F29, Seq F31 VB4 Seq F18, Seq F19, Seq F20, Seq F21, Seq F22, Seq F23, Seq F24, Seq F29, Seq F30 VB5 Seq F19, Seq F20, Seq F21, Seq F23, Seq F24, Seq F25, Seq F26, Seq F29, Seq F30

[0243] plasmid extraction

[0244] The bacterial culture obtained overnight was used to extract plasmids using a plasmid extraction kit. The obtained plasmids were then identified by enzyme digestion, sequencing, and other methods to obtain the correct plasmids.

[0245] plasmid linearization

[0246] The aforementioned plasmid was digested with SapI and reacted at 37°C for 1 hour. The linearized plasmid was then purified using a DNA purification kit to obtain the correct linearized plasmid.

[0247] mRNA preparation

[0248] The linearized plasmid obtained above was added to the transcription system and reacted at 37°C for 2 hours to obtain crude mRNA. After removing the DNA template, the mRNA sample was purified by magnetic beads.

[0249] Western Blot (WB) Experiment

[0250] The prepared mRNA was added to 1 μg of mRNA per well of a 12-well plate containing 293T cells (10E6 cells / well). The cells were incubated at 37°C with 5% CO2 for 24 h, then lysed. The supernatant was collected by centrifugation, placed on ice, and labeled accordingly. The supernatant was thoroughly mixed, and 4×SDS loading was added. The mixture was heated to denature the supernatant and then placed on ice. The pre-cast gel was assembled and placed in an electrophoresis tank. Buffer was added, and the sample was loaded. Electrophoresis was performed at 150 mV for 1 h. A 0.45 μm PVDF membrane of appropriate size was cut and immediately transferred after electrophoresis at 250 mA for 45 min. After electrophoresis, the membrane was removed with forceps and placed in a container. After rinsing, blocking buffer was added, and the membrane was blocked at room temperature for 1 h. The membrane was washed twice with TBST and incubated overnight with primary antibody. After washing three times with TBST, the membrane was incubated at room temperature for 1.5 h with secondary antibody. After 3 TBST washes, ECL Western blot substrate was added to the membrane, and luminescence imaging was performed to observe protein expression results. The results are as follows: Figure 8 As shown, the protein expression in Examples 1-5 was stronger than that in the control group.

[0251] After luminescence imaging, grayscale scanning analysis was performed using ImageJ software. Following grayscale scanning, GAPDH was used as a reference to normalize all samples, and then the fold change of each group relative to the control group was calculated.

[0252] Data processing results are as follows Figure 9 As shown, Example 5 showed a 4-fold increase in expression compared to the comparative example, Examples 1 and 2 showed a 5-fold increase in expression compared to the comparative example, and Examples 3 and 4 showed an 8-fold increase in expression compared to the WT group.

[0253] Cell transfection

[0254] Take the prepared mRNA and adjust the input amount according to the relative protein expression fold change results obtained from the Western blot experiment, as shown in the table below. To ensure that the total input amount of mRNA is the same, use nonsense RNA to supplement the input amount, so that the total amount is 1 μg. Use 293T cells in 12-well plates (10E6 cells / well), and add 1 μg mRNA and 1 μg g gRNA (as mentioned above, gEMX1) to each well. Collect cells 48 h after transfection, remove as much supernatant as possible, and store at -80℃.

[0255] Experimental Group SpCas9 mRNA Nonsense RNA 1 WT mRNA: 1.000 pg 0.000 pg 2 VB1 mRNA: 0.250 pg 0.750 pg 3 VB2 mRNA: 0.250 pg 0.750 pg 4 VB3 mRNA: 0.125 pg 0.875 pg 5 VB3 mRNA: 0.250 pg 0.875 pg 6 VB3 mRNA: 0.500 pg 0.500 pg 7 VB4 mRNA: 0.125 pg 0.875 pg 8 VB4 mRNA: 0.250 pg 0.750 pg 9 VB4 mRNA: 0.500 pg 0.500 pg 10 VB5 mRNA: 0.250 pg 0.750 pg

[0256] Cell genome extraction

[0257] Extract the aforementioned cell genome using a cell genomic DNA extraction kit, following the kit instructions. After extraction, take 2 μL to determine the concentration and store for later use.

[0258] Fragments required for PCR amplification

[0259] Based on the following components and corresponding volumes, the extracted cell genome was used to perform PCR amplification of the target fragment:

[0260] Component Volume Taq Enzyme (5 U / µl) 1 µL 10 x Taq Buffer 10 µL dNTP Mixture (2.5 mM each) 16 µL Template (Cell Genomic, from the aforementioned extraction step) 1 µg Primer 1 (F, as aforementioned) 4 µL Primer 2 (R, as aforementioned) 4 µL Deionized Water up to 100 µL

[0261] The PCR amplification procedure includes:

[0262] 94℃, 1min; → [98℃, 10s → 68℃, 60s] × 30 cycles → 72℃, 10min → 4℃, save temporarily.

[0263] PCR product purification and recovery

[0264] The PCR products obtained from the aforementioned PCR amplification were purified and recovered using a standard DNA product purification kit, following the instructions. After purification, 2 μL was taken to determine the concentration and kept for later use.

[0265] T7E1 enzyme digestion

[0266] The PCR purified product obtained from the aforementioned PCR product purification and recovery steps was used to identify its gene editing efficiency by T7E1 restriction enzyme digestion. The specific operation is as follows.

[0267] The annealing system for PCR purification products is as follows:

[0268] Component Volume PCR Purified Product (from the aforementioned step) 400 ng 10 x Buffer 2 µL Nuclease-Free Water Add to 19μL

[0269] The procedure is as follows:

[0270] 95℃, 5min→95-85℃, -2℃ / s→85-25℃, -0.1℃ / s→4℃, temporarily stored.

[0271] Add 1 μL of T7 Endonuclease I to the annealing reaction system, digest at 37℃ for 15 min, and analyze by 2% agarose gel electrophoresis.

[0272] The results are as follows Figure 10 As shown, in functional verification, reducing the proportion of mRNA input in the disclosed embodiments still achieves editing efficiency similar to that of the comparative embodiments.

[0273] Image J was used to perform grayscale scanning analysis on the agarose gel electrophoresis images, and the gene editing efficiency of each experimental group was calculated according to the publicly published calculation method (see: Guschin, DY, et al. (2010) A rapid and general assay for monitoring endogenous gene modification. Methods Mol Biol, 649, 247–256.).

[0274] Data processing results are as follows Figure 11 As shown, in Example 3 of this disclosure, when the amount of mRNA used is 1 / 8 of that in the comparative example, a similar editing efficiency is still achieved. This significantly reduces the amount of mRNA used. When the amount used in Examples 1, 2, and 5 is 1 / 4 of that in the comparative example, the editing efficiency is close to that of the comparative example.

[0275] In summary, the hSpCas9 protein expression ratio provided in the embodiments of this disclosure is improved compared to the control group. This facilitates better application in in vivo experiments and also enables the production of products with market competitiveness.

[0276] The above description is only a preferred embodiment of this disclosure. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure.

Claims

1. A nucleic acid of CRISPR / Cas9, characterized in that, comprises: (1) a nucleotide sequence as shown in SEQ ID NO: 1 or SEQ ID NO: 2 or SEQ ID NO: 3 or SEQ ID NO: 4 or SEQ ID NO: 5; or (2) a nucleotide sequence functionally identical or similar to the nucleotide sequence shown in (1) obtained by one or more base substitutions, insertions, deletions or inversions; or (3) a nucleotide sequence having at least 80% identity to the nucleotide sequence shown in (1) or (2).

2. A recombinant vector of CRISPR / Cas9, characterized by, comprises the nucleic acid as claimed in claim 1 and other acceptable elements.

3. Recombinant engineering bacteria of CRISPR / Cas9, characterized in that, comprises: the recombinant vector as claimed in claim 2; a host transformed and / or transfected with the recombinant vector.

4. Method for the preparation of mRNA, characterized in that, comprises: the mRNA obtained by treating the recombinant vector as claimed in claim 2 with steps including linearization, in vitro transcription and purification.

5. The mRNA obtained by the preparation method of claim 4.

6. A CRISPR / Cas9 recombinant protein characterized in that, the CRISPR / Cas9 recombinant protein is translated from the mRNA as claimed in claim 5 and can bind with gRNA.

7. A CRISPR / Cas9 composition characterized in that, comprises the CRISPR / Cas9 recombinant protein as claimed in claim 6 and gRNA.

8. A method of expressing CRISPR / Cas9, characterized by, comprises the following steps: preparing the recombinant vector as claimed in claim 2 into mRNA; transfecting the mRNA into host cells to induce protein expression.

9. A method of CRISPR / Cas9 gene editing, characterized in that, comprises the following steps: preparing the recombinant vector as claimed in claim 2 into mRNA; transfecting the mRNA and gRNA into host cells to achieve gene editing.

10. A gene editing product, characterized in that, comprises: the nucleic acid as claimed in claim 1, the recombinant vector as claimed in claim 2, the recombinant engineering bacteria as claimed in claim 3, the mRNA as claimed in claim 5, the recombinant protein as claimed in claim 6 and / or the composition as claimed in claim 7, and acceptable adjuvants, carriers or devices.