Codon optimization method
By adjusting the parameters such as GC content, repeat sequence and secondary structure of the nucleic acid sequence, combined with the codon preference of the host, the DNA sequence is optimized to improve gene expression efficiency, and the problem of unstable optimization results in the prior art is solved, and higher protein expression and mRNA stability are achieved.
Patent Information
- Application Number
- PCT/CN2024/116834
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-21
- Filing Date
- 2024-09-04
- Publication Date
- 2025-07-03
AI Technical Summary
Existing codon optimization tools have limitations in improving gene expression efficiency and cannot provide multiple options. The optimization results vary greatly depending on the host cell, and sequences with high CAI values are not expressed well in actual operations.
According to the codon preference of the host, the GC content, repeat sequence, secondary structure and free energy of the nucleic acid sequence were adjusted, and the DNA sequence of highly expressed proteins was obtained through experimental screening. The optimized sequence showed higher expression effects in specific host cells.
It significantly improves the expression amount of protein and mRNA stability, enhances the expression effect in specific host cells, provides more optimization protocol choices, and reduces the energy demand for DNA replication.
Smart Images

Figure PCTCN2024116834-APPB-I100001 
Figure PCTCN2024116834-APPB-I100002 
Figure PCTCN2024116834-APPB-I100003
Abstract
Description
Codon optimization methods
[0001] This application claims priority from the following cases, the entire contents of which are incorporated herein by reference.
[0002] A Chinese patent application, application number 202311851197.5, titled “Nucleic Acids, mRNA and Their Applications,” was submitted to the China Patent Office on December 28, 2023;
[0003] A Chinese patent application, filed with the China Patent Office on May 21, 2024, with application number 202410635658.3 and the title of the invention being “Nucleic acid encoding chicken ovalbumin and expression vector”;
[0004] A Chinese patent application, filed with the China Patent Office on May 21, 2024, with application number 202410635650.7 and titled “Cre recombinase encoding nucleic acid and its application”;
[0005] A Chinese patent application filed with the China Patent Office on May 21, 2024, with application number 202410636918.9, titled “Gaussia luciferase nucleic acid, recombinant expression vector and system, recombinant engineered bacteria, mRNA and preparation method, expression method, application and detection product”;
[0006] A Chinese patent application was submitted to the China Patent Office on May 21, 2024, with application number 202410636900.9 and invention name: "CRISPR / Cas9 nucleic acids, recombinant vectors, recombinant engineered bacteria, mRNA and preparation methods, recombinant proteins, compositions, expression and gene editing methods and products". Technical Field
[0007] The present invention relates to the technical field of molecular biology, and in particular to a method for codon optimization. Background Art
[0008] Codons are specific sequences of three adjacent nucleotides on messenger RNA (mRNA) that determine the type of amino acids and, in turn, the sequence of proteins. Codon optimization, also known as codon engineering or codon adaptation, is a technique that optimizes gene expression efficiency in a specific host by altering codon usage within a gene sequence.
[0009] Codon optimization methods primarily include computer simulation optimization, experimental screening optimization, and big data-based deep learning methods. Computer simulation optimization uses algorithms to predict optimal codon combinations, experimental screening optimization experimentally compares the expression efficiency of different codon combinations, and big data-based deep learning methods use large-scale bioinformatics data to train models and predict optimal codons. Optimization strategies can include local or global codon substitutions within gene sequences based on the codon preferences of host cells to improve expression efficiency. Currently, many companies have automated codon optimization tools available on their websites.
[0010] Existing codon optimization tools primarily use IT algorithms to optimize GC content, codon usage frequency, mRNA secondary structure, RNase splicing sites, and RNA-stabilizing trans-acting elements for a specific cell, thereby increasing protein expression at the plasmid level. Their downstream applications primarily target recombinant proteins: 1) plasmid transfection into cells, with the expressed protein identified through Western blotting, ELISA, and other assays; 2) plasmid transformation into hosts (such as E. coli, yeast, insect, or mammalian cells), inducing large-scale protein expression and purifying the protein to obtain a product. Furthermore, IT algorithms often provide unique codon-optimized sequences, making them inconvenient for clients to screen.
[0011] As can be seen, existing codon optimization techniques have achieved significant results in improving gene expression efficiency. However, this technology also has certain limitations. For example, IT algorithms only input a unique sequence and cannot provide a wide range of options. Different host cells have different codon preferences, so optimization results may vary depending on the host cell. Furthermore, the sequences output by IT algorithms often rely solely on the CAI value to judge the quality of the sequence. However, in practice, sequences with high CAI values often do not show good expression results.
[0012] Summary of the Invention
[0013] In view of this, the technical problem to be solved by the present invention is to provide a method for codon optimization in order to further improve the expression level.
[0014] In the present invention, the codon optimization method comprises:
[0015] Obtaining a protein-coding nucleic acid sequence based on the host's codon preference, and then adjusting the GC content, repeat sequence, secondary structure, and / or free energy in the resulting sequence to obtain an optimized coding nucleic acid;
[0016] The GC content includes: the GC content of the entire length of the encoding nucleic acid and the GC content within a unit length range.
[0017] The method of the present invention replaces codons according to set parameters, which can reduce the vast number of codon sorting schemes to only dozens or even a few codon sorting schemes. The optimized sequence can be screened experimentally to obtain a DNA sequence with high protein expression.
[0018] The GC content in a nucleic acid sequence refers to the ratio of guanine (G) to cytosine (C). The method of the present invention not only adjusts the GC content over the entire length, but also controls the local GC content in the sequence. When the local GC content is too high, the codons in that part are replaced. By dually controlling the local GC content and the full-length GC content, the energy required for the target gene during DNA replication can be further reduced, reducing the energy demand caused by high GC or low GC sequences, thereby improving the stability of mRNA and thus increasing protein production.
[0019] In the present invention, the local GC content, that is, the GC content within a unit length range, refers to the proportion of G and C in a nucleic acid sequence fragment of a specific length. For example, the GC content within a unit length range in the present invention is the GC content within a length of 10bp to 100bp. Preferably, it is the GC content within a length of 20bp to 50bp. More preferably, it is the GC content within a length of 30bp. Taking a unit length of 10bp as an example, the GC content within the unit length range refers to the GC content within the range of 1 to 10bp, 2 to 11bp, 3 to 12bp...n to n+9bp (and so on) in the sequence. In a sequence to be optimized, the unit length is at least one of 10bp to 100bp, for example, its length can be at least one of 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100bp.
[0020] In some embodiments, the GC content per unit length is 20% to 95%. Preferably, the GC content is 26% to 63% or 40% to 83%. In some specific embodiments, the GC content per unit length is 26% to 63% or 43% to 80%, or 46% to 83%, or 40% to 80%.
[0021] In some embodiments, the full-length GC content is higher than 0% to 15% of the total GC content of the host species.
[0022] In some specific embodiments, taking human cells as the host, the full-length GC content is adjusted to 52.27% to 67.27%. Preferably, using human cells as the host, the full-length GC content is adjusted to 56% to 63% by codon optimization.
[0023] In other specific embodiments, taking cynomolgus macaque cells as the host, the full-length GC content is adjusted to 49.64% to 64.64%. Preferably, taking cynomolgus macaque cells as the host, the full-length GC content is adjusted to 54% to 61%.
[0024] The repetitive sequence in the nucleic acid sequence refers to the identical or symmetrical sequence fragments that appear at different positions in the nucleic acid sequence encoding the protein. If there are repetitive fragments in the amino acid sequence of the protein, the encoding nucleic acid is likely to also have corresponding repetitive sequences. However, previous optimization methods have paid less attention to the repetitive sequences in the optimized nucleic acid sequence. In the present invention, the longest repetitive sequence is controlled within 30bp, which can improve the stability of the target gene during DNA replication and reduce the loss of fragments caused by homologous recombination or other reasons that occur during replication of the target gene. In an embodiment of the present invention, the length of the repetitive sequence is not more than 30bp, and preferably, the repetitive sequence length is not more than 20bp.
[0025] Nucleic acid secondary structure involves the DNA double helix and RNA folding. The secondary structure of RNA refers to the three-dimensional spatial structure formed by the spatial folding of single-stranded RNA molecules. It is usually manifested as a hairpin-shaped single-stranded structure, in which the single strands fold back to form a local small double helix, also known as a stem-loop structure or a globular loop structure. Avoiding the formation of RNA secondary structure is conducive to more accurate and efficient protein translation, and is also more conducive to mRNA to more accurately and efficiently exert its physiological activity. In the present invention, the adjustment of the secondary structure includes reducing the continuous base pairing region.
[0026] mRNA free energy refers to the energy absorbed or released when an mRNA molecule transitions from one conformation to another under specific conditions. It reflects the structural stability of mRNA and is an important factor in evaluating its folding state and function. Codon-optimized nucleic acid sequences are used for protein expression or in the preparation of mRNA transfection reagents. Therefore, the present invention adjusts the free energy of the optimized sequence. This adjustment includes adjusting the minimum free energy of the sequence to 80% to 100% of the lowest free energy. For example, the minimum free energy of the sequence can be 85% to 95% of the lowest free energy. In specific embodiments, the minimum free energy of the sequence is 80% to 85%, 85% to 90%, 90% to 95%, or 95% to 100% of the lowest free energy.
[0027] In the present invention, the adjustment includes: replacing the codon with the first percentage with the codon with the second, third, and fourth percentages of codon preference, and the percentages of the codon with the first, second, third, and fourth percentages of usage frequency in the full length of the encoding nucleic acid are: 40% to 100%, 0% to 50%, 0% to 25%, and 0% to 15%, respectively.
[0028] In some specific embodiments, the host is a human cell, and the adjusting comprises:
[0029] Among Phe codons, UUU is used at a frequency of 1% to 25%, and UUC is used at a frequency of 75% to 99%;
[0030] Among the codons for Leu, the usage frequency of CUG is 86% to 99%, the usage frequency of CUC is 1% to 10%, the usage frequency of UUA is 0% to 1%, the usage frequency of UUG is 0% to 1%, the usage frequency of CUU is 0% to 1%, and the usage frequency of CUA is 0% to 1%;
[0031] Among the codons of Ile, the usage frequency of AUU is 5% to 25%, the usage frequency of AUC is 75% to 94%, and the usage frequency of AUA is 0% to 1%;
[0032] Among Val codons, the usage frequency of GUU is 0% to 5%, the usage frequency of GUC is 5% to 15%, the usage frequency of GUA is 0% to 1%, and the usage frequency of GUG is 79% to 95%;
[0033] Among the codons for Ser, the usage frequency of UCU is 0% to 20%, the usage frequency of UCC is 10% to 25%, the usage frequency of UCA is 0% to 1%, and the usage frequency of UCG is 0% to 1%; the usage frequency of AGU is 0% to 1%, and the usage frequency of AGC is 52% to 90%;
[0034] Among the codons of Pro, the usage frequency of CCU is 20% to 50%, the usage frequency of CCC is 44% to 72%, the usage frequency of CCA is 0% to 8%, and the usage frequency of CCG is 0% to 6%;
[0035] Among Thr codons, the frequency of ACU is 0% to 1%, the frequency of ACC is 50% to 78%, the frequency of ACA is 20% to 48%, and the frequency of ACG is 0% to 1%;
[0036] Among Ala codons, GCU is used at a frequency of 3% to 20%, GCC at a frequency of 73% to 90%, GCA at a frequency of 0% to 6%, and GCG at a frequency of 0% to 1%;
[0037] Among Tyr codons, UAU is used at a frequency of 1% to 25%, and UAC at a frequency of 75% to 99%;
[0038] Among His codons, the usage frequency of CAU is 1% to 15%, and the usage frequency of CAC is 85% to 99%;
[0039] Among Gln codons, CAA is used at a frequency of 1% to 15%, and CAG at a frequency of 85% to 99%;
[0040] Among Asn codons, the usage frequency of AAU is 5% to 30%, and the usage frequency of AAC is 70% to 95%;
[0041] Among Lys codons, AAA is used at a frequency of 1% to 25%, and AAG at a frequency of 75% to 99%;
[0042] Among the codons of Asp, the usage frequency of GAU is 3% to 35%, and the usage frequency of GAC is 65% to 97%;
[0043] Among the codons for Glu, the usage frequency of GAA is 10% to 35%, and the usage frequency of GAG is 65% to 90%;
[0044] Among Cys codons, the usage frequency of UGU is 15% to 40%, and the usage frequency of UGC is 60% to 85%;
[0045] Among the codons for Arg, the usage frequency of CGU is 0% to 1%, the usage frequency of CGC is 1% to 8%, the usage frequency of CGA is 0% to 1%, the usage frequency of CGG is 44% to 72%, the usage frequency of AGA is 15% to 50%, and the usage frequency of AGG is 1% to 8%;
[0046] Among Gly codons, the usage frequency of GGU is 0% to 5%, the usage frequency of GGC is 65% to 98%, the usage frequency of GGA is 1% to 20%, and the usage frequency of GGG is 1% to 10%.
[0047] In addition, the start codon is AUG; the stop codon is UAA, UAG or UGA, and the codon for Trp is GUU.
[0048] In the scheme of the present invention, the codon usage frequency suitable for the host can be obtained based on the protein expression amount under different codon usage attempts. As mentioned above, the codon usage frequency is obtained after a lot of optimization. Studies have shown that compared with the scheme of codon optimization based on the natural usage frequency of human codons in the prior art, the codon usage frequency in the coding nucleic acid is adjusted to meet the above percentages, which can more effectively improve the protein expression effect and obtain more protein. When the host is a cell from other biological sources, the codon usage frequency is different from the above data and needs to be optimized for a specific species. In the method described in the present invention, the step of obtaining the protein encoding nucleic acid sequence according to the codon preference of the host includes: obtaining the amino acid sequence to be optimized and determining the species direction to be optimized; obtaining the codon usage frequency table of the corresponding species in nature; according to the adjusted codon table, the amino acid sequence to be optimized is randomly distributed and optimized to obtain the encoding nucleic acid sequence.
[0049] In the methods described herein, if, after codon adjustment, the GC content, repeat sequence, secondary structure, and / or free energy still fail to meet the aforementioned criteria, the protein-encoding nucleic acid sequence obtained based on the host's codon preference is adjusted. Such adjustments include lowering the screening criteria and widening the range of codon usage frequencies to ensure a wider range of codon options for subsequent analysis to obtain appropriate results.
[0050] The method of the present invention further includes the steps of translating the optimized sequence and comparing the translated protein sequence with the target protein sequence to determine whether the translated protein sequence is consistent with the target protein sequence.
[0051] Furthermore, the present invention provides a codon optimization system, which includes a module for implementing the above-mentioned method.
[0052] Furthermore, the present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of the above-mentioned method when executed by a processor.
[0053] Furthermore, the present invention also provides a codon-optimized electronic device, which includes a storage medium, a processor, and a computer program stored in the storage medium, wherein the processor executes the computer program to implement the method as described above.
[0054] Furthermore, the present invention also provides a biomaterial comprising at least one of the following I) to V):
[0055] 1) Optimizing the nucleic acid obtained as described above;
[0056] II), an expression unit comprising the nucleic acid described in I);
[0057] III), a plasmid vector containing the nucleic acid described in I), or containing the expression unit described in II);
[0058] IV), mRNA comprising a 5' end cap structure and the nucleic acid described in I), or comprising a 5' end cap structure and the expression unit described in II);
[0059] V) A host cell having the nucleic acid described in I) integrated into its genome, or having the expression unit described in II) integrated into its genome, or having been transformed or transfected with the plasmid vector described in III).
[0060] In the present invention,
[0061] The nucleic acid encoding chicken ovalbumin has a sequence as shown in any one of SEQ ID NOs: 1 to 4;
[0062] or a nucleic acid sequence encoding chicken ovalbumin in which one or more bases are substituted, deleted, added and / or replaced in the nucleic acid sequence shown in any one of SEQ ID NOs: 1 to 4;
[0063] Or a nucleic acid sequence that is at least 80% identical to a nucleic acid sequence as described above.
[0064] This study optimizes the codons of the nucleic acid encoding chicken ovalbumin, yielding four optimized sequences. The resulting nucleic acid sequences demonstrate significantly higher OVA expression levels in hosts compared to wild-type and other control groups. mRNA prepared from these nucleic acids exhibits excellent transfection efficiency, achieving higher expression levels in cells, particularly 293T cells, thereby better mimicking the immune response of the human immune system.
[0065] In the present invention,
[0066] The nucleic acid encoding the Cre recombinase has a sequence as shown in any one of SEQ ID NOs: 10 to 13;
[0067] or a nucleic acid sequence encoding a Cre recombinase in which one or more bases are substituted, deleted, added, and / or replaced in the nucleic acid sequence shown in any one of SEQ ID NOs: 10 to 13;
[0068] Or a nucleic acid sequence that is at least 80% identical to a nucleic acid sequence as described above.
[0069] The present invention performs codon optimization on four versions of Cre recombinase, and the expression level of the resulting Cre recombinase protein is significantly improved compared with the wild type and other control ratios. The mRNA transfection reagent prepared from these nucleic acids can have good transfection efficiency and can be transferred into cells, especially 293T cells, to more effectively exert the function of the Cre recombinase system.
[0070] In the present invention,
[0071] The nucleic acid encoding Gaussia luciferase has a sequence as shown in any one of SEQ ID NOs: 15 to 18;
[0072] or a nucleic acid sequence encoding Gaussia luciferase, wherein one or more bases are substituted, deleted, added, and / or replaced in the nucleic acid sequence shown in any one of SEQ ID NOs: 15 to 18;
[0073] Or a nucleic acid sequence that is at least 80% identical to a nucleic acid sequence as described above.
[0074] Gaussia luciferase expression is low when used in human HEK293T cells. The inventors independently developed an artificial codon optimization method. By optimizing the ratio of codons for the same amino acid, GC content, and sequence repetitiveness in the wild-type Gaussia luciferase DNA sequence, they achieved a DNA sequence that overexpresses Gaussia luciferase. The optimized codons showed expression levels 1.4-16.2 times higher than the wild-type sequence.
[0075] In the present invention,
[0076] The nucleic acid encoding the firefly luciferase has a sequence as shown in any one of SEQ ID NOs: 23 to 24;
[0077] or a nucleic acid sequence encoding a firefly luciferase in which one or more bases are substituted, deleted, added, and / or replaced in the nucleic acid sequence shown in any one of SEQ ID NOs: 23 to 24;
[0078] Or a nucleic acid sequence that is at least 80% identical to a nucleic acid sequence as described above.
[0079] The firefly luciferase protein of the embodiment of the present invention has an expression level 2 to 201 times higher than that of the wild type and other control examples in cell experiments, and has a significant increase in fluorescence brightness in mouse experiments, which greatly enhances the application value of firefly luciferase as a reporter gene.
[0080] In the present invention,
[0081] The nucleic acid encoding SpCas9 has a sequence as shown in any one of SEQ ID NOs: 33 to 37.
[0082] or a nucleic acid sequence encoding SpCas9, wherein one or more bases are substituted, deleted, added, and / or replaced in the nucleic acid sequence shown in any one of SEQ ID NOs: 33 to 37;
[0083] Or a nucleic acid sequence that is at least 80% identical to a nucleic acid sequence as described above.
[0084] The CRISPR / Cas9 protein expression level provided in the technical solution of the present invention is increased by 4-8 times compared with the wild type or other control ratios, and has a higher gene editing efficiency, thus greatly enhancing the application value of Cas9 protein in gene editing.
[0085] In the present invention, the at least 80% identity includes: at least 85% identity, at least 90% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, at least 99% identity, at least 99.5% identity, at least 99.6% identity, at least 99.7% identity, at least 99.8% identity or at least 99.9% identity.
[0086] Compared with the wild type, the expression level of the sequence optimized in the embodiment of the present invention is significantly improved. Compared with other optimization schemes in the prior art, the expression level of the sequence optimized in the embodiment of the present invention is also significantly improved.
[0087] The expression unit provided by the present invention comprises the nucleic acid and promoter as described above.
[0088] For different expression vectors, the expression unit of the present invention may further include other elements, such as enhancers, terminators and / or nuclear localization signals.
[0089] In some embodiments, the promoter is selected from a eukaryotic promoter and a prokaryotic promoter, which is not limited in the present invention. For example, the promoter is a T7 promoter, a CMV promoter, a CAG promoter, an EF1a promoter, a PGK promoter, a U6 and H1 promoter, an EFS promoter, a CBh promoter, a SFFV promoter, a MSCV promoter, a SV40 promoter, a UBC promoter, or a TRE promoter.
[0090] In some embodiments, the enhancer is selected from SV40 enhancer, CMV enhancer, SV-1 enhancer, ROSA26 enhancer, EF1α enhancer, HARE5 enhancer, UBC enhancer, EF1A enhancer, PGK enhancer, CAGG enhancer, COPIA enhancer, or ACT5C enhancer.
[0091] In some embodiments, the terminator is selected from T7 phage terminator, TO phage terminator, λ phage terminator, SV40 terminator, CMV terminator, rrnB terminator, bGH terminator, hGH terminator, or rbGlob terminator.
[0092] In some embodiments, the nuclear localization signal is SV40 NLS, whose amino acid sequence is PKKKRKV, and the nuclear localization signal is located between the promoter and the nucleic acid as described above.
[0093] In some embodiments, the expression unit comprises, from the 5' end to the 3' end, a promoter, a 5' UTR, a Kozak sequence, a nucleic acid as described above, a 3' UTR, a poly A tail, and a restriction enzyme site. In some specific embodiments, the poly A tail is 110 bp long.
[0094] The plasmid vector provided by the present invention comprises a backbone vector and at least one of the following i) to ii):
[0095] i) the nucleic acid as described above;
[0096] ii) an expression unit as described above.
[0097] The plasmid vector described in the present invention is used for the amplification, preservation, transformation and / or transfection of nucleic acid, and the present invention does not limit this. The plasmid vector described in the present invention may be circular or linear, and the present invention does not limit this. The present invention does not limit the insertion site of the nucleic acid and / or expression unit in the plasmid vector. Preferably, it can be inserted at the multiple cloning site of the plasmid vector or other regions. In some embodiments, it is a cloning vector, an expression vector, a herpes simplex virus vector, a retroviral vector, a lentiviral vector, an adenoviral vector or an adeno-associated virus vector. For example, the cloning vector is a pUC series plasmid vector, a pBR322 plasmid vector, a pGEM series plasmid vector, a pET series plasmid vector, a Yeast series plasmid vector or a Gateway plasmid vector, etc. For example, the adeno-associated virus vector is an rAAV vector, and its serotype includes AAV1, AAV2, AAV5, AAV6, AAV8 or AAV9. In some embodiments, the pmRVac retroviral vector is used as an example for expression verification, and its effect is better than other vectors.
[0098] Furthermore, the present invention provides an mRNA comprising a 5' end cap structure and the aforementioned nucleic acid, or comprising a 5' end cap structure and the aforementioned expression unit.
[0099] In the present invention, the mRNA includes a 5'UTR, a Kozak sequence, the aforementioned nucleic acid, a 3'UTR, a polyA tail, and a restriction enzyme cleavage site, which are sequentially connected.
[0100] In the present invention, the sequences of the 5'UTR, 3'UTR or Kozak are not limited.
[0101] The length of the polyA tail is 50 to 200 bp, for example, 50 bp, 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp or 200 bp. Alternatively, the length of the polyA tail can be any intermediate value between the first two values.
[0102] The sequence of the restriction enzyme site is repeated once in the plasmid vector as described above. In the embodiment of the invention, Sap I is used as the restriction enzyme site for plasmid linearization. In addition, any other restriction endonuclease site that can linearize the plasmid can be used.
[0103] Furthermore, the mRNA preparation method includes linearizing the plasmid vector as described above, capping, and purifying the mRNA.
[0104] Specifically, the mRNA preparation method includes digesting the plasmid vector as described above with a restriction endonuclease, adding a capping enzyme and other raw materials required for capping, and performing a capping reaction at 37°C for 1 hour; then obtaining an mRNA sample after chromatography purification, sterilizing and filtering for later use.
[0105] Furthermore, the present invention also provides a host cell, which is transformed or transfected with the aforementioned plasmid vector, or has the aforementioned nucleic acid integrated into its genome, or has the aforementioned expression unit integrated into its genome.
[0106] In the present invention, the host is a human cell or a mammalian cell. Alternatively, the host may be a prokaryotic microorganism or a eukaryotic microorganism. Eukaryotic hosts include, but are not limited to, yeast and insect cells, and prokaryotic hosts include, but are not limited to, Escherichia coli. For example, the host is E. coli BL21(DE3), BL21(DE3)pLysS, DH5α, JM109, JM110, TOP10, HB101, or XL1-Blue. The human cell is a 293T cell.
[0107] Furthermore, the present invention also provides a method for constructing a host cell, which comprises treating the cell with the mRNA and lipids as described above. In the embodiments of the present invention, the treatment comprises transfection and / or transformation.
[0108] The present invention also provides a product obtained by culturing the host cell as described above.
[0109] Furthermore, the present invention also provides a method for preparing a recombinant protein, which comprises:
[0110] After obtaining the nucleic acid molecule optimized by the method described above, the nucleic acid molecule is introduced into the host cell, and then cultured to obtain a product containing the recombinant protein.
[0111] Furthermore, the present invention also provides a method for preparing an mRNA transfection preparation or vaccine, which comprises: obtaining a nucleic acid molecule optimized by the method described above, connecting the nucleic acid molecule to a vector, linearizing it, capping it, purifying it, and encapsulating it to obtain an mRNA transfection preparation or vaccine.
[0112] The present invention provides a codon optimization method that replaces codons based on parameters such as codon type and number, local GC content, local repetitive sequences, mRNA secondary structure, and mRNA free energy. The optimized sequence can be experimentally screened to obtain a DNA sequence that highly expresses proteins. The downstream application of the present method is mainly in the direction of IVT mRNA: the optimized DNA sequence is constructed into a vector, and the linearized plasmid is used as a template for in vitro transcription of mRNA. The mRNA is then encapsulated into LNPs or LNPs coupled to other molecules, resulting in significant protein expression improvement in 293T cells (not limited to 293T cells) or mice. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 shows a schematic diagram of the standard codon representation;
[0114] Figure 2 shows the schematic representation of the codons for the species "Homo sapiens";
[0115] FIG3 shows a GC content analysis diagram of the first version of the DNA sequence in Example 1;
[0116] FIG4 is a schematic diagram showing the analysis of the duplication of the first version of the DNA sequence in Example 1;
[0117] FIG5 is a schematic diagram showing the RNA secondary structure of the first version of the DNA sequence in Example 1;
[0118] FIG6 shows a GC content analysis chart of OVA (1) in Example 1;
[0119] FIG7 is a schematic diagram showing the sequence duplication analysis of OVA (1) in Example 1;
[0120] Figure 8 shows a schematic diagram of the RNA secondary structure of OVA (1) in Example 1;
[0121] FIG9 shows a GC content analysis chart of OVA (2) in Example 1;
[0122] Figure 10 is a schematic diagram showing the sequence duplication analysis of OVA (2) in Example 1;
[0123] Figure 11 shows a schematic diagram of the RNA secondary structure of OVA (2) in Example 1;
[0124] FIG12 shows the GC content analysis of OVA (3) in Example 1;
[0125] Figure 13 is a schematic diagram showing the sequence duplication analysis of OVA (3) in Example 1;
[0126] Figure 14 shows a schematic diagram of the RNA secondary structure of OVA (3) in Example 1;
[0127] Figure 15 shows the GC content analysis of OVA (4) in Example 1;
[0128] Figure 16 is a schematic diagram showing the sequence duplication analysis of OVA (4) in Example 1;
[0129] Figure 17 shows a schematic diagram of the RNA secondary structure of OVA (4) in Example 1;
[0130] FIG18 shows the GC content ratio analysis results of the first version sequence in Example 2;
[0131] FIG19 shows the dot plot analysis results of the first version sequence duplication in Example 2;
[0132] Figure 20 shows the secondary structure of the RNA sequence corresponding to the first version sequence in Example 2;
[0133] FIG21 shows the GC content ratio analysis results of the Co1 sequence in Example 2;
[0134] FIG22 shows the dot plot analysis results of the Co1 sequence duplication in Example 2;
[0135] Figure 23 shows the secondary structure of the RNA sequence corresponding to the Col sequence in Example 2;
[0136] FIG24 shows the GC content ratio analysis results of the Co2 sequence in Example 2;
[0137] FIG25 shows the results of dot plot analysis of the Co2 sequence duplication in Example 2;
[0138] Figure 26 shows the secondary structure of the RNA sequence corresponding to the Co2 sequence in Example 2;
[0139] FIG27 shows the GC content ratio analysis results of the Co3 sequence in Example 2;
[0140] FIG28 shows the results of a dot plot analysis of the Co3 sequence duplication in Example 2;
[0141] Figure 29 shows the secondary structure of the RNA sequence corresponding to the Co3 sequence in Example 2;
[0142] FIG30 shows the GC content ratio analysis results of the Co4 sequence in Example 2;
[0143] FIG31 shows the results of dot plot analysis of the Co4 sequence duplication in Example 2;
[0144] Figure 32 shows the secondary structure of the RNA sequence corresponding to the Co4 sequence in Example 2;
[0145] FIG33 is a graph showing the GC content analysis of the GLuc1 DNA sequence in Example 3;
[0146] FIG34 is a diagram showing sequence duplication analysis of the GLuc1 DNA sequence in Example 3;
[0147] FIG35 is a schematic diagram of the RNA secondary structure of the GLuc1 DNA sequence in Example 3;
[0148] FIG36 is a graph showing the GC content analysis of the GLuc2 DNA sequence in Example 3;
[0149] FIG37 is a diagram showing sequence duplication analysis of the GLuc2 DNA sequence in Example 3;
[0150] Figure 38 is a schematic diagram of the RNA secondary structure of the GLuc2 DNA sequence in Example 3;
[0151] FIG39 is a graph showing the GC content analysis of the GLuc3 DNA sequence in Example 3;
[0152] FIG40 is a diagram showing the sequence duplication analysis of the GLuc3 DNA sequence in Example 3;
[0153] FIG41 is a schematic diagram of the RNA secondary structure of the GLuc3 DNA sequence in Example 3;
[0154] FIG42 is a graph showing the GC content analysis of the GLuc4 DNA sequence in Example 3;
[0155] FIG43 is a diagram showing sequence duplication analysis of the GLuc4 DNA sequence in Example 3;
[0156] Figure 44 is a schematic diagram of the RNA secondary structure of the GLuc4 DNA sequence in Example 3;
[0157] FIG45 is a schematic diagram showing the GC content analysis of the first version of the luciferase DNA sequence in Example 4;
[0158] FIG46 is a schematic diagram showing the sequence duplication analysis of the first version of the luciferase DNA sequence in Example 4;
[0159] FIG47 is a schematic diagram showing the RNA secondary structure of the first version of the luciferase DNA sequence in Example 4;
[0160] FIG48 is a schematic diagram showing the GC content analysis of LUC-1 in Example 4;
[0161] FIG49 is a schematic diagram showing the sequence duplication analysis of LUC-1 in Example 4;
[0162] Figure 50 is a schematic diagram of the RNA secondary structure of LUC-1 in Example 4;
[0163] FIG51 is a schematic diagram showing the GC content analysis of LUC-2 in Example 4;
[0164] FIG52 is a schematic diagram showing the sequence duplication analysis of LUC-2 in Example 4;
[0165] Figure 53 is a schematic diagram of the RNA secondary structure of LUC-2 in Example 4;
[0166] Figure 54 is a schematic diagram of DNA sequence analysis of WT in Example 5;
[0167] Figure 55 is a schematic diagram of the DNA sequence analysis of VB1 in Example 5;
[0168] Figure 56 is a schematic diagram of the DNA sequence analysis of VB2 in Example 5;
[0169] Figure 57 is a schematic diagram of DNA sequence analysis of VB3 in Example 5;
[0170] Figure 58 is a schematic diagram of the DNA sequence analysis of VB4 in Example 5;
[0171] Figure 59 is a schematic diagram of DNA sequence analysis of VB5 in Example 5;
[0172] Figure 60 shows the pmRVac vector map
[0173] Figure 61 shows western blot analysis of cell expression in each group in Example 1;
[0174] Figure 62 shows SDS-PAGE electrophoresis images of cells expressed in each group in Example 1;
[0175] Figure 63 shows grayscale analysis of western blot gel images after cell expression in each group of Example 1 (****P<0.0001);
[0176] In Figure 64, 64-1 shows the WB experimental results (NC: Negative Control), and 64-2 shows the analysis of CreRecombinase expression levels in different samples (WB image band gray value analysis, ***P<0.001, ****P<0.0001);
[0177] In FIG65 , 65-1 shows the results of tdTomato fluorescence imaging of a liver slice in vivo (magnification 200x, exposure 5ms, gain 5x), and 65-2 shows the results of tdTomato fluorescence imaging of a spleen slice in vivo (magnification 200x, exposure 5ms, gain 5x);
[0178] FIG66 is a graph showing the expression intensity of four Gaussia luciferase genes in HEK293T cells in Example 3;
[0179] Figure 67 is a schematic diagram of the firefly luciferase gene cell experiment data in Example 4; columns 1 to 9 respectively represent the expression levels of control sequence 1, control sequence 2, control sequence 3, control sequence 4, control sequence 5, control sequence 6, control sequence 7, LUC-1, and LUC-2 nucleic acids in the cells;
[0180] Figure 68 shows the in vivo imaging results 6 hours after injection of LNP-mRNA in Example 4; wherein: the left side shows the 10 μg dose group; the right side shows the 30 μg dose group, from left to right: control sequence 1, control sequence 3, control sequence 5, LUC-1 and LUC-2;
[0181] Figure 69 shows the in vivo imaging results 24 hours after injection of LNP-mRNA in Example 4; wherein: the left side shows the 10 μg dose group; the right side shows the 30 μg dose group, from left to right: control sequence 1, control sequence 3, control sequence 5, LUC-1 and LUC-2;
[0182] Figure 70 shows the in vivo imaging results 48 hours after injection of LNP-mRNA in Example 4; wherein: the left shows the 10 μg dose group; the right shows the 30 μg dose group, from left to right: control sequence 1, control sequence 3, control sequence 5, LUC-1 and LUC-2;
[0183] Figure 71 shows the in vivo imaging results 72 hours after injection of LNP-mRNA in Example 4; wherein: the left side shows the 10 μg dose group; the right side shows the 30 μg dose group, from left to right: control sequence 1, control sequence 3, control sequence 5, LUC-1 and LUC-2;
[0184] Figure 72 is an immunoblot of the Cas9 protein expressed in Example 5;
[0185] Figure 73 is an analysis of the relative expression levels of the Cas9 protein expressed in Example 5;
[0186] FIG74 is a comparison of the functional verification of the Cas9 protein expressed in cells in Example 5;
[0187] Figure 75 is a visual data analysis diagram comparing the functional verification of Cas9 proteins expressed in cells obtained in Example 5. DETAILED DESCRIPTION
[0188] The present invention provides methods for codon optimization, and those skilled in the art can refer to the contents herein and appropriately improve the process parameters to achieve the desired results. It should be noted that all similar substitutions and modifications will be apparent to those skilled in the art and are considered to be included in the present invention. The methods and applications of the present invention have been described through preferred embodiments, and it is apparent that those skilled in the art can modify or appropriately alter and combine the methods and applications herein to implement and apply the technology of the present invention without departing from the content, spirit, and scope of the present invention.
[0189] Unless otherwise defined herein, scientific and technical terms related to the present invention shall have the meanings that are understood by those of ordinary skill in the art.
[0190] In this application, "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0191] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items.
[0192] In this application, "including," "comprising," and "having" are used interchangeably to indicate the inclusiveness of a solution, meaning that the solution may contain other elements in addition to the listed elements. It should also be understood that the use of "including," "comprising," and "having" in this document also provides solutions that "consist of" or "as shown in."
[0193] In the present invention, the term "codon" refers to a triplet nucleotide residue sequence on RNA (or DNA) that encodes a specific amino acid. The most common start codons are methionine or valine codons.
[0194] In this context, the term "codon table" refers to the traditional representation of the genetic code in the form of an RNA codon table. This is because when cellular ribosomes produce proteins, messenger RNA (mRNA) directs protein synthesis. The sequence of mRNA is determined by genomic DNA. With the rise of computational biology and genomics, most genes can now be discovered at the DNA level, making DNA codon tables increasingly useful.
[0195] As used herein, "identity" can be calculated by determining the percent "identity" of two amino acid sequences or two nucleic acid sequences by aligning the sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of the first and second amino acid sequences or nucleic acid sequences for optimal alignment, or non-homologous sequences can be discarded for comparison purposes). The amino acid residues or nucleotides at corresponding amino acid positions or nucleotide positions are then compared. When a position in the first sequence is occupied by the same amino acid residue or nucleotide at the corresponding position in the second sequence, the molecules are identical at that position.
[0196] The percent identity between the two sequences will vary with the number of identical positions shared by the sequences, taking into account the number of gaps introduced and the length of each gap for optimal alignment of the two sequences. In the present invention, the expression "80% or greater identity" includes: 85% or greater, 90% or greater, 95% or greater, 96% or greater, 97% or greater, 98% or greater, 99% or greater, 99.5% or greater, 99.6% or greater, 99.7% or greater, 99.8% or greater, or 99.9% or greater.
[0197] In this application, "nucleic acid" includes any compound and / or substance that comprises a polymer of nucleotides. Each nucleotide is composed of a base, particularly a purine or pyrimidine base (i.e., cytosine (C), guanine (G), adenine (A), thymine (T), or uracil (U)), a sugar (i.e., deoxyribose or ribose) and a phosphate group. Typically, nucleic acid molecules are described by a sequence of bases, whereby the bases represent the primary structure (linear structure) of the nucleic acid molecule. The sequence of bases is typically expressed as 5' to 3'.
[0198] In the present application, the nucleic acid molecules encompass deoxyribonucleic acid (DNA), including, for example, complementary DNA (cDNA) and genomic DNA, ribonucleic acid (RNA), particularly synthetic forms of messenger RNA (mRNA), DNA or RNA, and polymers comprising a mixture of two or more of these molecules. Nucleic acid molecules can be linear or cyclic. In addition, the term nucleic acid molecule includes both sense and antisense strands, as well as single-stranded and double-stranded forms. Moreover, the nucleic acid molecules described herein can contain naturally occurring or non-naturally occurring nucleotides. Examples of non-naturally occurring nucleotides include modified nucleotide bases with derived sugar or phosphate backbone linkages or chemically modified residues.
[0199] As used herein, "plasmid vector" refers to a nucleic acid molecule capable of amplifying another nucleic acid to which it is linked. The term includes vectors that are self-replicating nucleic acid structures as well as vectors that integrate into the genome of a host cell into which the vector has been introduced.
[0200] In this application, "host cell" refers to a cell into which an exogenous nucleic acid has been introduced, including the offspring of such a cell. Host cells include "transformants" and "transformed cells," which include the original transformed cell and the offspring derived therefrom, without considering the number of passages. Offspring may not be completely identical to the parent cell in nucleic acid content, but may contain mutations. Mutant offspring having the same function or biological activity as that screened or selected in the initially transformed cell are included herein.
[0201] In addition, the numerical ranges and parameters used to define the present invention are approximate values. The relevant numerical values in the specific examples have been presented as accurately as possible. However, any numerical value inherently inevitably contains standard deviations due to individual testing methods. Therefore, unless otherwise expressly stated, all ranges, amounts, values, and percentages used in this disclosure should be understood to be modified by the word "about." As used herein, "about" generally means that the actual value is within plus or minus 10%, 5%, 1%, or 0.5% of a specified value or range.
[0202] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. Some or all of the steps can be executed in parallel or sequentially. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0203] The test materials used in the present invention are all common commercial products and can be purchased in the market.
[0204] The codon table involved in the embodiment is shown in Table 1:
[0205] Table 1
[0206] Taking human cells as the host, the frequency of human codon usage is as follows:
[0207] Table 2
[0208] The term "CAI" as used herein refers to the codon usage frequency of highly expressed genes in a particular species. CAI values range from 0 to 1. A high CAI value for a target gene indicates that it is more efficiently expressed in that species.
[0209] Conventional IT algorithms only adjust parameters, such as raising the CAI value or controlling the GC content value at around 60%, and the calculation process is carried out by issuing instructions one by one. Artificial intelligence will make a personalized judgment based on the entire sequence, and then selectively choose to lower or raise a certain value, allowing for the existence of local values that do not meet the IT algorithm parameters. At the same time, artificial intelligence can also determine whether the sequence needs to be optimized, and it does not require further optimization for the optimized sequence. Artificial intelligence can provide multiple optimized versions, and then screen out the better version based on the experimental results. The IT algorithm generally only provides a codon-optimized version sequence. The relevant optimization data of the present invention show that for some verified genes, the protein expression level has increased by several times or even hundreds of times.
[0210] The sequences involved in the embodiment are as follows:
[0211] The test materials used in the present invention are all common commercial products and can be purchased on the market. The present invention is further described below with reference to the following examples:
[0212] Example 1 Chicken Ovalbumin (OVA) Codon Optimization
[0213] The optimized sequence was derived from the wild-type CDS region sequence of the species Gallus gallus.
[0214] Related information
[0215] Sequence source: https: / / www.ncbi.nlm.nih.gov / nuccore / 2099367237
[0216] The DNA sequence optimization method of the present invention is as follows.
[0217] (1) Amino acid sequence analysis
[0218] The present invention relates to a wild-type version of chicken ovalbumin amino acid sequence, as shown in Table 3.
[0219] Table 3 Chicken ovalbumin amino acid count
[0220] Amino acid sequence 1: MGSIGAASMEFCFDVFKELKVHHANENIFYCPIAIMSALAMVYLGAKDSTRTQINKVVRFDKLPGFGDSIEAQCGTSVNVHSSLRDILNQITKPNDVYSFSLASRLYAEERYPILPEYLQCVKELYRGGLEPINFQTAADQARELINSWVESQTNGIIRNVLQPSSVDSQTAMVLVNAIVFKGLWEKAFK DEDTQAMPFRVTEQESKPVQMMYQIGLFRVASMASEKMKILELPFASGTMSMLVLLPDEVSGLEQLESIINFEKLTEWTSSNVMEERKIKVYLPRMKMEEKYNLTSVLMAMGITDVFSSSANLSGISSAESLKISQAVHAAHAEINEAGREVVGSAEAGVDAASVSEEFRADHPFLFCIKHIATNAVLFFGRCVSP*.
[0221] (2) Conversion of amino acid sequence into DNA sequence
[0222] Convert the amino acid sequence to the most commonly used codon sequence for the species to be used. For example, for Homo sapiens, replace all "alanine" with "GCC." You can find the most commonly used codon sequence for each species at the following website (https: / / www.kazusa.or.jp / codon / cgi-bin / showcodon.cgi?species=9606).
[0223] After following the above procedure, the first version of the DNA sequence (SEQ ID NO: 7) was obtained:
[0224] (3) DNA sequence analysis
[0225] The GC content of the first version of the DNA sequence obtained in step 2 was analyzed, as shown in FIG3 .
[0226] GC content analysis tool URL: https: / / www.vectorbuilder.cn / tool / gc-content-calculator.html
[0227] The first version of the DNA sequence obtained in step 2 was analyzed for sequence duplication, as shown in FIG4 .
[0228] Sequence duplication analysis tool URL: https: / / www.vectorbuilder.cn / tool / sequence-dot-plot.html
[0229] The first version of the DNA sequence was subjected to sequence RNA secondary structure and free energy analysis. The minimum free energy (MFE) of the RNA sequence was -475.50 kcal / mol, and the RNA secondary structure was shown in FIG5 .
[0230] RNA secondary structure and free energy analysis tool URL: http: / / rna.tbi.univie.ac.at / / cgi-bin / RNAWebSuite / RNAfold.cgi
[0231] (4) DNA sequence adjustment
[0232] By combining parameters such as codon preference, GC content ratio, sequence repetition, RNA secondary structure, and RNA free energy, an optimized version of the DNA sequence is obtained.
[0233] The optimization of each parameter is described as follows.
[0234] Codon bias: After obtaining the first version of the DNA sequence based on the most frequently used codons, the codons with the second, third, and fourth highest frequencies are replaced, respectively. The principle is to control the proportion of the codons with the first, second, third, and fourth highest frequencies in the protein sequence to be within the following ranges: 40-100%, 0-50%, 0-25%, and 0-15%. As shown in Figure 1 (https: / / zh.wikipedia.org / wiki / DNA%E5%AF%86%E7%A0%81%E5%AD%90%E8%A1%A8), the codons corresponding to amino acids are identified using the standard codon table. This is combined with the codon table for Homo sapiens (Figure 2) to determine the codon bias. In Homo sapiens, the codons for glycine GGC account for 40-100% of the time, GGA for 0-50%, GGG for 0-25%, and GGT for 0-15%.
[0235] 1) GC content: GC content should be higher than the 0% to 15% GC content of the species. For example, the GC content of Homo sapiens is 52.27%, so the optimized DNA sequence should have a GC content within the range of 52.27% to 67.27%. GC content varies across species. For example, the GC content of Macaca fascicularis is 49.64%, so the optimized DNA sequence should have a GC content within the range of 49.64% to 64.64%. The GC content of the DNA sequence should be controlled between 30% and 95%.
[0236] 2) Sequence repetition: Optimize the number of repeated sequences and sequence length to minimize the number of repeated sequences. For example, replace a codon in a sequence segment with a repeated sequence greater than 20 bases with a synonymous codon to reduce the number of bases in the repeated sequence.
[0237] 3) RNA secondary structure and free energy: Reduce the continuous base-pairing region to lower the minimum free energy. If both cannot be achieved simultaneously, prioritize reducing the continuous base-pairing region.
[0238] According to the above optimization method, the optimized sequences 1 to 4 of OVA were obtained, and the free energies were: -446.09 kcal / mol, -444.51 kcal / mol, -450.96 kcal / mol, and -445.90 kcal / mol, respectively.
[0239] The proportions of the first, second, third, and fourth codons of OVA (1) (SEQ ID NO: 1) in the total sequence of the protein range from 44 to 100%, 0 to 50%, 0 to 14%, and 0 to 3%. The GC content analysis is shown in Figure 6 , the sequence repetition is shown in Figure 7 , and the RNA secondary structure is shown in Figure 8 .
[0240] The proportions of the first, second, third, and fourth codons of OVA (2) (SEQ ID NO: 2) in the total sequence of the protein range from 44 to 100%, 0 to 50%, 0 to 14%, and 0 to 3%. The GC content analysis is shown in Figure 9 , the sequence repetition is shown in Figure 10 , and the RNA secondary structure is shown in Figure 11 .
[0241] The proportions of the first, second, third, and fourth codons of OVA (3) (SEQ ID NO: 3) in the total sequence of the protein range from 44 to 100%, 0 to 50%, 0 to 14%, and 0 to 3%. The GC content analysis is shown in Figure 12 , the sequence repetition is shown in Figure 13 , and the RNA secondary structure is shown in Figure 14 .
[0242] The proportions of the first, second, third, and fourth codons of OVA (4) (SEQ ID NO: 4) in the total sequence of the protein range from 44 to 100%, 0 to 50%, 0 to 14%, and 0 to 3%. The GC content analysis is shown in Figure 15 , the sequence repetition is shown in Figure 16 , and the RNA secondary structure is shown in Figure 17 .
[0243] Table 4 Sequence comparison information of this experiment
[0244] Example 2 Codon Optimization of Cre Recombinase (Cre)
[0245] Four versions of the DNA sequence encoding Cre recombinase were artificially codon-optimized: Co1 is shown in SEQ ID NO: 1, Co2 is shown in SEQ ID NO: 2, Co3 is shown in SEQ ID NO: 3, and Co4 is shown in SEQ ID NO: 4. The wild-type (WT) version of the Cre enzyme (with the SV40 NLS amino acid sequence (PKKKRKV) following the promoter to facilitate nuclear recombination of DNA fragments) is shown in SEQ ID NO: 5. The nucleic acid sequence encoding the Cre recombinase optimized by other companies is shown in SEQ ID NO: 6.
[0246] The DNA sequence optimization method of the present invention is as follows.
[0247] (1) Amino acid sequence analysis
[0248] Table 5 Amino acid sequence analysis
[0249] Amino acid sequence:
[0250] (2) Convert amino acid sequence to DNA sequence
[0251] The entire amino acid sequence is decoded into the most commonly used codon sequence for the species to be used (the reference website for the codon usage table is https: / / www.kazusa.or.jp / codon / cgi-bin / showcodon.cgi?species=9606, as shown in Figure 2). For example, for the species "Homo sapiens", all "glycine" is decoded as "GGC", and the codon sequences corresponding to other amino acids are decoded in sequence to obtain the first version of the DNA sequence.
[0252] DNA sequence of the first version (SEQ ID NO: 14):
[0253] (3) DNA sequence analysis
[0254] The first version of the DNA sequence was subjected to GC content ratio analysis (GC content analysis website is https: / / www.vectorbuilder.cn / tool / gc-content-calculator.html), as shown in FIG18 .
[0255] The first version of the DNA sequence was subjected to sequence duplication dot plot analysis (the sequence duplication dot plot analysis website is https: / / www.vectorbuilder.cn / tool / sequence-dot-plot.html), as shown in FIG19 .
[0256] The first version of the DNA sequence was subjected to RNA sequence free energy and secondary structure analysis (analysis URL: http: / / rna.tbi.univie.ac.at / / cgi-bin / RNAWebSuite / RNAfold.cgi). The minimum free energy (MFE) of the RNA sequence was -490.9 kcal / mol, and the RNA secondary structure is shown in FIG20 .
[0257] (4) DNA sequence adjustment
[0258] The DNA sequence is adjusted based on parameters such as codon preference, sequence repetition, GC content ratio, RNA free energy, and RNA secondary structure to obtain an optimized version of the DNA sequence.
[0259] The optimization of each parameter is described as follows.
[0260] 1) Codon Bias: After obtaining the first version of the DNA sequence based on the most frequently used codons, the codons with the second, third, and fourth highest percentages are replaced with the codons with the highest percentages. The general principle is to control the percentages of the codons with the highest, second, third, and fourth highest percentages in the total sequence to be within the following ranges: 40-100%, 0-50%, 0-25%, and 0-15%. As shown in Figure 1, the codons corresponding to the amino acids are identified using the standard codon table, and then combined with the codon table for Homo sapiens in Figure 2 to determine the codon bias.
[0261] GC content: The GC content should be higher than the GC content of the species (0% to 15%). For Homo sapiens, the GC content is 52.27%, so the optimized DNA sequence should have a GC content between 52.27% and 67.27%. GC content varies across species. Additionally, the GC content of a DNA sequence should be controlled between 30% and 95%.
[0262] 3) Sequence repetition: Optimize the number of repeated sequences and sequence length to minimize the number of repeated sequences. For example, replace a codon in a sequence segment with a repeated sequence greater than 20 bases with a synonymous codon to reduce the number of bases in the repeated sequence.
[0263] 4) RNA secondary structure and free energy: Reduce the continuous base pairing region and lower the minimum free energy. If both cannot be achieved simultaneously, reduce the continuous base pairing region first.
[0264] According to the above optimization method, the optimized Cre recombinase coding sequences were obtained and recorded as Co1-4.
[0265] As shown in Figures 21 to 23, the codon usage percentages of Co1 are as follows: the first, second, third, and fourth codons in the total sequence range from 68% to 100%, 0% to 29%, 0% to 7%, and 0% to 7%. The GC content of Co1 is 64.5%. The minimum free energy of the Co1 RNA sequence is -447.70 kcal / mol.
[0266] As shown in Figures 24 to 26, the codon usage percentages of the first, second, third, and fourth codons in the Co2 sequence range from 44% to 100%, 0% to 50%, 0% to 14%, and 0% to 3%, respectively. The GC content of the Co2 sequence is 60.4%. The minimum free energy of the Co2 RNA sequence is -406.70 kcal / mol.
[0267] As shown in Figures 27 to 29, the codon usage percentages of the first, second, third, and fourth codons in the Co3 sequence range from 68% to 100%, 0% to 25%, 0% to 10%, and 0% to 7%, respectively. The GC content of the Co3 sequence is 65.8%. The minimum free energy of the Co3 RNA sequence is -461.80 kcal / mol.
[0268] As shown in Figures 30 to 32, the codon usage percentages of Co4 are as follows: the first, second, third, and fourth codons in the total sequence range from 50% to 100%, 0% to 33%, 0% to 9%, and 0% to 3%. The GC content of the Co4 sequence is 62.5%. The minimum free energy of the Co4 sequence RNA is -424.50 kcal / mol.
[0269] Table 6. Sequence information alignment of this experiment
[0270] Example 3 Codon Optimization of Gaussia Luciferase
[0271] The wild-type CDS sequence from Gaussia princeps was used as experimental material (Sequence ID: AY015993.1). Four versions of the Gaussia luciferase target gene sequence were optimized using an independently developed artificial codon optimization method. The DNA sequence optimization method in the present invention is as follows.
[0272] (1) Amino acid sequence analysis
[0273] The amino acid sequence of the Gaussia luciferase protein involved in the present invention is shown in Table 7.
[0274] Table 7
[0275] Amino acid sequence:
[0276] (2) Conversion of amino acid sequence to DNA sequence
[0277] Convert all amino acid sequences to the most commonly used codon sequence for the species to be used, as shown in Figure 2. The most commonly used codon sequence can be obtained from the following website:
[0278] https: / / www.kazusa.or.jp / codon / cgi-bin / showcodon.cgi?species=9606.
[0279] For example, for the species "Homo sapiens," all "alanine" would be replaced with "GCC." This yields the first version of the DNA sequence.
[0280] First version of the DNA sequence:
[0281] (3) DNA sequence analysis
[0282] [Corrected 25.10.2024 in accordance with Rule 91] The first version of the DNA sequence was analyzed for GC content, as shown in Figure 39.
[0283] GC content ratio analysis link:
[0284] https: / / www.vectorbuilder.cn / tool / gc-content-calculator.html
[0285] [Corrected 25.10.2024 in accordance with Rule 91] The first version of the DNA sequence was analyzed for sequence duplications, as shown in Figure 40.
[0286] DNA sequence duplication analysis link:
[0287] https: / / www.vectorbuilder.cn / tool / sequence-dot-plot.html
[0288] [Corrected 25.10.2024 according to Rule 91] The first version of the DNA sequence was subjected to RNA secondary structure and free energy analysis. The minimum free energy (MFE) of the RNA sequence was -241.70 kcal / mol, and the RNA secondary structure is shown in Figure 41.
[0289] RNA secondary structure and free energy analysis link:
[0290] http: / / rna.tbi.univie.ac.at / / cgi-bin / RNAWebSuite / RNAfold.cgi
[0291] (4) DNA sequence adjustment
[0292] By combining parameters such as codon preference, GC content ratio, sequence repetition, RNA secondary structure, and RNA free energy, an optimized version of the DNA sequence is obtained.
[0293] The optimization of each parameter is described below.
[0294] 1) Codon Bias: After obtaining the first version of the DNA sequence based on the most frequently used codons, the codons with the second, third, and fourth highest usage percentages are used to replace the codons with the highest usage percentages. As a general rule, the percentages of the codons with the highest, second, third, and fourth highest usage percentages in the protein's total sequence are controlled within the following ranges: 40-100%, 0-50%, 0-25%, and 0-15%. As shown in Figure 1, the codons corresponding to the amino acids are identified using the standard codon table. This is combined with the codon table for the Homo sapiens species shown in Figure 2 to determine the codon bias. In Homo sapiens, the amino acid GGC accounts for 40-100% of the time, GGA for 0-50%, GGG for 0-25%, and GGT for 0-15%.
[0295] Standard codon table link URL:
[0296] https: / / zh.wikipedia.org / wiki / DNA%E5%AF%86%E7%A0%81%E5%AD%90%E8%A1%A8
[0297] 2) GC content: GC content should be higher than the 0% to 15% GC content of the species. For example, the GC content of Homo sapiens is 52.27%, so the optimized DNA sequence should have a GC content between 52.27% and 67.27%. GC content varies across species. For example, the GC content of Macaca fascicularis is 49.64%, so the optimized DNA sequence should have a GC content between 49.64% and 64.64%. The GC content of a DNA sequence should be controlled between 30% and 95%.
[0298] 3) Sequence repetitions: Optimize the number and length of repeats to minimize the number of repeats. For example, replace a codon in a sequence segment with a repeat length greater than 20 bases with a synonymous codon to reduce the number of repeat bases. For example, the repeat sequence "ATGGAGGACGCCAAGAACATCAAG" before optimization has a total of 24 bases and occurs three times. By mutating the codons at two different positions to synonymous codons, two different DNA sequences are obtained: "ATGGAAGACGCCAAGAACATCAAG" and "ATGGAGGATGCCAAGAACATCAAG." The amino acid sequence for all three DNA sequences is "MEDAKNIK."
[0299] 4) RNA secondary structure and free energy: Reduce the continuous base-pairing region to lower the minimum free energy. If both cannot be achieved simultaneously, prioritize reducing the continuous base-pairing region.
[0300] According to the above optimization method, the optimized sequences GLuc1, 2, 3, and 4 of the present invention were obtained.
[0301] [Corrected 25.10.2024 according to Rule 91] As shown in Figures 33-35, the usage percentages of the first, second, third, and fourth codons in the total sequence of GLuc1 protein range from 68% to 100%, 0% to 29%, 0% to 7%, and 0% to 7%. The minimum free energy (MFE) of the GLuc1 RNA sequence is -226.96 kcal / mol.
[0302] [Corrected 25.10.2024 per Rule 91] As shown in Figures 36-38, the usage percentages of the first, second, third, and fourth codons in the total sequence of GLuc2 range from 44% to 100%, 0% to 50%, 0% to 14%, and 0% to 3%, respectively. The minimum free energy of the GLuc2 RNA sequence is -202.70 kcal / mol.
[0303] [Corrected 25.10.2024 per Rule 91] As shown in Figures 39-41, the usage percentages of the first, second, third, and fourth codons in the total sequence of GLuc3 are controlled within the following ranges: 70-100%, 0-40%, 0-15%, and 0-5%. The minimum free energy of the GLuc3 RNA sequence is -241.70 kcal / mol.
[0304] [Corrected 25.10.2024 per Rule 91] As shown in Figures 42-44, the usage percentages of the first, second, third, and fourth codons in the total sequence of GLuc4 range from 44% to 100%, 0% to 40%, 0% to 12%, and 0% respectively. The minimum free energy of the GLuc4 RNA sequence is -210.10 kcal / mol.
[0305] Four codon-optimized DNA sequences were used as examples, and wild-type DNA sequences and DNA sequences from vectors obtained from other companies were used as comparisons to conduct cell experiments. They are named as shown in Table 8:
[0306] Table 8 Comparative analysis of Gaussia luciferase protein and DNA sequences
[0307] Example 4 Codon Optimization of Firefly Luciferase
[0308] (1) Amino acid sequence analysis
[0309] The present invention involves two versions of the amino acid sequences of Photinus firefly luciferase, as shown in Table 9.
[0310] Table 9
[0311] Amino acid sequence 1:***.
[0312] Amino acid sequence 2:
[0313] (2) Convert amino acid sequence to DNA sequence
[0314] Convert the entire amino acid sequence to the most commonly used codon sequence for the species to be used. For example, for Homo sapiens, replace all "alanine" with "GCC." This yields the first version of the DNA sequence, as shown in Figure 2.
[0315] Link URL:
[0316] https: / / www.kazusa.or.jp / codon / cgi-bin / showcodon.cgi?species=9606
[0317] https: / / www.kazusa.or.jp / codon /
[0318] First version of the DNA sequence:
[0319] (3) DNA sequence analysis
[0320] The operation of step (2) is performed using amino acid sequence 2, and the obtained DNA sequence is the first version of the DNA sequence.
[0321] [Corrected 25.10.2024 according to Rule 91] The first version of the DNA sequence was analyzed for GC content. As shown in Figure 45
[0322] Link: https: / / www.vectorbuilder.cn / tool / gc-content-calculator.html
[0323] [Corrected 25.10.2024 in accordance with Rule 91] The first version of the DNA sequence was subjected to sequence duplication analysis, as shown in Figure 46.
[0324] Link: https: / / www.vectorbuilder.cn / tool / sequence-dot-plot.html
[0325] [Corrected 25.10.2024 according to Rule 91] The first version of the DNA sequence was subjected to RNA secondary structure and free energy analysis. The minimum free energy (MFE) of the RNA sequence was -687.5 kcal / mol, and the RNA secondary structure is shown in Figure 47.
[0326] Link: http: / / rna.tbi.univie.ac.at / / cgi-bin / RNAWebSuite / RNAfold.cgi
[0327] (4) DNA sequence adjustment
[0328] By combining parameters such as codon preference, GC content ratio, sequence repetition, RNA secondary structure, and RNA free energy, an optimized version of the DNA sequence is obtained.
[0329] The optimization of each parameter is described below.
[0330] 1) Codon Bias: After obtaining the first version of the DNA sequence based on the most frequently used codons, the codons with the second, third, and fourth highest usage percentages are used to replace the codons with the highest usage percentages. As a general rule, the percentages of the codons with the highest, second, third, and fourth highest usage percentages in the protein sequence should be controlled within the following ranges: 40-100%, 0-50%, 0-25%, and 0-15%. As shown in Figure 5, the codons corresponding to amino acids are identified using the standard codon table. This is combined with the codon table for Homo sapiens in Figure 1 to determine the codon bias. In Homo sapiens, the percentage of glycine GGC is 40-100%, GGA is 0-50%, GGG is 0-25%, and GGT is 0-15%.
[0331] Link URL: https: / / zh.wikipedia.org / wiki / DNA%E5%AF%86%E7%A0%81%E5%AD%90%E8%A1%A8
[0332] GC content: The GC content should be higher than the GC content of the species (0% to 15%). For example, the GC content of Homo sapiens is 52.27%, so the optimized DNA sequence should have a GC content of 52.27% to 67.27%. GC content varies between species. For example, the GC content of Macaca fascicularis is 49.64%, so the optimized DNA sequence should have a GC content of 49.64% to 64.64%. The GC content of a DNA sequence should be controlled between 30% and 95%.
[0333] 3) Sequence repetitions: Optimize the number and length of repeats to minimize the number of repeats. For example, replace a codon in a sequence segment with a repeat length greater than 20 bases with a synonymous codon to reduce the number of bases in the repeat sequence. Example: Before optimization, the repeat sequence "ATGGAGGACGCCAAGAACATCAAG" has a total of 24 bases appearing three times. By mutating the codons at two different positions to synonymous codons, two different DNA sequences are obtained: "ATGGAAGACGCCAAGAACATCAAG" and "ATGGAGGATGCCAAGAACATCAAG." The amino acid sequence for all three DNA sequences is "MEDAKNIK."
[0334] 4) RNA secondary structure and free energy: Reduce the continuous base pairing region and lower the minimum free energy. If both cannot be achieved simultaneously, reduce the continuous base pairing region first.
[0335] According to the above optimization method, No. 8 and No. 9 of the present invention were obtained.
[0336] [Corrected 25.10.2024 according to Rule 91] As shown in Figures 48-50, the usage percentages of the first, second, third, and fourth codons in the total sequence of LUC-1 range from 68% to 100%, 0% to 29%, 0% to 7%, and 0% to 7%. The minimum free energy (MFE) of the RNA sequence No. 8 is -639.90 kcal / mol.
[0337] [Corrected 25.10.2024 according to Rule 91] As shown in Figures 51-53, the usage percentages of the first, second, third, and fourth codons in the total sequence of LUC-2 range from 44% to 100%, 0% to 50%, 0% to 14%, and 0% to 3%, respectively. The minimum free energy of RNA sequence No. 9 is -655.40 kcal / mol.
[0338] Table 10
[0339] Example 5 SpCas9 codon optimization
[0340] (1) Analysis of amino acid sequence
[0341] The present invention involves only one amino acid sequence (SpCas9):
[0342] (2) DNA sequence conversion
[0343] According to the codon preference table, all amino acid sequences are converted into the most commonly used codon sequences of the species to be applied. The frequency of human codon usage is shown in the following table:
[0344] Table 11
[0345] For example, for the species "Homo sapiens," all "alanine" residues were replaced with "GCC." This yielded the first version of the DNA sequence, hSpCas9-WT.
[0346]
[0347] (3) DNA sequence analysis
[0348] DNA sequence analysis of hSpCas9-WT was performed, including GC content analysis (https: / / www.vectorbuilder.cn / tool / gc-content-calculator.html), sequence repetition analysis (https: / / www.vectorbuilder.cn / tool / sequence-dot-plot.html), and sequence RNA secondary structure and free energy analysis (http: / / rna.tbi.univie.ac.at / / cgi-bin / RNAWebSuite / RNAfold.cgi).
[0349] (4) DNA sequence adjustment
[0350] By combining parameters such as codon preference, GC content ratio, sequence repetition, RNA secondary structure, and RNA free energy, an optimized version of the DNA sequence is obtained.
[0351] The optimization of each parameter is described as follows.
[0352] 1) Codon Bias: After deriving the first version of the DNA sequence based on the most frequently used codons, the codons with the second, third, and fourth highest usage percentages are used in that order. As a general rule, the percentages of the codons with the highest, second, third, and fourth highest usage percentages in the total protein sequence should be controlled within the following ranges: 40-100%, 0-50%, 0-25%, and 0-15%.
[0353] 2) GC content: higher than the 0% to 15% of the total GC content of the species. The total GC content of the species "Homo sapiens" is 52.27%, so the total GC content of the optimized DNA sequence should be "52.27% to 67.27%".
[0354] 3) Sequence repetition: Optimize the number of repeated sequences and sequence length to minimize the number of repeated sequences.
[0355] 4) RNA secondary structure and free energy: Reduce the continuous base-pairing region to lower the minimum free energy. If both cannot be achieved simultaneously, prioritize reducing the continuous base-pairing region.
[0356] According to the above optimization method, a total of 5 versions of nucleic acid sequences were generated, including hSpCas9 in the prior art as a comparative example, and they are named as follows:
[0357] [Corrected 25.10.2024 in accordance with Rule 91] Table 11
[0358] Effect verification
[0359] 1. Expression verification of the sequences of Examples 1 to 5 was performed. Taking OVA involved in Example 1 as an example, the construction of other protein expression vectors is similar to this step:
[0360] 1. Vector Construction
[0361] [Corrected 25.10.2024 according to Rule 91] The DNA sequence of chicken ovalbumin was inserted into the pmRVac backbone as an antigen coding region using the Gibson method (the DNA sequences of all elements except the target sequence remained consistent with those of the comparative example and the example. The pmRVac vector was modified from the pUC19k vector, as shown in Figure 60). A Gibson reaction buffer containing the pmRVac-OVA vector (Figure 60) was obtained. The reaction buffer was then transferred to VB UltraStable competent cells by electroporation, SOC culture medium was added, and the cells were incubated on a shaker for 1 hour. A portion of the culture medium was then removed and spread onto LB plates containing Kana antibiotics, streaked, and incubated overnight in a 37°C incubator.
[0362] 2. Clone identification
[0363] Pick several single colonies from the overnight cultured Kana plates in step 1 and dissolve them in sterile water to obtain a bacterial suspension. Perform PCR amplification on this suspension and identify the correct clones by gel electrophoresis. Add the bacterial suspension of the correct clones to LB medium containing Kana antibiotics and incubate overnight at 37°C in a shaker.
[0364] 3. Plasmid Extraction
[0365] Take the overnight culture solution in step 2 and use a plasmid extraction kit to extract the plasmid. Perform enzyme digestion, sequencing and other identification on the obtained plasmid to obtain the correct plasmid.
[0366] 4. Plasmid linearization
[0367] The plasmid from step 3 was digested with SapI and purified using a PCR product purification kit to obtain a purified linearized plasmid.
[0368] 5. In vitro transcription of mRNA
[0369] Take the linearized plasmid in step 4 and add it to the transcription system, react at 37°C to obtain a crude mRNA sample, purify it, heat denature it, add it to the capping system, cap it at 37°C, and obtain the mRNA sample after magnetic bead purification.
[0370] 6. Cell transfection
[0371] mRNA transfection dose: 1ug / well.
[0372] Take the mRNA in step 5 and add it to 293T cells in a 12-well plate (the cell density during transfection is 80%). After culturing at 37°C for 24 hours, collect the supernatant and place it on ice for labeling.
[0373] 2. Expression effect detection:
[0374] 2.1. Determination of protein concentration by BCA assay
[0375] Dilute the BCA protein standard according to a gradient of concentrations and add an appropriate amount to a 96-well microtiter plate. Mix BCA working solution Reagents A and B in a 50:1 ratio. Add the BCA working solution to each well of the microtiter plate immediately before use. Oscillate the microtiter plate on a shaker and incubate at 37°C. Measure absorbance at 562 nm using a microplate spectrophotometer. Plot a standard curve using protein content (µg) on the horizontal axis and absorbance on the vertical axis, and fit a linear equation.
[0376] Dilute the sample to be tested to the appropriate concentration, add BCA working solution, mix thoroughly, incubate at 37°C, and measure the absorbance. Calculate the protein concentration from the absorbance value using the standard curve and multiply by the sample dilution factor to obtain the sample concentration.
[0377] 2.2 Western blot experiment and SDS-PAGE gel staining and destaining
[0378] Mix the supernatant obtained in step 2.1 thoroughly, add 4x SDS loading, mix thoroughly, heat to denature, and place on ice. Assemble two precast gels and place them in an electrophoresis tank. Add buffer and apply samples in the same order and volume. Run electrophoresis at 150mV for 1 hour. Cut a 0.45μm PVDF membrane of appropriate size and transfer one precast gel immediately after electrophoresis. Transfer the membrane at 250mA for 45 minutes. After transfer, remove the membrane with tweezers and place it in a container. Rinse with TBST, add blocking buffer, and block at room temperature for 1 hour. Rinse twice with TBST and incubate with primary antibody overnight. Rinse three times with TBST and incubate with secondary antibody at room temperature for 1.5 hours. After rinsing three times with TBST, add ECL Western blotting substrate to the membrane and perform luminescence imaging using a gel scanner to observe protein expression results.
[0379] For another precast gel, remove the gel and soak it in a glass dish containing an appropriate amount of Coomassie Brilliant Blue stain. Place the dish on an orbital shaker and incubate at room temperature for 1 hour. Discard the staining solution and add an appropriate amount of destaining solution. Destain at room temperature for 4-24 hours, changing the destaining solution 2-4 times.
[0380] Take 100 μg / well sample for WB and SDS-PAGE experiments.
[0381] 2.3 In vivo experiments
[0382] According to the results of the WB experiment, the Cre recombinase obtained in Example 2 was used for the experiment. The mRNA sample obtained in step 6 was encapsulated with LNP and injected into the tail vein of Ai9 mice at a low dose (group marked as L, 0.1 mg / kg mouse body weight) and a high dose (group marked as H, 0.4 mg / kg mouse body weight). After 48 hours of injection, the sample was perfused by heart, and then the liver and spleen were dissected for fixation, dehydration, and frozen sectioning to observe the tdTomato red fluorescence intensity to compare the effect. The negative control group was the PBS injection group.
[0383] 3. Test Conclusion
[0384] [Corrected 25.10.2024 according to Rule 91] 3.1 In Example 1, the results of the cell expression assays in each group are shown in Figures 61 to 63. The results show that, at the same dose in Example 1, OVA(1) protein expression was the strongest, OVA(G) was weaker, and OVA(3) was the worst. No significant protein expression was observed in the other samples. The SDS-PAGE gel bands were generally consistent across the samples, indicating consistent WB loading.
[0385] [Corrected 25.10.2024 according to Rule 91] 3.2 In Example 2, the test results after expression of each group of cells are shown in Figure 64. As shown in Figure 64, in the WB experiment, the size of the Cre recombinase band is consistent; and at different doses (the mRNA transfection dosage is 1ug / well and 2ug / well respectively), similar trends are observed; and the expression levels of the internal reference protein Vinculin in different groups are similar, which proves that the experimental results are credible. As shown in Figure 64, in the WB experiment, at low doses, the expression level of Co3 is relatively the highest, followed by the expression level of Co1, both of which are significantly better than the wild type or other companies' control group. At high doses, the expression level of Co1 is relatively the highest, followed by the expression level of Co3, both of which are significantly better than the wild type or other companies' control group. As shown in Figure 65, in the in vivo experiment, the tdTomato red fluorescence intensity of Co3 is relatively the highest, slightly better than the control group.
[0386] [Corrected 25.10.2024 according to Rule 91] 3.3 In Example 3, at four time points, namely 6h, 24h, 48h, and 72h after transfection, 20ul of cell culture supernatant was aspirated and mixed with 0.5ul of substrate coelenterazine and 50ul of reaction buffer. The expression product Gaussia luciferase came into contact with its substrate coelenterazine to produce an enzymatic reaction and chemiluminescence. The intensity of the chemiluminescence was measured using a microplate chemiluminescence analyzer, and the intensity was converted into a numerical value to show the expression effect of the product. The expression efficiency of the mRNA optimized by different codons can be corresponding to the intensity of the chemiluminescence reaction of the expression product. The experimental results are shown in Table 12 and Figure 66.
[0387] Table 12 Comparison of expression effect intensity between the embodiment and the comparative example cell experiment
[0388] The results showed that GLuc2 and GLuc3 were significantly superior to GLuc-wt and slightly superior to GLuc-other. GLuc1 and GLuc4 were slightly superior to GLuc-wt. The mRNA expression vector constructed with the Gaussia luciferase DNA sequence of the present invention showed higher expression efficiency in HEK293T cells than the mRNA expression vector constructed with the wild-type Gaussia luciferase DNA sequence.
[0389] 3.4 In Example 4, cell experiments and animal experiments were conducted respectively:
[0390] [Corrected 25.10.2024 according to Rule 91] In the cell-based assay, two experimental groups were set up for dose comparison, with mRNA dosages of 0.25 μg / well and 0.5 μg / well, respectively. The expressed mRNA was added to 293T cells in a 12-well plate (106 cells / well) and incubated at 37°C. Samples were processed at 6, 24, 48, and 72 hours. The mRNA-transfected cells were removed from the incubator, and the cell waste was removed in a clean bench. Cell lysis buffer was added. After incubation at room temperature for 1 minute, the cells were gently shaken to dislodge. The cell lysate was pipetted into a clean 1.5 mL centrifuge tube and centrifuged at 12,000 rpm at 4°C for 3 minutes. An equal amount of the supernatant from the centrifuged cell lysate was added to a 96-well plate containing D-Luciferin solution. Mix thoroughly by gently pipetting. The tubes were placed in a microplate reader protected from light and the fluorescence intensity was read. The results showed that the expression level of Luciferase reached its peak at 24 h, and the corresponding chemiluminescence intensity of LUC-1 and LUC-2 samples was higher ( Figure 67 ).
[0391] In animal experiments,
[0392] The mRNAs corresponding to control sequences 1, 3, and 5, as well as LUC-1 and LUC-2, were encapsulated in LNPs and subsequently tested in mice. Six- to eight-week-old c57BL / 6J female mice were used. The total injection dose was 10 μg and 30 μg (total dose in both legs). Sample concentrations and injection volumes are shown in the table below. For injection volumes less than 50 μL, 1× PBS solution was used to make up the total injection volume to 50 μL.
[0393] Table 13
[0394] Before sample injection, mice were marked, captured, and secured. After disinfecting the injection area with 70% alcohol, the sample was injected intramuscularly into both hind legs. After injection, mice were returned to their original cages and housed normally. In vivo imaging was performed 6, 24, 48, and 72 hours after injection. Before in vivo imaging, a 15 mg / mL D-Luciferin solution (substrate) was prepared in D-PBS, sterilized by filtration through a 0.22 μm filter, and stored in a dark place. Before substrate injection, each mouse was weighed and recorded. D-Luciferin solution was injected intraperitoneally at 150 mg / kg body weight. After substrate injection, mice were placed in an induced anesthesia chamber (isoflurane gas anesthesia) to achieve full anesthesia. Approximately 15 minutes after luciferin substrate injection was allowed to proceed with in vivo imaging.
[0395] [Corrected 25.10.2024 according to Rule 91] From the results in Figures 68 to 71 and Tables 14 to 17, it can be seen that the in vivo expression effects of LUC-1 and LUC-2 samples are relatively strong.
[0396] Table 14 Chemiluminescence detection data of Firefly Luciferase transfection for 6 hours
[0397] Table 15 Chemiluminescence detection data of Firefly Luciferase transfection 24h
[0398] Table 16 Chemiluminescence detection data of Firefly Luciferase transfection 48h
[0399] Table 17 Chemiluminescence detection data of Firefly Luciferase transfection 72h
[0400] 3.5 In Example 5,
[0401] [Corrected 25.10.2024 according to Rule 91] 3.5.1 Western Blot (WB) results are shown in Figure 72. The results show that VB1-5 protein expression is stronger than WT. Luminescence imaging was followed by grayscale scanning analysis using Image J software. After grayscale scanning, all samples were normalized using GAPDH as a reference, and the fold change of each group relative to the control was calculated. The data processing results are shown in Figure 73. VB5 expression increased 4-fold compared to WT, VB1 and VB2 expression increased 5-fold compared to WT, and VB3 and VB4 expression increased 8-fold compared to WT.
[0402] 3.5.2 Functional verification of the encoding nucleic acid of Example 5 was also performed
[0403] 3.5.2.1 Cell transfection
[0404] Take the mRNA prepared above and adjust the mRNA input amount as shown in the table below based on the relative protein expression fold change results obtained in the Western blotting experiment. To ensure the same total amount of mRNA input, use nonsense RNA to supplement the input amount to a total of 1 μg. Use 293T cells in a 12-well plate (10E6 cells / well). Add 1 μg of mRNA and 1 μg of gRNA (such as gEMX1, as described above) to each well. Harvest cells 48 hours after transfection, remove as much supernatant as possible, and store at -80°C.
[0405] Table 18
[0406] 3.5.2.2 Cell genome extraction
[0407] Extract the genome of the cells using a genomic DNA extraction kit as described in the kit instructions. After extraction, take 2 μL of the sample for concentration measurement and set aside.
[0408] 3.5.2.3 PCR amplification of the desired fragment
[0409] The target fragment was amplified by PCR using the previously extracted cell genome based on the following components and corresponding volumes:
[0410] Table 19
[0411] 3.5.2.4 PCR amplification procedures include:
[0412] 94°C, 1 min; → [98°C, 10 s → 68°C, 60 s] × 30 cycles → 72°C, 10 min → 4°C, temporarily store.
[0413] 3.5.2.5 PCR product purification and recovery
[0414] Purify and recover the PCR product obtained by PCR amplification using a standard DNA product purification kit according to the instructions. After completion, take 2 μL of the aliquot to measure the concentration and set aside.
[0415] 3.5.2.6 T7E1 digestion
[0416] The PCR product obtained from the previous PCR product purification and recovery step was digested with T7E1 to assess its gene editing efficiency. The specific steps are as follows.
[0417] The annealing system for PCR purification products is as follows:
[0418] Table 20
[0419] The procedure is as follows:
[0420] 95℃, 5min→95-85℃, -2℃ / s→85-25℃, -0.1℃ / s→4℃, temporarily store.
[0421] Add 1 μL T7 Endonuclease I to the annealing reaction system, digest at 37°C for 15 min, and analyze by 2% agarose gel electrophoresis.
[0422] [Corrected 25.10.2024 according to Rule 91] The results are shown in Figure 74. In the functional verification, reducing the mRNA input ratio of the disclosed example can still achieve an editing efficiency similar to that of the comparative example.
[0423] Grayscale scanning of agarose gel electrophoresis images was performed using Image J, and the gene editing efficiency of each experimental group was calculated according to a publicly available calculation method (see: Guschin, DY, et al. (2010) A rapid and general assay for monitoring endogenous gene modification. Methods Mol Biol, 649, 247–256.).
[0424] [Corrected 25.10.2024 in accordance with Rule 91] The data processing results are shown in Figure 75. When the amount of VB3 mRNA input in the disclosed embodiment was 1 / 8 of that in the comparative example, similar editing efficiency was still achieved. This significantly reduced the amount of mRNA input. When the amounts of VB1, VB2, and VB5 were 1 / 4 of that in the comparative example, the editing efficiency was close to that in the comparative example.
[0425] In summary, the hSpCas9 protein expression provided in the embodiments of the present disclosure is improved compared to the control ratio, which is conducive to better application in in vivo experiments and can also produce market-competitive products.
[0426] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for codon optimization, comprising: Obtaining a coding nucleic acid sequence of a protein according to the codon preference of a host, and then adjusting the GC content, repeat sequence, secondary structure, and / or free energy in the obtained sequence to obtain an optimized coding nucleic acid; The GC content includes: the full-length GC content of the coding nucleic acid and the GC content within a unit length range.
2. The method according to claim 1, wherein The GC content within the unit length range is the GC content within 10 bp to 100 bp; preferably, it is the GC content within 20 bp to 50 bp.
3. The method according to claim 2, wherein The GC content within the unit length range is 20% to 95%; preferably, the GC content is 26% to 63% or 40% to 83%.
4. The method according to claim 1, wherein The full-length GC content is higher than 0% to 15% of the total GC content ratio of the host species.
5. The method according to claim 4, wherein The host is a human cell, and the full-length GC content is 52.27% to 67.27%, preferably, the full-length GC content is 56% to 63%; The host is a cynomolgus monkey cell, and the full-length GC content is 49.64% to 64.64%, preferably, the full-length GC content is 54% to 61%.
6. The method according to claim 1, characterized in that The length of the repeat sequence is not greater than 20 bp, preferably, the length of the repeat sequence is not greater than 30 bp.
7. The method according to claim 1, wherein The adjustment of the secondary structure includes reducing the continuous base pairing region.
8. The method according to claim 1, wherein The adjustment of the free energy includes making the minimum free energy of the sequence be 80% to 100% of the lowest free energy.
9. The method according to any one of claims 1 to 8, characterized in that, The adjustment includes: sequentially using the codons with the second, third, and fourth usage ratios of codon preference to replace the codons with the first usage ratio, and the percentages of the codons with the first, second, third, and fourth usage frequencies in the full length of the coding nucleic acid are: 40% to 100%, 0% to 50%, 0% to 25%, 0% to 15%, respectively.
10. The method according to claim 9, wherein The adjustment includes making: Among the codons of Phe, the usage frequency of UUU is 1% to 25%, and the usage frequency of UUC is 75% to 99%; Among the codons of Leu, the usage frequency of CUG is 86% to 99%, the usage frequency of CUC is 1% to 10%, the usage frequency of UUA is 0% to 1%, the usage frequency of UUG is 0% to 1%, the usage frequency of CUU is 0% to 1%, and the usage frequency of CUA is 0% to 1%; Among the codons of Ile, the usage frequency of AUU is 5% to 25%, the usage frequency of AUC is 75 to 94%, and the usage frequency of AUA is 0% to 1%; Among the codons of Val, the usage frequency of GUU is 0% to 5%, the usage frequency of GUC is 5% to 15%, the usage frequency of GUA is 0% to 1%, and the usage frequency of GUG is 79% to 95%; Among the codons of Ser, the usage frequency of UCU is 0% to 20%, the usage frequency of UCC is 10 to 25%, the usage frequency of UCA is 0% to 1%, the usage frequency of UCG is 0% to 1%; the usage frequency of AGU is 0% to 1%, and the usage frequency of AGC is 52% to 90%; Among the codons of Pro, the usage frequency of CCU is 20% - 50%, the usage frequency of CCC is 44% - 72%, the usage frequency of CCA is 0% - 8%, and the usage frequency of CCG is 0% - 6%; Among the codons of Thr, the usage frequency of ACU is 0% - 1%, the usage frequency of ACC is 50% - 78%, the usage frequency of ACA is 20% - 48%, and the usage frequency of ACG is 0% - 1%; Among the codons of Ala, the usage frequency of GCU is 3% - 20%, the usage frequency of GCC is 73% - 90%, the usage frequency of GCA is 0% - 6%, and the usage frequency of GCG is 0% - 1%; Among the codons of Tyr, the usage frequency of UAU is 1% - 25%, and the usage frequency of UAC is 75% - 99%; Among the codons of His, the usage frequency of CAU is 1% - 15%, and the usage frequency of CAC is 85% - 99%; Among the codons of Gln, the usage frequency of CAA is 1% - 15%, and the usage frequency of CAG is 85% - 99%; Among the codons of Asn, the usage frequency of AAU is 5% - 30%, and the usage frequency of AAC is 70% - 95%; Among the codons of Lys, the usage frequency of AAA is 1% - 25%, and the usage frequency of AAG is 75% - 99%; Among the codons of Asp, the usage frequency of GAU is 3% - 35%, and the usage frequency of GAC is 65% - 97%; Among the codons of Glu, the usage frequency of GAA is 10% - 35%, and the usage frequency of GAG is 65% - 90%; Among the codons of Cys, the usage frequency of UGU is 15% - 40%, and the usage frequency of UGC is 60% - 85%; The codon of Trp is GUU; Among the codons of Arg, the usage frequency of CGU is 0% - 1%, the usage frequency of CGC is 1% - 8%, the usage frequency of CGA is 0% - 1%, the usage frequency of CGG is 44% - 72%, the usage frequency of AGA is 15% - 50%, and the usage frequency of AGG is 1% - 8%; Among the codons of Gly, the usage frequency of GGU is 0% - 5%, the usage frequency of GGC is 65% - 98%, the usage frequency of GGA is 1% - 20%, and the usage frequency of GGG is 1% - 10%.
11. A codon-optimized system, characterized in that, It includes a module for implementing the method according to any one of claims 1 to 10.
12. A codon-optimized electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the method according to any one of claims 1 to 10.
13. A computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.
14. A biological material, which includes at least one of the following I) - V): I) Nucleic acid optimized by the method according to any one of claims 1 to 10; II) An expression unit containing the nucleic acid described in I); III) A plasmid vector containing the nucleic acid described in I), or containing the expression unit described in II); IV), mRNA, which contains a 5'-terminal cap structure and the nucleic acid described in I), or contains a 5'-terminal cap structure and the expression unit described in II); V), a host cell, which includes the nucleic acid described in I integrated into its genome, or the expression unit described in II integrated into its genome, or is transformed or transfected with the plasmid vector described in III).
15. The biomaterial according to claim 14, wherein In the said nucleic acid: The nucleic acid encoding chicken ovalbumin has a sequence shown in any one of SEQ ID NO: 1 to 4; The nucleic acid encoding Cre recombinase has a sequence shown in any one of SEQ ID NO: 10 to 13; The nucleic acid encoding Gauss luciferase has a sequence shown in any one of SEQ ID NO: 15 to 18; The nucleic acid encoding North American firefly luciferase has a sequence shown in any one of SEQ ID NO: 23 to 24; The nucleic acid encoding SpCas9 has a sequence shown in any one of SEQ ID NO: 33 to 37.
16. A method for preparing a recombinant protein, which includes, After obtaining the optimized nucleic acid molecule of the method described in any one of claims 1 to 10, introducing the nucleic acid molecule into a host cell, and obtaining a product containing the recombinant protein after cultivation.
17. A method for preparing an mRNA transfection agent or vaccine, comprising: After obtaining the optimized nucleic acid molecule of the method described in any one of claims 1 to 10, connecting the nucleic acid molecule with a vector, linearizing it, and then encapsulating it after capping and purification to obtain an mRNA transfection preparation or a vaccine.
Citation Information
Patent Citations
Codon optimization method for increasing expression level of target gene in host body
CN110423769A
Codon optimization
CN112513989A
Codon optimization method
US20070292918A1
Codon optimization of a synthetic gene(s) for protein expression
US20140244228A1
Compact promoters for gene expression
WO2022212766A2