Base sequence conversion device, base sequence conversion method, polynucleotide, method for producing polynucleotide, and program

By reducing high-frequency codons like ATG and GTA in the nucleotide sequence without changing the translated polypeptide, the nucleotide sequence conversion method enhances polypeptide synthesis efficiency.

JP7714207B2Active Publication Date: 2025-07-29NUPROTEIN CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021066901
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-12
Publication Date
2025-07-29
Estimated Expiration
2041-04-12

Smart Images

  • Figure 0007714207000007
    Figure 0007714207000007
  • Figure 0007714207000008
    Figure 0007714207000008
  • Figure 0007714207000009
    Figure 0007714207000009
Patent Text Reader

Abstract

To provide a base sequence conversion device, a base sequence conversion method, a polynucleotide fabricated from a base sequence obtained by the base sequence conversion method, a method for producing a polynucleotide, and a program for base sequence conversion.SOLUTION: A base sequence conversion device includes: an input part which inputs a base sequence encoding polypeptide; and a conversion part which converts an input base sequence, where the conversion part converts codon including at least any one of A, T and G in ATG and / or GTA in 5' to 3' directions in a base sequence without changing a polypeptide translated from the input base sequence, and reduces ATG and / or GTA in a base sequence.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosure in the present application relates to a base sequence conversion device, a base sequence conversion method, a polynucleotide, a method for producing a polynucleotide, and a program.

Background Art

[0002] Many studies have been conducted to clarify the structure and function of proteins. In protein research, it is necessary to obtain proteins, and in this process, techniques such as protein synthesis in living cells using recombinant DNA and cell-free protein synthesis using cell-derived factors have been developed. Although these techniques can obtain proteins more efficiently than extracting natural proteins from biological tissues and cells, there is a further demand to obtain even more proteins. Patent Document 1 and Patent Document 2 disclose optimizing the base sequence to increase the amount of synthesized protein.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] Patent Document 1 discloses optimizing a nucleotide sequence for protein expression using quality functions based on codon usage frequency, GC content, DNA motifs, repetitive sequences, secondary structure, and inverted repeats. Patent Document 2 also discloses improving and optimizing gene sequences for enhancing protein expression of genes in bacteria, yeast, insect, and mammalian cells using a particle swarm optimization algorithm, considering parameters and factors affecting protein expression, including codon usage frequency, tRNA usage, GC content, ribosome binding sequences, promoters, 5'-UTR, ORF, and 3'-UTR sequences.

[0005] The optimization of the nucleotide sequences disclosed in Patent Document 1 and Patent Document 2 is performed based on a lot of information such as codon usage frequency and GC content. Therefore, the applicant has conducted intensive research and newly found that by converting codons of the nucleotide sequence encoding a polypeptide without using a lot of information as disclosed in Patent Document 1 and Patent Document 2, the synthesis amount of the polypeptide can be increased.

[0006] That is, the object of the disclosure of the present application is to provide a nucleotide sequence conversion device, a nucleotide sequence conversion method, a polynucleotide, a method for producing a polynucleotide, and a program. Other optional additional effects of the disclosure of the present application will be clarified in the mode for carrying out the invention.

Means for Solving the Problems

[0007] (1) An input unit that inputs a nucleotide sequence encoding a polypeptide, A conversion unit that converts the input nucleotide sequence, comprising The conversion unit without changing the polypeptide translated from the input nucleotide sequence, converts a codon containing at least one of A, T, and G in ATG and / or GTA in the 5'-to-3' direction in the nucleotide sequence to reduce ATG and / or GTA in the nucleotide sequence, A nucleotide sequence conversion device. (2) The codon to be converted is the codon with the highest codon usage frequency among the codons that can be selected for conversion. The nucleotide sequence conversion device according to (1) above. (3) A nucleotide sequence conversion method including a conversion step of converting a nucleotide sequence encoding a polypeptide, The conversion step is Without changing the polypeptide translated from the nucleotide sequence before conversion, Convert a codon containing at least one of A, T, and G in ATG and / or GTA in the 5' to 3' direction in the nucleotide sequence, and reduce ATG and / or GTA in the nucleotide sequence. Nucleotide sequence conversion method. (4) The codon to be converted is the codon with the highest codon usage frequency among the codons that can be selected for conversion. The nucleotide sequence conversion method according to (3) above. (5) A polynucleotide prepared from a nucleotide sequence obtained by any one of the nucleotide sequence conversion devices according to (1) and (2) above and the nucleotide sequence conversion methods according to (3) and (4) above. (6) A method for preparing a polynucleotide from a nucleotide sequence obtained by any one of the nucleotide sequence conversion devices according to (1) and (2) above and the nucleotide sequence conversion methods according to (3) and (4) above. (7) A process of inputting a nucleotide sequence encoding a polypeptide, A process of converting the input nucleotide sequence, A program for causing a computer to execute, The process of converting the input nucleotide sequence is Without changing the polypeptide translated from the input nucleotide sequence, Convert a codon containing at least one of A, T, and G in ATG and / or GTA in the 5' to 3' direction in the nucleotide sequence, and reduce ATG and / or GTA in the nucleotide sequence. Program. (8) The codon to be converted is the codon with the highest codon usage frequency among the codons that can be selected for conversion. The program according to the above (7).

Advantages of the Invention

[0008] The base sequence with the codon converted increases the amount of polypeptide synthesized by translation compared to the base sequence without codon conversion.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0010] Hereinafter, with reference to the drawings, a base sequence conversion device, a base sequence conversion method, a polynucleotide, a method for producing a polynucleotide, and a program will be described. In this specification, parts having the same kind of functions are denoted by the same or similar reference numerals. And, the repeated explanations for the parts denoted by the same or similar reference numerals may be omitted.

[0011] (Embodiment of the Base Sequence Conversion Device) With reference to FIGS. 1 to 4, the base sequence conversion device 1 according to the embodiment will be described. FIG. 1 is a schematic diagram of the base sequence conversion device 1. FIG. 2 shows five patterns of codons to be converted. FIG. 3 shows a codon table with codon usage frequencies in wheat. FIG. 4 shows conversion examples for each pattern of codons to be converted.

[0012] The base sequence conversion device 1 according to the embodiment includes at least an input unit 2 and a conversion unit 3. In the example shown in FIG. 1, a storage unit 4 and a display unit 5 are additionally provided optionally.

[0013] The base sequence conversion device 1 according to the embodiment may be configured by a computer. The computer includes a control unit (CPU). When the control unit reads a predetermined program, the base sequence conversion device 1 is provided with the conversion unit 3.

[0014] The input unit 2 is not particularly limited as long as it can input a base sequence encoding a polypeptide into the base sequence conversion device 1. Examples of the input unit 2 include a keyboard, a mouse, or a touch panel. Alternatively, a base sequence encoding a polypeptide may be input to the input unit 2 via a network (for example, LAN, Internet, etc.), and in this case, the input unit 2 may be configured in the form of a network interface. Further alternatively, a gene sequence may be input to the input unit 2 using a scanner or storage means. The base sequence input to the input unit 2 may be a DNA sequence or an RNA sequence. In this specification, thymine (T) in the DNA sequence and uracil (U) in the RNA sequence are equivalent as base sequence information. Therefore, when the base sequence is described below, thymine may be uracil, and uracil may be thymine.

[0015] The nucleotide sequence encoding the polypeptide input to the input unit 2 is not particularly limited as long as it contains a sequence encoding the polypeptide, and may also contain a sequence that does not encode the polypeptide. The polypeptide translated from the nucleotide sequence with codons converted by the conversion unit 3 described later is the same as that before the codons are converted. Therefore, it is known what polypeptide the nucleotide sequence encoding the polypeptide input to the input unit 2 encodes. Thus, for example, when there is an arbitrary sequence on the 5'-side of the input nucleotide sequence, since the sequence encoding the polypeptide is known, triplets may be used as codons starting from the location where the polypeptide in the nucleotide sequence is encoded. In the case where the input nucleotide sequence is only a sequence encoding a polypeptide, triplets starting from the beginning on the 5'-side of the nucleotide sequence may be used as codons.

[0016] The conversion unit 3 converts the codons of the nucleotide sequence encoding the polypeptide without changing the polypeptide translated from the nucleotide sequence encoding the input polypeptide.

[0017] The general process of translation initiation for synthesizing a protein, which is a polypeptide, is carried out as follows. First, a transcription initiation factor recognizes the cap structure at the 5'-end of mRNA, and a tRNA bound to a small ribosome and methionine (Met) corresponding to the start codon binds to the mRNA. Next, the small ribosome and tRNA bound to the mRNA move in the 3'-direction to the codon corresponding to methionine (start codon), and the anticodon of the tRNA binds to the corresponding start codon on the mRNA. Then, a large ribosome associates therewith to form a complex of the ribosome, mRNA, and tRNA bound to methionine. And a tRNA having an anticodon corresponding to the next codon of methionine binds to the corresponding codon on the mRNA and the elongation reaction proceeds. Therefore, in the process of translation initiation, mRNA having a cap structure can recognize the start codon of the sequence encoding the polypeptide of the mRNA.

[0018] However, when synthesizing a polypeptide in a cell-free system, mRNA without a cap structure is used. Since proteins are synthesized even in a cell-free system, the binding of tRNA bound to a small ribosome and methionine to mRNA is not due to the cap structure. Therefore, if there are multiple sequences of ATG, which is the codon for methionine, in the mRNA encoding the protein, translation may start from a location that is not the original start position, depending on the binding location of tRNA bound to a small ribosome and methionine. Also, although translation proceeds in the 5' to 3' direction, since the mRNA to which the ribosome binds is single-stranded, there is also a possibility that tRNA bound to a small ribosome and methionine binds to a sequence that becomes ATG (GTA in the 5' to 3' direction) in the 3' to 5' direction. Therefore, it is considered that by reducing the ATG sequence in either the 5' to 3' direction or the 3' to 5' direction, it is possible to suppress the binding of tRNA bound to a small ribosome and methionine to an incorrect location, even for mRNA without a cap structure.

[0019] Therefore, in order to reduce ATG and / or GTA in the 5' to 3' direction in the nucleotide sequence encoding the input polypeptide, the conversion unit 3 converts a codon containing at least one of A, T, and G among ATG and / or GTA in the 5' to 3' direction without changing the polypeptide translated from the nucleotide sequence. Hereinafter, "ATG and / or GTA in the 5' to 3' direction" may also be referred to as "ATG and / or GTA". Also, the nucleotide sequence with the codon converted in the conversion unit 3 is useful for mRNA without a cap structure as described above, but is also useful for mRNA with a cap structure. Even for mRNA with a cap structure, unnecessary ATG and / or GTA in the sequence encoding the polypeptide can be reduced, and translation errors can be suppressed for mRNA with a cap structure.

[0020] Among the ATG and / or GTA to be converted by the conversion unit 3, there are five patterns of codons containing at least one of A, T, and G. Five patterns to be converted are shown in FIG. 2. The frames in FIG. 2 represent one codon, and the circles within the frames are any of A, T, G, or C. The five patterns are: (1) a pattern in which in two consecutive codons, ATG spans the two codons, the 5'-side codon has A and T, and the 3'-side codon has G; (2) a pattern in which in two consecutive codons, ATG spans the two codons, the 5'-side codon has A, and the 3'-side codon has T and G; (3) a pattern in which in two consecutive codons, GTA spans the two codons, the 5'-side codon has G and T, and the 3'-side codon has A; (4) a pattern in which in two consecutive codons, GTA spans the two codons, the 5'-side codon has G, and the 3'-side codon has T and A; (5) a pattern in which the codon is GTA.

[0021] In addition, in the nucleotide sequence encoding the polypeptide, those in which ATG exists as a codon are not converted. Although ATG encodes methionine, as shown in the codon table in FIG. 3, methionine is not translated by codons other than ATG. Therefore, if the codon of ATG is converted, it will be converted to something other than methionine, changing the polypeptide translated from the input nucleotide sequence. Thus, when ATG is a codon, it is not converted. Therefore, "without changing the polypeptide translated from the input nucleotide sequence" in this specification means not converting when there is ATG as a codon.

[0022] The codons to be converted are not particularly limited as long as they do not change the polypeptide translated from the nucleotide sequence encoding the input polypeptide. As shown in the codon table in Figure 3, there are multiple codons for some of the translated amino acids. When converting such codons, for example, they may be converted to the codon with the highest codon usage frequency among the selectable codons, or to codons with other frequencies. Also, when ATG and / or GTA span two codons as in the above patterns (1) to (4), either the 5'-side or 3'-side codon may be converted, or both may be converted. For example, by comparing the codon usage frequencies, the codon with the higher codon usage frequency may be converted. Note that the codon usage frequency varies depending on the host and can be appropriately obtained according to the cells and cytokines used when synthesizing the polypeptide. Codon usage frequency tables for various hosts are available from http: / / www.kazusa.or.jp / codon / .

[0023] In a specific example, the conversion of codons will be described below. Figure 4 shows examples of codons to be converted in each of the above patterns (1) to (5).

[0024] Pattern (1) shows an example where two codons, GAT-GGT and CAT-GCC, are arranged continuously. The example of GAT-GGT shows the codons for aspartic acid (Asp) and glycine (Gly) arranged together. In the sequence of GAT-GGT, the codons are converted so that ATG does not appear in the nucleotide sequence without changing the translated amino acids. From the codon table shown in Figure 3, the codons for aspartic acid are GAT and GAC. Also, the codons for glycine are GGT, GGC, GGA, and GGG. Therefore, the codon for aspartic acid is converted from GAT to GAC. On the other hand, the codon for glycine may be converted to one other than GGT, but since the first base of all the codons for glycine is G, it may not be necessary to convert from GGT. By converting GAT to GAC, the translated amino acids remain as aspartic acid and glycine, and the nucleotide sequence does not contain ATG.

[0025] The example of CAT-GCC has the codons of histidine (His) and alanine (Ala) arranged together. The codons of histidine are CAT and CAC. Also, the codons of alanine are GCT, GCC, GCA, and GCG. Therefore, similar to GAT-GGT above, in CAT-GCC, by converting the CAT of histidine to CAC, without changing the translated amino acid, the base sequence will not contain ATG.

[0026] Pattern (2) shows an example where two codons, CCA-TGG and CAA-TGC, are arranged consecutively. The example of CCA-TGG has the codons of proline (Pro) and tryptophan (Trp) arranged together. The codons of proline are CCT, CCC, CCA, and CCG. Also, the codon of tryptophan is TGG. Then, since there is only one codon for tryptophan, it cannot be converted. However, since there are four codons for proline, proline can be converted to any of the three, CCT, CCC, or CCG, so that ATG does not appear in the base sequence. The numerical values attached to the codon table shown in Figure 3 are the codon usage frequencies in wheat. For example, if the codon of proline to be converted is to be the one with a high codon usage frequency, it is converted to CCG, which has the highest codon usage frequency among CCT, CCC, and CCG. Note that the numerical values of the codon usage frequencies in Figure 3 are normalized so that the total frequency of each amino acid is 1.

[0027] The example of CAA-TGC has the codons of glutamine (Gln) and cysteine (Cys) arranged together. As shown in Figure 4, convert the codon CAA of glutamine to CAG.

[0028] Pattern (3) shows an example where two codons, AGT-AAA and GGT-ATT, are arranged consecutively. AGT-AAA has codons for serine (Ser) and lysine (Lys) arranged. The codons for serine are AGT, AGC, TCT, TCC, TCA, and TCG. In this case, for codon conversion, as long as the translated amino acid remains the same, not only AGC which converts the T in AGT, but also any of TCT, TCC, TCA, and TCG can be used for conversion. Therefore, as shown in Figure 4, it can be converted to any of the arrangements of the five codons.

[0029] GGT-ATT is an example of glycine (Gly) and isoleucine (Ile), and as shown in Figure 4, it can be converted to any of the arrangements of the three codons.

[0030] Pattern (4) shows an example where two codons, CTG-TAT and GCG-TAC, are arranged consecutively. CTG-TAT is leucine (Leu) and tyrosine (Tyr), and as shown in Figure 4, it can be converted to any of the arrangements of the four codons. Also, GCG-TAC is alanine (Ala) and tyrosine (Tyr), and as shown in Figure 4, it can be converted to any of the arrangements of the three codons.

[0031] Pattern (5) is the case where the codon is GTA. GTA is the codon for valine (Val). Since there are four codons for valine, as shown in Figure 4, it can be converted to any of the three codons other than GTA.

[0032] The nucleotide sequence converted by the conversion unit 3 increases the amount of polypeptide synthesis compared to the nucleotide sequence before conversion.

[0033] In the base sequence conversion device 1 according to the embodiment, the storage unit 4 and the display unit 5 are optional components. The storage unit 4 stores a program that performs a process of inputting a base sequence encoding a polypeptide and a process of converting the input base sequence. Further, the storage unit 4 may store data such as a base sequence encoding the input polypeptide and a base sequence converted by the conversion unit 3. Examples of the storage unit 4 include flash memories such as RAM, ROM, and SSD, and HDDs.

[0034] The display unit 5 is not particularly limited as long as it can display the base sequence encoding the polypeptide input by the input unit 2 and the base sequence converted by the conversion unit 3. Examples of the display unit 5 include a liquid crystal display, a CRT display, an organic EL display, and an LED display.

[0035] (Embodiment of the base sequence conversion method) An embodiment of the base sequence conversion method will be described. The base sequence conversion method according to the embodiment includes a conversion step of converting a base sequence encoding a polypeptide.

[0036] In the conversion step of converting the base sequence encoding the polypeptide, without changing the polypeptide translated from the base sequence before conversion, a codon containing at least one of A, T, and G in ATG and / or GTA in the base sequence is converted, and ATG and / or GTA in the base sequence is reduced. The conversion of a codon containing at least one of A, T, and G in ATG and / or GTA in the base sequence is the same as the codon conversion performed by the conversion unit 3 of the base sequence conversion device 1 according to the above-described embodiment.

[0037] The base sequence conversion device 1 and the base sequence conversion method according to the embodiment have the following effects. (1) Without changing the translated polypeptide, a base sequence in which codons are converted so as to reduce ATG or / and GTA in the base sequence can increase the amount of polypeptide synthesized by translation compared to the base sequence before codon conversion.

[0038] (Embodiments of polynucleotides and methods for producing polynucleotides) Using the base sequence obtained by converting the codon by the base sequence conversion device 1 and / or the base sequence conversion method according to the above embodiment, a polynucleotide with the codon converted can be produced. The polynucleotide may be, for example, DNA or RNA. The polynucleotide can be produced by a known method such as the phosphoramidite method.

[0039] A polypeptide can be synthesized using the produced polynucleotide. The polypeptide can be synthesized by a known method such as polypeptide synthesis in living cells by recombinant DNA or cell-free polypeptide synthesis using cell-derived factors.

[0040] (Embodiments of the program) The base sequence conversion device 1 according to the above embodiment can be configured by a computer. In that case, an existing computer can be used as it is. That is, by providing a program that causes the computer to execute a process of inputting a base sequence encoding a polypeptide and a process of converting the input base sequence, the computer can be used as the base sequence conversion device 1.

[0041] Examples are given below to specifically explain the embodiments disclosed in this application, but these examples are for the purpose of explaining the embodiments only. They do not represent a limitation or restriction of the scope of the invention disclosed in this application.

Example

[0042] <Example 1> [Codon conversion of green fluorescent protein (GFP)] Using the base sequence encoding GFP, codon conversion was performed to reduce ATG and GTA in the base sequence without changing the GFP synthesized by translation.

[0043] The results are shown in Figure 5. Figure 5A shows the original base sequence of GFP before conversion. In Figure 5A, ATG or GTA that corresponds to any of the five patterns described above is shown in a larger font. In addition, ATG as a codon translated into methionine is boxed. Figure 5B shows the base sequence of GFP after codon conversion (hereinafter, sometimes referred to as "converted GFP"). In Figure 5B, the underlined codons are the converted codons.

[0044] As shown in Figure 5A, the original GFP sequence contained 23 ATGs, including 15 ATGs of pattern (1), 2 ATGs of pattern (2), and 6 ATGs that are translated into methionine. It also contained 5 GTAs, including 2 GTAs of pattern (3) and 3 GTAs of pattern (5).

[0045] In the case of patterns (1) to (3), the 5' codon of the GFP-encoding sequence was converted. When multiple codons were available for conversion, the most frequently used codon was selected. As shown in Figure 5B, the converted GFP sequence contained no ATG or GTA codons other than the six ATG codons translated into methionine.

[0046] <Example 2> [Creating DNA with the converted GFP sequence and synthesizing GFP] The base sequence of the converted GFP was used to prepare the DNA (Optimized GFP) shown in Table 1. Optimized GFP was prepared by a custom synthesis service (Eurofins Genomics, Inc.).

[0047] [Table 1]

[0048] Using the prepared Optimized GFP, GFP was synthesized in a cell-free system according to the following procedure.

[0049] (1) Preparation of Transcription Template DNA by PCR For the transcription template DNA, a set of primers shown in Table 1 (forward: FW-E02-OptGFP, reverse: RV-OptGFP) was designed, and the transcription template DNA was prepared by PCR. The reaction solution composition for PCR is shown in Table 2. Also, the reaction cycles are shown in Table 3. The primers were prepared by a contract synthesis service (Eurofins Genomics K.K.).

[0050]

Table 2

[0051]

Table 3

[0052] The reagents and equipment used are as follows. ·PCR enzyme: Toyobo Co., Ltd. KOD-Plus-Neo ·Thermal Cycler: eppendorf Mastecycler X50s

[0053] (2) Transcription Reaction Next, using the prepared transcription template DNA, the translation template mRNA was prepared. The transcription reaction was carried out at 37°C for 3 hours using 2.5 μl of the previously prepared PCR reaction solution (containing transcription template DNA) and the reaction solution shown in Table 4 of PSS4050 manufactured by NUProtein.

[0054]

Table 4

[0055] To 25 μl of the transcription reaction solution, 10 μl of 4 M ammonium acetate was added and mixed well. Then, 100 μl of 100% ethanol was added and mixed by inverting. After centrifuging for several seconds using a tabletop centrifuge, it was left standing at -20°C for 10 minutes. Then, centrifugation was performed (12,000 rpm, 15 minutes, 4°C). After removing the supernatant, centrifugation was performed for several seconds using a tabletop centrifuge. The supernatant was removed again, and it was left standing until the precipitate dried. Then, 40 μl of RNase free water (DEPC water) was added to 25 μl of the transcription reaction solution, and the precipitate was suspended well with a tip. According to the PSS4050 protocol, the nucleic acid concentration was measured so that the amount of mRNA in 110 μl of the translation solution would be 35 μg, and it was filled up to 80 μl, which was used as the translation template mRNA solution.

[0056] (3) Translation reaction Next, using a translation reaction solution with the following composition, it was placed in an incubator at 16°C and reacted for 10 hours. Among the compositions shown in Table 5, a composition solution excluding the translation template mRNA was prepared. Then, after returning this composition solution to room temperature, the translation template mRNA was added, and it was pumped without foaming and reacted. Wheat germ extract and amino acid mix were used from NUProtein's PSS4050.

[0057] [Table 5]

[0058] After the reaction, the reaction solution was collected in an Eppendorf tube, and centrifugation was performed (15,000 rpm, 15 minutes, 4°C), and the supernatant was used as the GFP solution after translation was completed.

[0059] <Comparative Example 1> GFP was synthesized based on the base sequence of the original GFP without codon conversion. The synthesis procedure was the same as in Example 2 except that Native GFP, FW-E02-GFP, and RV-GFP shown in Table 6 were used.

[0060] [Table 6]

[0061] <Example 3> [Fluorescence measurement of synthesized GFP] Fluorescence measurement of the GPF synthesized in Example 2 and Comparative Example 1 was performed. 220 μl of a solution containing the GFP synthesized in Example 2 or Comparative Example 1 was used as a sample, irradiated with excitation light having a wavelength of 475 nm, and the fluorescence from the GFP was measured with a plate reader. As the plate reader, a GloMax (registered trademark) plate reader (Promega) was used.

[0062] The results are shown in Fig. 6. The GFP synthesized in Example 2 and Comparative Example 1 emitted fluorescence. And the amount of fluorescence in Example 2 was larger than that in Comparative Example 1. The GFP of Example 2 and the GFP of Comparative Example 1 have the same polypeptide sequence. Therefore, since the amount of fluorescence in Example 2 is larger than that in Comparative Example 1, it was shown that the amount of GFP synthesized in Example 2 is larger than the amount of GFP synthesized in Comparative Example 1.

[0063] From the above examples, it was shown that by converting the codons so as to reduce ATG and / or GTA in the base sequence without changing the polypeptide to be translated, the synthesis amount of the polypeptide by translation can be increased compared to the base sequence before codon conversion.

Industrial Applicability

[0064] Using the base sequence conversion device, base sequence conversion method, polynucleotide, method for producing a polynucleotide, and program disclosed in the present application, the synthesis amount of a polypeptide can be increased. Therefore, it is useful for those who handle polypeptides such as proteins.

Explanation of Signs

[0065] 1... Base sequence conversion device, 2... Input unit, 3... Conversion unit, 4... Storage unit, 5... Display unit

Claims

1. An input unit that inputs a nucleotide sequence encoding a polypeptide, A conversion unit that converts the input nucleotide sequence, comprising, The conversion unit is, Without changing the polypeptide translated from the input nucleotide sequence, Convert a codon containing at least one of A, T, and G in GTA in the 5' to 3' direction in the nucleotide sequence, and reduce GTA in the nucleotide sequence, Nucleotide sequence conversion device.

2. The conversion unit, in addition to the GTA, converts a codon containing at least one of A, T, and G in ATG in the 5' to 3' direction in the nucleotide sequence, and reduces ATG in the nucleotide sequence, The nucleotide sequence conversion device according to claim 1.

3. The converted codon is the codon with the highest codon usage frequency among the codons that can be selected during conversion, The nucleotide sequence conversion device according to claim 1 or 2.

4. A nucleotide sequence conversion method including a conversion step of converting a nucleotide sequence encoding a polypeptide, The conversion step is, Without changing the polypeptide translated from the nucleotide sequence before conversion, Convert a codon containing at least one of A, T, and G in GTA in the 5' to 3' direction in the nucleotide sequence, and reduce GTA in the nucleotide sequence, Nucleotide sequence conversion method.

5. The conversion step, in addition to the GTA, converts a codon containing at least one of A, T, and G in ATG in the 5' to 3' direction in the nucleotide sequence, and reduces ATG in the nucleotide sequence, The nucleotide sequence conversion method according to claim 4.

6. The converted codon is the codon with the highest codon usage frequency among the codons that can be selected during conversion, The nucleotide sequence conversion method according to claim 4 or 5.

7. A method for producing a polynucleotide, comprising a step of producing a polynucleotide from a nucleotide sequence obtained by any one of the nucleotide sequence conversion device according to claims 1 to 3 and the nucleotide sequence conversion method according to claims 4 to 6.

8. A process of inputting a nucleotide sequence encoding a polypeptide, A process of converting the input nucleotide sequence, A program for causing a computer to execute, The process of converting the input nucleotide sequence is, Without changing the polypeptide translated from the input nucleotide sequence, Converting a codon containing at least one of A, T, and G in GTA in the 5'-to-3' direction in a nucleotide sequence to reduce GTA in the nucleotide sequence. Program.

9. The process of converting the input nucleotide sequence, in addition to GTA, converts a codon containing at least one of A, T, and G in ATG in the 5'-to-3' direction in the nucleotide sequence to reduce ATG in the nucleotide sequence. The program according to claim 8.

10. The converted codon is the codon with the highest codon usage frequency among the codons that can be selected during conversion. The program according to claim 8 or 9.

Citation Information

Patent Citations

  • Methods and apparatus for optimizing nucleotide sequences for protein expression

    JP2006512649A

  • Method and apparatus for optimizing nucleotide sequences for protein expression.

    JP4510640B2

  • Method of Sequence Optimization for Improved Recombinant Protein Expression using a Particle Swarm Optimization Algorithm

    US20110081708A1

  • Method of designing base sequence

    WO2007102578A1

  • Cyclic RNA and protein production method

    WO2013118878A1