Nucleic acid sequence optimization method using a peptide barcode library with high measurement accuracy

JP2026125181APending Publication Date: 2026-08-03HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
HITACHI LTD
Filing Date
2025-01-22
Publication Date
2026-08-03

AI Technical Summary

Benefits of technology

【0012】 本発明により、核酸の配列を最適化するための方法及びキットが提供される。本発明の方法及びキットを使用することにより、質量分析計におけるペプチドバーコードの定量精度が向上する。これにより、核酸配列の発現量の定量的な比較が可能になり、核酸、例えばmRNA医薬の塩基配列を、高精度に最適化することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026125181000001_ABST
    Figure 2026125181000001_ABST
Patent Text Reader

Abstract

To provide a method and means for improving the measurement accuracy of peptide barcodes in mass spectrometers and for optimizing nucleic acid sequences with high precision. [Solution] A method for optimizing nucleic acid sequences, comprising the steps of: preparing a nucleic acid sequence comprising a candidate nucleic acid sequence including a target protein nucleic acid which is a sequence of an untranslated region and a sequence encoding a target protein; and a barcode nucleic acid which is a nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein; expressing a protein from the nucleic acid sequence; separating the peptide barcode from the protein; analyzing the separated peptide barcode; and selecting the optimal candidate nucleic acid sequence based on the results of the analysis, wherein the peptide barcode includes multiple types of peptide barcodes included in a peptide barcode library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods and means for optimizing the sequence of nucleic acids, such as mRNA, using a peptide barcode library.

Background Art

[0002] Nucleic acid drugs such as mRNA drugs are attracting expectations in terms of treating diseases different from conventional small molecule drugs, having a simple manufacturing method, and a rapid drug discovery process.

[0003] The performance of nucleic acid drugs is determined by their base sequences. The base sequence of an mRNA drug generally consists of a 5'untranslated region (5'UTR), an open reading frame (ORF) encoding a protein, and a 3'untranslated region (3'UTR). The 5'UTR is a site recognized by ribosomes and is involved in the control of expression. The expression efficiency of the ORF changes depending on the pattern of its codons. The 3'UTR is involved in the stability of mRNA.

[0004] For the performance design of nucleic acid drugs, that is, sequence optimization, methods based on deep learning that have developed significantly since 2006 are expected. However, there is no high-throughput method for measuring the expression level of each nucleic acid sequence, and the lack of learning data has become a bottleneck. As a method for solving this, it is conceivable to apply the high-throughput measurement method of the reaction efficiency to the base material of drug candidates (Non-Patent Document 1, Non-Patent Document 2, and Patent Document 1), which appeared in the 2010s for small molecule drugs and antibody drugs, to nucleic acid drugs.

[0005] In the measurement methods for small molecule drugs and antibody drugs, a recognition tag is added to the drug candidate to cause a reaction with a base material such as an antigen-antibody reaction, and the recognition tag of the reacted drug candidate is cleaved and extracted and quantified using an apparatus according to the type of recognition tag.

Prior Art Documents

Patent Documents

[0006] [Patent Document 1] Patent No. 7185929 [Non-patent literature]

[0007] [Non-Patent Document 1] Egloff, P., et al. Engineered peptide barcodes for in-depth analyzes of binding protein libraries, Nat Methods, 16(5):421-428 (2019) [Non-Patent Document 2] Gironda-Martinez A., et al. DNA-encoded chemical libraries; a comprehensive review with successful stories and future challenges, ACS Pharmacol. Transi. Sci. 4(4):1265-1279 (2021) [Non-Patent Document 3] Dincer, AB, et al. Reducing peptide bias in quantitative mass spectrometry data with machine learning, J Proteome Res 1;21(7):1771 (2022) [Overview of the Initiative] [Problems that the invention aims to solve]

[0008] In the case of nucleic acid drugs, the recognition tag is implemented using a peptide (peptide barcode), and quantification is performed by mass spectrometry, but mass spectrometry is a measurement method with large variability in sensitivity (Non-Patent Literature 3). This is due to the large variability in the ionization efficiency of the substance being measured. [Means for solving the problem]

[0009] This invention provides a method and kit for optimizing nucleic acid sequences by creating a peptide barcode library (peptide barcode library) with low variability in ionization efficiency in a mass spectrometer. Specifically, a peptide barcode library with low variability in ionization efficiency in a mass spectrometer is realized by ensuring that all peptide barcodes contained in the peptide barcode library have the same amino acid sequence (fixed sequence).

[0010] In one embodiment, the present invention is a method for optimizing the sequence of nucleic acids, A step of preparing a nucleic acid sequence comprising a candidate nucleic acid sequence containing a non-translated region sequence and a target protein nucleic acid sequence encoding the target protein, and a barcode nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein. A step of expressing a protein from the nucleic acid sequence, A step of separating the peptide barcode from the protein, A step of analyzing the separated peptide barcode, Based on the results of the above analysis, the relationship between the expression of the target protein and the candidate nucleic acid sequence is obtained, and the optimal candidate nucleic acid sequence is selected. The present invention relates to a method comprising a peptide barcode comprising multiple types of peptide barcodes contained in a peptide barcode library, wherein all of the peptide barcodes contained in the peptide barcode library contain a fixed sequence which is a common amino acid sequence.

[0011] In another embodiment, the present invention is: An insertion site for inserting a candidate nucleic acid sequence containing the sequence encoding the target protein and the sequence of the untranslated region, A sequence encoding a peptide barcode containing two or more amino acids and A kit for nucleic acid sequence optimization comprising multiple expression cassettes, The plurality of expression cassettes each have sequences that encode different peptide barcodes, When the expression cassette is expressed, the target protein inserted at the insertion site and the peptide barcode are linked and expressed. The different peptide barcodes include a plurality of types of peptide barcodes included in a peptide barcode library, and all the peptide barcodes included in the peptide barcode library include a fixed sequence that is a common amino acid sequence. It relates to a kit.

Effect of the Invention

[0012] The present invention provides a method and a kit for optimizing the sequence of a nucleic acid. By using the method and kit of the present invention, the quantification accuracy of peptide barcodes in a mass spectrometer is improved. Thereby, quantitative comparison of the expression levels of nucleic acid sequences becomes possible, and the base sequence of a nucleic acid, for example, an mRNA drug, can be optimized with high precision.

Brief Description of the Drawings

[0013] [Figure 1] It is a diagram showing an outline of a process in the first embodiment of the present invention. [Figure 2] It is a diagram showing an example of a preparation process in the first embodiment of the present invention. [Figure 3] It is a diagram showing a configuration example of the amino acid sequence of a protein obtained by an expression process in the first embodiment of the present invention. [Figure 4] It is a conceptual diagram showing the hydrophilicity retention coefficient and mass distribution of peptides constituting a peptide barcode library measured by an analysis process in the first embodiment of the present invention. [Figure 5] It is a graph showing the difference in the frequency distribution of measurement efficiency of peptides constituting a peptide barcode library measured by an analysis process in the first embodiment of the present invention, depending on the presence or absence of a fixed sequence. [Figure 6] It is a diagram showing an example of a nucleic acid obtained by a preparation process in the second embodiment of the present invention. [Figure 7]A graph showing the adequacy of the accuracy of the converter used when obtaining the fixed array and the insertion array in the third embodiment of the present invention.

Embodiments for Carrying Out the Invention

[0014] The present invention relates to a method and a kit for optimizing a nucleic acid sequence. In the present invention, a sequence encoding a target protein and a sequence affecting expression (such as a sequence in an untranslated region) are ligated and expressed with a sequence encoding a distinguishable peptide barcode, and by analyzing the peptide barcode portion in the expressed protein, a nucleic acid optimal for the expression of the target protein can be selected based on the peptide barcode indicating desired expression (such as high expression or long-term expression). The peptide barcode used includes a plurality of types of peptide barcodes constituting a peptide barcode library, and all the peptide barcodes included in the peptide barcode library include a fixed sequence consisting of a common amino acid sequence.

[0015] In one aspect, the present invention is a method for optimizing a nucleic acid sequence, comprising: a step of preparing a nucleic acid sequence (preparation step) including a candidate nucleic acid sequence containing a sequence in an untranslated region and a target protein nucleic acid which is a sequence encoding a target protein, and a barcode nucleic acid which is a nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein; a step of expressing a protein from the nucleic acid sequence (expression step); a step of separating the peptide barcode from the protein (separation step); a step of analyzing the separated peptide barcode (analysis step); a step of obtaining the relationship between the expression of the target protein and the candidate nucleic acid sequence based on the result of the analysis and selecting an optimal candidate nucleic acid sequence (selection step) The present invention provides a method comprising a peptide barcode comprising multiple types of peptide barcodes contained in a peptide barcode library, wherein all of the peptide barcodes contained in the peptide barcode library contain a fixed sequence which is a common amino acid sequence.

[0016] Nucleic acid sequence optimization refers to optimizing the sequence of a nucleic acid in order to express the target protein encoded by that nucleic acid. Protein expression can vary depending on the sequence of the protein-coding sequence and the untranslated region connected to it. It is preferable to obtain a nucleic acid having an optimized sequence that can achieve the desired protein expression.

[0017] The nucleic acid may be RNA or DNA, as long as it is a nucleic acid for which sequence optimization for protein expression is desired. Preferably, the nucleic acid is RNA, for example, an mRNA drug. The target protein encoded by the nucleic acid is not particularly limited as long as it is a protein for which expression is desired. In the case of nucleic acid drugs, for example mRNA drugs, the target protein may include, for example, an immunogenic protein for a vaccine for infectious diseases (viruses, bacteria, fungi, etc.), or a protein specifically expressed in cancer cells for a cancer vaccine. Alternatively, the target protein may be a protein intended to be expressed on a large scale in cells or in a cell-free expression system. In this specification, "nucleic acid sequence" may be either DNA or RNA, and for example, "nucleic acid sequence containing a sequence encoding a target protein" includes both the DNA sequence and the mRNA sequence obtained by further reverse transcription therefrom.

[0018] The method according to the present invention prepares a nucleic acid sequence comprising a candidate nucleic acid sequence including a non-translated region sequence and a target protein nucleic acid sequence encoding the target protein, and a barcode nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein (preparation step).

[0019] The untranslated region (UTR) includes the 5'UTR and 3'UTR. The untranslated region is a region that affects protein expression. Therefore, it is preferable to optimize the sequence of the untranslated region for desired protein expression. Optimization may be performed on only the 5'UTR, only the 3'UTR, or both the 5'UTR and 3'UTR. Candidate nucleic acid sequences for the 5'UTR and 3'UTR may be known UTR sequences or their variants, or random sequences.

[0020] In the method according to the present invention, the sequence encoding the target protein may be the same nucleic acid sequence. Alternatively, because the same amino acid can be encoded by different codons due to genetic degeneracy, even when encoding the same protein, different nucleic acid sequences may be present, and different sequences may result in different protein expression. Therefore, the sequence encoding the target protein can also be subject to optimization (codon optimization).

[0021] In the method according to the present invention, a candidate nucleic acid sequence, which includes at least an untranslated region sequence and a target protein nucleic acid sequence encoding the target protein, is linked to a barcode nucleic acid sequence encoding a peptide barcode. A peptide barcode is a peptide consisting of two or more amino acids, and each peptide barcode is identifiable in the step of analyzing the peptide barcode. A peptide barcode contains, for example, 5 to 40, preferably 10 to 30, amino acids. It is preferable that the peptide barcode has a length and composition that does not significantly affect the expression of the target protein.

[0022] In this invention, the peptide barcode includes multiple types of peptide barcodes contained in a peptide barcode library, and all peptide barcodes contained in the peptide barcode library include a fixed sequence which is a common amino acid sequence. Here, the peptide barcode library refers to a group of peptides composed of multiple different types of peptide barcodes.

[0023] In one embodiment, the peptide barcode includes a random sequence consisting of random amino acid sequences and a fixed sequence. The random sequence is not particularly limited as long as it has the length and composition necessary to identify the candidate nucleic acid sequence linked to the peptide barcode nucleic acid. For example, the random sequence of the peptide barcode consists of, for example, 2 to 40 amino acids, preferably 6 to 30 amino acids. It is preferable that the random sequences of each peptide barcode in the peptide barcode library are different sequences without overlap, but even if there is overlap, the overlapping peptide barcodes can be excluded in the analysis and selection steps described later.

[0024] In one embodiment, the random sequence of the peptide barcode contains the amino acids used in the random sequence with equal probability. In one embodiment, the random sequence contains amino acids D, F, W, and / or P, each with a probability of 3% or less.

[0025] The fixed sequence of a peptide barcode is an amino acid sequence common to all peptide barcodes in the peptide barcode library, consisting of one or more amino acids, preferably two or more. The fixed sequence is not limited to this purpose, but is added to suppress the sensitivity variability that may occur during the analysis of peptide barcodes by mass spectrometer, as described later, due to differences in the characteristics of the amino acids that make up the sequence of each peptide barcode in the peptide barcode library. Since the variability in the sensitivity of the mass spectrometer can be caused by the magnitude of the ionization efficiency variability of the peptide barcode being measured, a fixed sequence that reduces the dispersion of ionization efficiency is added to all peptide barcodes in the peptide barcode library in common. By using a peptide barcode library containing a group of peptides with low ionization efficiency dispersion, the quantitative accuracy of the peptide barcode is improved.

[0026] In one embodiment, the fixed sequence of the peptide barcode is a sequence in which the ionization efficiency of the peptide group containing the fixed sequence is small in a converter that takes an amino acid sequence or an amino acid sequence and charge as input and outputs the ionization efficiency of the amino acid sequence in a mass spectrometer. Specific fixed sequences that can be used in the present invention are not limited to those containing R at the N terminus and RP at the C terminus, and those containing RPR at the C terminus.

[0027] In one embodiment, the peptide barcode includes a random sequence consisting of random amino acid sequences, an insertion sequence consisting of amino acid sequences common to some peptide barcodes in the peptide barcode library, and a fixed sequence.

[0028] The insertion sequence of a peptide barcode is an amino acid sequence common to some of the peptide barcodes in the peptide barcode library, consisting of one or two amino acids, but may also consist of three or more amino acids. By using amino acids that have little effect on ionization efficiency and have a significant effect on the hydrophilicity retention coefficient and mass-to-charge ratio that define the measurement range of the analyzer (e.g., a liquid chromatography-mass spectrometer), the peptide barcode library can be widely and low-densityly distributed within the measurement range of the analyzer, enabling high-throughput measurement. The peptide barcode library may include peptide barcodes without an insertion sequence, peptide barcodes with a certain insertion sequence, and peptide barcodes with another insertion sequence.

[0029] In one embodiment, the insertion sequence of the peptide barcode is a sequence that has low dispersion in the ionization efficiency of the peptide group including the insertion sequence, in a converter that takes an amino acid sequence or an amino acid sequence and charge as input and outputs the ionization efficiency of the amino acid sequence in a mass spectrometer. By making the insertion sequence, in addition to the fixed sequence described above, a sequence with low dispersion in ionization efficiency, the quantitative accuracy of the peptide barcode is further improved.

[0030] In one embodiment, the insertion sequence is selected from the group consisting of L, Q, E, EQ, QL, and WP. In one embodiment, the insertion sequence is (i) L, Q, and E; (ii) Permutations of length 2 or 3 that include one or more of Q, E, D, P, G, and A; and (iii) Permutations of length 2 to 4 that include one or more of Q, E, D, P, G, and A, and one or more of L, W, F, V, and Y. It is selected from the group consisting of the following.

[0031] In one embodiment, the peptide barcode does not contain K in particular in its fixed sequence.

[0032] In one embodiment, when a mass spectrometer is used in the step of analyzing peptide barcodes, as described later, the peptide barcodes included in the peptide barcode library may include sequences that achieve both low dispersion and a large average in the ionization efficiency of the peptide barcode group. That is, the dispersion and average of the ionization efficiency of the peptide barcode group can be optimized. Here, it is preferable that such sequences are longer than the lengths of other sequences in the peptide barcode. For example, at least a portion of the peptide barcodes included in the peptide barcode library may include the same sequence, and that sequence can achieve both low dispersion and a large average in the ionization efficiency of the peptide barcode group. That is, multiple peptide barcodes (peptide barcode groups) may have a common sequence in part that achieves both low dispersion and a large average in the ionization efficiency. When using a peptide group containing such a sequence, not only the quantitative sensitivity but also the sensitivity of the peptide barcode is improved.

[0033] The untranslated region sequence, the sequence encoding the target protein, and the sequence encoding the peptide barcode can be prepared by methods known in the art.

[0034] Furthermore, methods for linking candidate nucleic acid sequences with nucleic acid sequences encoding peptide barcodes are also well known in the art. For example, the sequence encoding the target protein and the sequence encoding the peptide barcode are linked directly or indirectly so that the target protein and the peptide barcode are expressed as a fusion protein. Indirect linking can be done via, for example, a sequence encoding an amino acid recognized by a protease, a spacer sequence, etc. Examples of amino acids recognized by proteases, though not limited to, include the amino acid DDDDK (SEQ ID NO: 18) recognized by enterokinase, an amino acid recognized by trypsin, an amino acid recognized by thrombin, and an amino acid recognized by factor Xa. In one embodiment, the sequence encoding the peptide barcode is linked to the sequence encoding the target protein via a sequence encoding an amino acid recognized by a protease.

[0035] A candidate nucleic acid sequence comprising a target protein nucleic acid, which is a sequence of an untranslated region and a sequence encoding the target protein, and a barcode nucleic acid, which is a nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein, may further include other sequences. In one embodiment, the nucleic acid sequence encoding the peptide barcode is linked to a sequence encoding a purification tag. The purification tag is not particularly limited as long as it is a tag conventionally used in the art, and examples include His tags, HQ tags, HN tags, etc. (purifiable by metal ions); FLAG tags, etc. (purifiable by affinity chromatography); Myc tags, etc. (purifiable by antibodies). By using the purification tag, the peptide barcode to which the purification tag is linked, or the peptide barcode and the target protein, can be easily purified and recovered. The linkage between the sequence encoding the peptide barcode and the purification tag may be direct or indirect; for example, the sequence encoding the peptide barcode is linked to the sequence encoding the purification tag via a sequence encoding an amino acid recognized by a protease. In other words, in one embodiment, a sequence encoding an amino acid recognized by a protease exists between the sequence encoding the peptide barcode and the sequence encoding the purification tag. This makes it possible to cleave the purification tag using a protease, thus avoiding the influence of the purification tag during the peptide barcode analysis process.

[0036] Alternatively, a nucleic acid sequence comprising a candidate nucleic acid sequence including a non-translated region sequence and a target protein nucleic acid sequence encoding the target protein, and a barcode nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein, may be capped and / or polyA-linked.

[0037] In one embodiment, in the preparation step of the method according to the present invention, a candidate nucleic acid sequence or a nucleic acid sequence encoding a target protein is amplified using a primer to which a nucleic acid sequence encoding a peptide barcode is attached, thereby preparing a nucleic acid sequence (including a candidate nucleic acid sequence containing a non-translated region sequence and a sequence encoding a target protein, and a nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein).

[0038] In one embodiment, in the preparation step of the method according to the present invention, a nucleic acid sequence (including a candidate nucleic acid sequence containing an untranslated region sequence and a sequence encoding the target protein, and a sequence encoding a peptide barcode directly or indirectly linked to the target protein) is prepared by amplifying the sequence encoding the target protein using a primer to which an untranslated region sequence has been attached.

[0039] In one embodiment, in the preparation step of the method according to the present invention, the nucleic acid sequence is prepared by amplifying a plasmid vector containing a nucleic acid sequence (a candidate nucleic acid sequence including a non-coding region sequence and a sequence encoding the target protein, and a sequence encoding a peptide barcode directly or indirectly linked to the target protein).

[0040] In one embodiment, in the preparation step of the method according to the present invention, Plasmid vectors containing the sequence of the untranslated region, the 5'UTR, plasmid vectors containing the sequence encoding the target protein, and plasmid vectors containing the sequence encoding the peptide barcode were constructed. From these vectors, plasmid vectors are constructed containing candidate nucleic acid sequences including the sequence of the untranslated region and the sequence encoding the target protein, and nucleic acid sequences including the sequence encoding the peptide barcode. The nucleic acid sequence is prepared by amplifying a plasmid vector containing the aforementioned nucleic acid sequence.

[0041] In addition to the plasmid vector described above, a plasmid vector containing the sequence of the 3'UTR, which is the untranslated region, may be used to further construct a plasmid vector containing the nucleic acid sequence described above. The sequence of the 5'UTR and / or the 3'UTR may contain a specific candidate nucleic acid sequence or may be a random sequence.

[0042] In one embodiment, in the preparation step of the method of the present invention, One or more sequences of the untranslated region, one or more sequences encoding the target protein, and sequences encoding multiple different peptide barcodes are attached to a plasmid vector using a homologous sequence DNA assembly method to create a plasmid vector containing multiple combinations of sequences. By amplifying a plasmid vector, a candidate nucleic acid sequence is prepared that includes a sequence of the untranslated region and a sequence encoding the target protein, as well as a sequence encoding a peptide barcode directly or indirectly linked to the target protein.

[0043] To construct plasmid vectors containing the above nucleic acid sequences, for example, the Gibson Assembly method using homologous sequences and the ligation method using blunt ends can be used, and both methods are known in the art. Preferably, the Gibson Assembly method is used to construct plasmid vectors containing candidate nucleic acid sequences including an untranslated region sequence and a sequence encoding the target protein, and nucleic acid sequences including a sequence encoding a peptide barcode, from DNA containing each vector or sequence. The Gibson Assembly method is a method used to ligate multiple DNA fragments, and can ligate DNA fragments having homologous sequences of a specific length (approximately 15-20 bases) at their ends. Therefore, when using the Gibson Assembly method, plasmid vectors or DNA containing each sequence are designed to contain homologous sequences in the ligation region.

[0044] In one embodiment, the homologous sequence of the 5'UTR sequence and the plasmid vector includes the T7 promoter sequence. In one embodiment, the homologous sequence of the 5'UTR sequence and the sequence encoding the target protein includes the start codon. In one embodiment, the homologous sequence of the sequence encoding the target protein and the sequence encoding the peptide barcode includes the protease recognition sequence (the sequence encoding the amino acid recognized by the protease), for example, the sequence encoding the amino acid recognized by enterokinase. In one embodiment, the homologous sequence of the sequence encoding the peptide barcode and the 3'UTR sequence includes the stop codon. In one embodiment, the homologous sequence of the 3'UTR sequence and the plasmid vector includes a portion of the plasmid vector sequence (for example, about 15-20 bases).

[0045] When randomly binding multiple sequences to a plasmid vector, it is preferable to perform the DNA assembly method under conditions where the number of peptide barcode types is greater than the product of the number of 5'UTR sequence types, the number of sequence types encoding the target protein, and the number of 3'UTR sequence types.

[0046] When the nucleic acid whose sequence is to be optimized is RNA such as mRNA, the nucleic acid sequence can be prepared as RNA (mRNA) containing the candidate nucleic acid sequence by reverse transcription (e.g., in vitro transcription (IVT)) of the amplified nucleic acid sequence (DNA).

[0047] In one embodiment, in the preparation step of the method according to the present invention, a plurality of nucleic acid sequences are prepared, each containing different untranslated region sequences and sequences encoding different peptide barcodes. The plurality of untranslated region sequences to be optimized are each associated with sequences encoding different peptide barcodes.

[0048] In one embodiment, the preparation step of the method according to the present invention involves preparing a plurality of nucleic acid sequences, each containing a different sequence encoding a target protein and a sequence encoding a different peptide barcode. The plurality of sequences encoding the target protein to be optimized are each associated with a sequence encoding a different peptide barcode.

[0049] Next, the prepared nucleic acid sequence is used to express the protein (expression step). In one embodiment, the expression step of the method according to the present invention is performed in a cell. The cell is not particularly limited and may be a prokaryotic cell, such as a bacterial cell (E. coli, etc.), or a eukaryotic cell, such as a fungal cell (yeast, etc.), an insect cell, or a mammalian cell (human cell, etc.). Preferably, it is a cell that is intended to express the target protein, and it is possible to select the sequence that is optimal for expression in that cell. For example, in the case of a nucleic acid drug for administration to humans, it is preferable to optimize the expression in human cells.

[0050] In another embodiment, the expression step of the method according to the present invention is carried out in a cell-free expression system. The cell-free expression system is also not particularly limited, and an appropriate system can be used from among those known in the art, such as expression systems derived from E. coli, wheat germ, rabbit reticulocytes, and insect cells. Preferably, a cell-free expression system that mimics the cells in which the target protein is ultimately to be expressed is used. For example, when the target protein is to be expressed in E. coli, nucleic acid sequence optimization can be performed simply and quickly by using an E. coli-derived expression system.

[0051] Next, the peptide barcode is separated from the expressed protein (separation step). Peptide barcodes can be separated, for example, by using a protease if the expressed protein contains amino acids recognized by a protease.

[0052] The expressed protein and / or isolated peptide barcode may be purified at an appropriate stage, for example, before the analysis of the peptide barcode. The purification step can be carried out using methods used for purifying proteins, such as purification using antibodies that bind to the target protein. If a sequence encoding a purification tag is concatenated to the nucleic acid sequence described above, the expressed protein and / or isolated peptide barcode can be easily purified using the purification tag.

[0053] Next, the separated peptide barcodes are analyzed (analysis step). In one embodiment, the peptide barcodes are analyzed using a mass spectrometer. In another embodiment, the peptide barcodes are analyzed using a known protein analysis method (e.g., immunoassay such as ELISA or immunoblotting). Apparatus and methods for analyzing peptide barcodes are well known in the art, and those skilled in the art can appropriately select the apparatus and method to be used depending on the composition and properties of the peptide barcodes used.

[0054] In one embodiment, in the analytical step of the method according to the present invention, the ionic intensity for each peptide barcode obtained by a mass spectrometer is normalized by the ionization efficiency for each peptide barcode obtained in advance, and the amount present for each peptide barcode is estimated from the normalized ionic intensity.

[0055] In one embodiment, an m / z peak list detected by a mass spectrometer is prepared in advance based on the amino acid sequence length of the peptide barcode, and in the analysis step of the method of the present invention, the amount of peptide barcode is estimated using the ionic intensity of the m / z values ​​in the m / z peak list.

[0056] In the method according to the present invention, the relationship between the expression of the target protein and candidate nucleic acid sequences is obtained based on the results of peptide barcode analysis, and the optimal candidate nucleic acid sequence is selected (selection step). That is, since the peptide barcode corresponds to the candidate nucleic acid sequence, the relationship between the expression of the target protein and the candidate nucleic acid sequence can be obtained by analyzing the peptide barcode. For example, information can be obtained regarding candidate nucleic acid sequences corresponding to target proteins with high or low expression levels, or candidate nucleic acid sequences corresponding to target proteins with long or short expression periods. The optimal candidate nucleic acid sequence is selected according to the expression level of the desired target protein. In one embodiment, a candidate nucleic acid sequence with a high expression level of the target protein is selected based on the results of peptide barcode analysis. In another embodiment, a candidate nucleic acid sequence with a long expression period of the target protein is selected based on the results of peptide barcode analysis.

[0057] The method according to the present invention may further include the step of identifying the selected candidate nucleic acid sequence by designing a primer based on the amino acid sequence of a peptide barcode corresponding to the selected candidate nucleic acid sequence, amplifying or reverse transcribing at least a portion of the nucleic acid sequence, and sequencing the obtained sequence. When a random sequence is used as the candidate nucleic acid sequence (e.g., 5'UTR and / or 3'UTR), the sequence of the selected candidate nucleic acid sequence may not be known by analysis of the peptide barcode. Therefore, the sequence of the selected candidate nucleic acid sequence can be determined by sequencing the sequence obtained by amplification or reverse transcription using a primer designed based on the amino acid sequence of the peptide barcode.

[0058] The method according to the present invention assigns an identifiable peptide barcode to the expressed protein, making it possible to distinguish expression products even when multiple nucleic acids containing candidate nucleic acid sequences for optimization are expressed simultaneously. By analyzing the separated peptide barcodes together, it becomes possible to simultaneously examine (screen) multiple candidate nucleic acid sequences.

[0059] In another embodiment, the present invention provides a kit for nucleic acid sequence optimization. Specifically, An insertion site for inserting a candidate nucleic acid sequence containing the sequence encoding the target protein and the sequence of the untranslated region, A sequence encoding a peptide barcode containing two or more amino acids and A kit for nucleic acid sequence optimization comprising multiple expression cassettes, The plurality of expression cassettes each have sequences that encode different peptide barcodes, When the expression cassette is expressed, the target protein inserted into the insertion site and the peptide barcode are linked and expressed together. The kit provides a set in which the different peptide barcodes include multiple types of peptide barcodes contained in a peptide barcode library, and all of the peptide barcodes contained in the peptide barcode library include a fixed sequence which is a common amino acid sequence.

[0060] In one embodiment, the peptide barcode may further include an insertion sequence consisting of an amino acid sequence common to some peptide barcodes in the peptide barcode library.

[0061] The expression cassette can be in a form suitable for cells expressing the target protein or for a cell-free expression system, and can be, for example, in the form of a linear nucleic acid or a vector. The expression cassette may be either DNA or RNA, but DNA is preferred for ease of handling. When RNA sequence optimization is performed, RNA can be obtained by performing a reverse transcription reaction from the DNA expression cassette. Such operations are well known in the art. The insertion site can be a site for inserting a sequence into the expression cassette, such as a restriction site, a multicloning site, or a homologous sequence.

[0062] By using the kit according to the present invention, it is possible to express a target protein linked to a peptide barcode simply by ligating a nucleic acid for which sequence optimization is desired to the insertion site of the expression cassette, and then easily examine the expression of the target protein based on that peptide barcode.

[0063] The kit according to the present invention is preferably a kit for carrying out the method of the present invention (method for optimizing nucleic acid sequences) described above. By using such a kit, the method of the present invention can be carried out simply and efficiently.

[0064] Hereinafter, embodiments for carrying out the present invention (hereinafter referred to as "embodiments") will be described with reference to the attached drawings. While these embodiments show specific examples in accordance with the principles of the present invention, they are for the purpose of understanding the present invention and are not intended to restrict its interpretation. Modifications resulting from combinations or substitutions of the following embodiments with known technologies are also included within the scope of the present invention. In all drawings used to illustrate the embodiments, components having the same function are denoted by the same reference numerals, and their repeated descriptions are omitted.

[0065] [First Embodiment] Hereinafter, a first embodiment of the present invention will be described with reference to Figures 1 to 5. Figure 1 is a diagram illustrating the overview of the process in the first embodiment of the present invention. In the first embodiment, an example is described in which the nucleic acid of a target protein, which includes a sequence encoding the target protein, is optimized in terms of its high expression level. Although the amino acid sequence of a protein is uniquely determined, due to degeneracy, the correspondence between amino acids and nucleic acid sequences is one-to-many, so multiple nucleic acid sequences correspond to the same protein. In the first embodiment, the expression levels of multiple nucleic acid sequences corresponding to the same protein are measured, and the nucleic acid sequence with the highest expression level is determined to be optimal.

[0066] 1 is an example of a nucleic acid sequence obtained by the preparation step, 110 is the target protein nucleic acid, 120 is the barcode nucleic acid, 2 is an example of a protein obtained by the expression step, 210 is the target protein, 220 is the peptide barcode, 310 is the target protein cleaved by the separation step, 320 is the peptide barcode similarly cleaved, 410 is the protein analysis method used in the analysis step, for example, the measured value of the peptide barcode by a mass spectrometer, 420 is the correspondence knowledge between barcode fragmentation patterns and peptide barcodes necessary to read the peptide barcode and its quantitative value from the measured value of the peptide barcode, 430 is the candidate nucleic acid sequence of the peptide barcode and its quantitative value, in this embodiment the correspondence knowledge between peptide barcodes and candidate nucleic acid sequences necessary to convert them into the target protein nucleic acid sequence and its quantitative value, 440 is the candidate nucleic acid sequence and its quantitative value obtained using 410 to 430, and 5 is an example of the optimal nucleic acid sequence obtained by the selection step.

[0067] The preparation process generates a nucleic acid sequence (1), in this embodiment, mRNA, by a process that is explained in detail using an example in Figure 1. Multiple types of nucleic acids to be optimized, in this embodiment, exist, including the target protein nucleic acid (110) and the barcode nucleic acid (120), and multiple nucleic acid sequences are generated in which one target protein nucleic acid and one barcode nucleic acid are linked. Here, the barcode nucleic acid is linked to the target protein nucleic acid in a way that satisfies the function of a barcode. That is, a one-to-many correspondence between the target protein nucleic acid and the barcode nucleic acid is acceptable, but a many-to-one correspondence is not acceptable. Also, in this embodiment, all of the target protein nucleic acids (110) are different nucleic acid sequences that encode the same protein.

[0068] The expression process generates multiple proteins (2) from a nucleic acid sequence (1), each containing a target protein and a peptide barcode. For example, the expression process involves introducing mRNA (1) into cells, followed by protein expression within the cells, and then extracting the multiple proteins (2) through processes such as cell lysis and centrifugation.

[0069] The separation step separates the peptide barcode (320=220) from the target protein (310=210). The separation step involves, for example, cleaving the peptide barcode from the target protein using a digestive enzyme. As illustrated in detail in Figure 3, the target protein (310=210), the recognition sequence of the digestive enzyme (230), and the peptide barcode (320=220) are linked together in that order from the N-terminus. The digestive enzyme recognizes the recognition sequence and cleaves the peptide at the cleavage site within the recognition sequence, resulting in the separation of the peptide barcode (320).

[0070] The analysis process involves quantifying the peptide barcode using a protein analysis method, such as a mass spectrometer (410), converting the quantified value of the peptide barcode to the quantified value of the candidate nucleic acid sequence using barcode fragmentation pattern-peptide barcode correspondence knowledge (420) and peptide barcode-candidate nucleic acid sequence correspondence knowledge (430), and obtaining a candidate nucleic acid sequence-quantified value correspondence result (440). The quantified value is, for example, the expression level. The barcode fragmentation pattern-peptide barcode correspondence knowledge (420) is, for example, existing publicly known information such as MASCOT. The peptide barcode-candidate nucleic acid sequence correspondence knowledge (430) is the correspondence knowledge between the target protein nucleic acid (110) and the barcode sequence (120), obtained by analyzing the nucleic acid sequence (1) obtained in the preparation process using, for example, NGS (Next Generation Sequencer).

[0071] The selection step involves selecting, for example, the candidate nucleic acid sequence with the highest expression level among the candidate nucleic acid sequence-quantification value correspondence results (440) generated by the analysis step, which in this embodiment is the target protein nucleic acid 2, as the optimal nucleic acid sequence (5).

[0072] Figure 2 shows an example of the preparation process in the first embodiment of the present invention. 130 is a candidate nucleic acid sequence library, which is a set of realizations of multiple candidate nucleic acid sequences; 131 is a DNA plasmid, which is a realization of a candidate nucleic acid sequence; 132 is a candidate nucleic acid sequence contained in the DNA plasmid, which in this embodiment is the target protein DNA, a realization of the target protein nucleic acid; 140 is a peptide barcode library, which is a set of realizations of multiple barcode nucleic acids; 141 is a DNA plasmid, which is a realization of a barcode nucleic acid; 142 is a barcode DNA, which is a barcode nucleic acid contained in the DNA plasmid; 150 is a DNA plasmid obtained by linking a DNA plasmid of a candidate nucleic acid sequence with a DNA plasmid of a barcode nucleic acid; 151 is the same target protein DNA as 132; 152 is the same barcode DNA as 142; 1 is an example of a nucleic acid sequence obtained by the preparation step; 110 is the target protein RNA, a realization of the target protein nucleic acid; and 120 is a barcode RNA, a realization of a barcode nucleic acid.

[0073] The preparation process begins by preparing multiple candidate nucleic acid sequences to be evaluated. In this embodiment, all candidate nucleic acid sequences correspond to the same protein. Next, the candidate nucleic acid sequences are synthesized as linear DNA, for example, by the phosphoramidite method, and ligated into a plasmid vector to produce a DNA plasmid (131). This process is repeated for the number of candidate nucleic acid sequences to produce a candidate nucleic acid sequence library (130). Subsequently, multiple barcode nucleic acids are prepared, and a peptide barcode library (140) is produced in the same manner. Here, all barcode nucleic acids correspond to different barcodes. The number of barcode nucleic acid types is set to be sufficiently greater than the number of candidate nucleic acid sequences to avoid a majority-to-one correspondence between the target protein nucleic acid and the barcode nucleic acid. Next, the candidate nucleic acid sequence library (130) and the peptide barcode library (140) are ligated to produce a plasmid (150) containing the ligated target protein DNA and barcode DNA. This is amplified, for example, by PCR (Polymerase Chain Reaction), and then transcribed, for example, by IVT (In Vitro Transcription), to produce a nucleic acid sequence (1) in which the target protein RNA (110) and barcode RNA (120) are linked.

[0074] Figure 3 shows an example of the amino acid sequence of a protein produced by the expression process in the first embodiment of the present invention. The protein comprises a target protein and a peptide barcode, with an enzyme recognition sequence between them. The peptide barcode consists of a random sequence, a fixed sequence, and possibly an inserted sequence.

[0075] 210 is the target protein, 220 is the peptide barcode, 230 is the recognition sequence of the digestive enzyme described as an example of the separation process, which is the recognition sequence (SEQ ID NO: 18) when the digestive enzyme is enterokinase, 221 is the random sequence, 222 is the fixed sequence, and 223 is the insertion sequence. 240 is an example of RPR with the fixed sequence at the C-terminus, 241 is an example of the fixed sequence with R at the N-terminus and RP at the C-terminus, 250 to 254 are examples of RPR with the fixed sequence at the C-terminus, with 250 being an example without an insertion sequence, 251 being an example of L, 252 being an example of Q, 253 being an example of EQ, and 254 being an example of QL.

[0076] The random sequence (221) is a sequence obtained by randomly selecting and ligating a specified number of uncharged amino acids, such as E, D, A, G, L, V, Q, P, Y, W, and F, with equal probability, for example, 6 amino acids, for use in peptide barcodes. Peptide barcodes (240) with a fixed sequence of RPR are suitable when the length of the random sequence (221) is long, typically 7 or more, while peptide barcodes (241) with a fixed sequence whose N-terminus is R and C-terminus is RP are suitable when the length of the random sequence (221) is short, typically 6 or less. Peptide barcodes (250) without an inserted sequence are suitable when the number of candidate nucleic acid sequences to be evaluated is small, specifically when the candidate nucleic acid sequences are distributed at a density below the measurement resolution on the measurement axis of a protein analysis method (410), such as a mass spectrometer. As the number of candidate nucleic acid sequences increases, peptide barcode libraries with added peptide barcodes of length 1 (250, 251, 252) and peptide barcode libraries with added peptide barcodes of length 1 and length 2 (250, 251, 252, 253, 254) are preferred. In this embodiment, where the digestive enzyme is enterokinase, the digestive enzyme cleaves the sequence at the C-terminus of its recognition sequence (230), so in the separation step, the amino acid sequence in which 221 and 222 are linked is separated as a peptide barcode.

[0077] Figure 4 is a conceptual diagram showing the distribution of hydrophilicity retention coefficient and mass of peptide barcodes constituting a peptide barcode library measured by the analytical step in the first embodiment of the present invention. In the analytical step of the method according to the present invention, an example is shown where the domain is two-dimensional, and one of the two dimensions is the hydrophilicity retention coefficient and the other is mass or mass-to-charge ratio, i.e., the analyzer is a liquid chromatography-mass spectrometer.

[0078] 400 corresponds to peptide barcode group 250 without an inserted sequence in Figure 3. Similarly, 401 corresponds to peptide barcode group 251 with the inserted sequence L, 402 corresponds to peptide barcode group 252 with the inserted sequence Q, 403 corresponds to peptide barcode group 253 with the inserted sequence EQ, and 404 corresponds to peptide barcode group 254 with the inserted sequence QL. Since the retention coefficients Rc for amino acids L, Q, and E are large, small, and small respectively (9.6, -0.9, and 0.0), peptide barcode groups 401 and 404, which contain L in the inserted sequence, have a distribution (401, 404) that is shifted by a certain amount on the horizontal axis representing the hydrophilicity retention coefficient compared to peptide barcode group 400 without an inserted sequence. Furthermore, peptide barcodes containing only Q or E, which have low retention coefficients, show a distribution shifted by a certain amount on the vertical axis representing mass compared to peptide barcode group 400, when the amino acid length of the inserted sequence is 1 (402), and when the amino acid length of the inserted sequence is 2 (403), the distribution shifted by approximately twice the aforementioned certain amount on the vertical axis representing mass compared to peptide barcode group 400, when the amino acid length of the inserted sequence is 2 (403). Since peptide barcode groups 400 to 404 each contain random sequences, they have a Gaussian distribution, with a high distribution density at the center of the distribution. However, the entire peptide barcode library consisting of peptide barcode groups 400 to 404 has a mixed Gaussian distribution, resulting in a distribution shape closer to a uniform distribution than a Gaussian distribution.

[0079] Figure 5 shows the difference in the frequency distribution of measurement efficiency of peptides constituting a peptide barcode library measured by the analytical step in the first embodiment of the present invention, depending on the presence or absence of a fixed sequence. An example is shown where the measuring instrument is a mass spectrometer, the measurement efficiency is the ionization tendency, the fixed sequence is R at the N-terminus and RP at the C-terminus (241), the random sequence is a sequence in which E, D, A, G, L, V, Q, P, Y, W, F are randomly linked with equal probability, and the insertion sequences are L, Q, EQ, QL. 510 to 514 are cases without a fixed sequence, and 520 to 524 are cases with a fixed sequence. 510 and 520 are peptide barcodes without an insertion sequence, 511 and 521 are peptide barcodes with insertion sequence L, 512 and 522 are peptide barcodes with insertion sequence Q, 513 and 523 are peptide barcodes with insertion sequence EQ, and 514 and 524 are peptide barcodes with insertion sequence QL. For example, 520 contains REGFGEARP (SEQ ID NO: 1), 521 contains REGFGEALRP (SEQ ID NO: 2), 522 contains REGFGEAQRP (SEQ ID NO: 3), 523 contains REGFGEAEQRP (SEQ ID NO: 4), and 524 contains REGFGEAQLRP (SEQ ID NO: 5). The ionization efficiency is a numerical calculation result obtained using a converter, which takes an amino acid sequence as input and outputs the ionization efficiency of the amino acid sequence in a mass spectrometer, as will be described later using Figure 7. The variance of ionization efficiency is smaller for 520 to 524, which have fixed sequences, than for 510 to 514, which do not have fixed sequences, even when comparing peptide barcode groups with the same insertion sequence, such as 511 and 521, or when comparing the union of 510 to 514 with the union of 520 to 524. In addition, in this embodiment, the average value of the ionization efficiency is larger.

[0080] The first embodiment has been described above using Figures 1 to 5, but this does not limit the embodiments.

[0081] Specifically, in the preparation process, the nucleic acid produced may be DNA or RNA. Among the candidate nucleic acid sequences, the sequences with different sequences may be 5'UTR sequences, 3'UTR sequences, or the target protein nucleic acid. The candidate nucleic acid sequences may also include sequences other than 5'UTR, 3'UTR, and the target protein nucleic acid, such as spacer sequences, purification tags, polyA, and capping structures. The target protein is not limited to immunogenic proteins, and may be proteins expressed in large quantities in cells or cell-free expression systems. The form of linking the candidate nucleic acid sequence to the sequence encoding the peptide barcode may be direct or indirect, and the linking method may be a method known in the art. Indirect linking may include, for example, sequences encoding amino acids recognized by proteases, spacer sequences, etc., and among the former, the protease may be, for example, enterokinase, trypsin, thrombin, or factor Xa.

[0082] In the expression process, the expression system can be any known system. For example, it can be a cell-based or cell-free expression system. If it is a cell-based system, it can be either a prokaryotic or eukaryotic cell. The former can be bacterial cells such as E. coli, and the latter can be fungal cells such as yeast, insect cells, or mammalian cells (such as human cells). If it is a cell-free expression system, it can be any expression system derived from E. coli, wheat germ, rabbit reticulocytes, or insect cells. Preferably, it is a system intended to express the target protein. For example, in the case of nucleic acid drugs intended for administration to humans, human cells are preferred.

[0083] In the analytical process, the protein analysis method may be any known method, such as mass spectrometry, ELISA, or immunoblotting. The barcode fragmentation pattern-peptide barcode correspondence knowledge may be existing publicly known information, such as MASCOT, if the protein analysis method is mass spectrometry, or it may be a value measured in advance for each peptide barcode using the protein analysis method. The quantitative value may be either expression level or expression period.

[0084] The selection process is not limited to candidate nucleic acid sequences considered optimal; for example, it may include candidate nucleic acid sequences with high expression levels, long expression periods, or short expression periods.

[0085] According to the above embodiments, by expressing multiple candidate nucleic acid sequences in the same expression system and analyzing the expression levels corresponding to multiple candidate nucleic acid sequences in parallel, high-throughput measurement of the expression levels of candidate nucleic acid sequences becomes possible. Furthermore, by including fixed sequences in the peptide barcode library that reduce the variance of the analytical values ​​of the peptide barcodes, the quantitative accuracy of the peptide barcodes is improved, and quantitative comparison of the expression levels of nucleic acid sequences becomes possible.

[0086] More specifically, by constructing the insertion sequence using amino acids that have little impact on ionization efficiency and have a significant impact on the hydrophilicity retention coefficient and mass-to-charge ratio that define the measurement range of an analyzer, such as a liquid chromatography-mass spectrometer, the peptide barcode library can be widely and low-densityly distributed within the analyzer's measurement range, enabling even higher throughput measurements. Furthermore, by constructing the peptide group to include fixed sequences and / or insertion sequences, a sequence group with low dispersion in ionization efficiency is realized, improving the quantitative accuracy of the peptide barcode. In addition, if a sequence group with a high average ionization efficiency is realized, the quantitative sensitivity of the peptide barcode improves.

[0087] [Second Embodiment] A second embodiment of the present invention will be described below with reference to Figure 6. In the first embodiment, an example was described in which the target protein nucleic acid, which is the sequence encoding the target protein, is optimized in terms of high expression level. However, the 5'UTR sequence of the untranslated region, which is a site involved in the regulation of expression, may also be optimized in terms of high expression level of the target protein.

[0088] In the first embodiment, for simplicity, the nucleic acid sequence (1) was shown as a link between the target protein nucleic acid (110) and the barcode nucleic acid (120). However, in the method according to the present invention, the preparation step involves preparing a nucleic acid sequence that includes a candidate nucleic acid sequence containing a 5'UTR sequence, a 3'UTR sequence, and the target protein nucleic acid, which is a sequence encoding the target protein, and a barcode nucleic acid.

[0089] 1 is an example of a nucleic acid sequence obtained by the preparation process, 160 is the 5'UTR sequence, 110 is the target protein nucleic acid, 120 is the barcode nucleic acid, and 161 is the 3'UTR sequence.

[0090] In this embodiment for optimizing the 5'UTR sequence, three candidate nucleic acid sequences are prepared such that only the 5'UTR sequence (160) and the barcode nucleic acid (120) are different.

[0091] According to the above embodiment, it is expected that by setting the nucleic acid sequences to be compared in the candidate nucleic acid sequence according to the functional unit of the nucleic acid, specifically according to the 5'UTR, ORF, and 3'UTR, it will be possible to optimize the nucleic acid sequence for higher performance.

[0092] [Third Embodiment] A third embodiment of the present invention will be described below with reference to Figure 7. The fixed sequence and insertion sequence are sequences that, in a converter that takes an amino acid sequence as input and outputs the ionization efficiency of the amino acid sequence in a mass spectrometer, exhibit low dispersion in the ionization efficiency of the peptide group including the fixed sequence and / or insertion sequence. The converter may use actual measured data or a converter created based on actual measured data. The method for creating the converter from actual measured data may involve information processing such as machine learning or deep learning.

[0093] Figure 7 shows the sufficiency of the accuracy of the converter used to determine the fixed and inserted sequences used in the third embodiment of the present invention.

[0094] The converter is a Deep Learning model called pepper (https: / / doi.org / 10.1021 / acs.jproteome.2c00211), which learns the network using the amino acid sequence, charge, and measurement intensity of a peptide as input and the ionization efficiency of the peptide as output. The measurement intensity from the mass spectrometer is used as dummy data, and the converter uses the amino acid sequence and charge of the peptide as input and the ionization efficiency of the peptide as output. Since different ionization methods are implemented for each mass spectrometer model, it is necessary to retrain the Deep Learning model or perform transfer learning for each ionization method. In this embodiment, the amino acid sequence and charge were input, the ionization efficiency for each amino acid sequence calculated numerically using the COSMO-RS method was output, and transfer learning was performed with the Mean Squared Error of the ionization efficiency as the loss function.

[0095] The amino acid sequences RPALWYFRP (SEQ ID NO: 6), FKVGEKPAY (SEQ ID NO: 7), RWGQPFWRP (SEQ ID NO: 8), PQLKRELLE (SEQ ID NO: 9), RPALWYFLRP (SEQ ID NO: 10), FKVGEKPAYL (SEQ ID NO: 11), RWGQPFWLRP (SEQ ID NO: 12), PQLKRELLEL (SEQ ID NO: 13), LFFFFDYRA (SEQ ID NO: 14), LFFFFDYRAL (SEQ ID NO: 15), RYLYEGARP (SEQ ID NO: 16), and RYLYEGALRP (SEQ ID NO: 17) were measured using a mass spectrometer. The measurement intensity was plotted on the x-axis and the ionization efficiency, which is the output obtained by inputting the amino acid sequence and charge into a converter, on the y-axis in a log-log plot in Figure 7. The correlation coefficient between the measurement intensity and the ionization efficiency calculated by the converter was 0.6. The two samples in the lower right of the figure, which are causing a decrease in accuracy, are thought to be outside the scope of transfer learning application, and the correlation coefficient excluding these two samples is 0.92.

[0096] The above embodiments demonstrate the sufficient accuracy of the converter used to determine the fixed and inserted sequences used in the present invention.

[0097] The above describes an example of a converter created based on actual measurement data, particularly in the case of Deep Learning. In the case of a converter created based on actual measurement data, it is possible to obtain the effects of the present invention regardless of the conditions of the analysis process. On the other hand, when the converter is based on actual measurement data, the sequence that reflects the conditions of the actual analysis process can be obtained, making it possible to obtain the effects of the present invention with high accuracy.

[0098] [Fourth Embodiment] A fourth embodiment of the present invention will be described below. The fixed sequence and insertion sequence are designed to achieve both low dispersion and high average ionization efficiency in the peptide barcode group in a converter that takes an amino acid sequence as input and outputs the ionization efficiency of the amino acid sequence in a mass spectrometer. The insertion sequence is used to increase the size of the peptide barcode library, i.e., to obtain high throughput.

[0099] By using insert sequences, high throughput becomes possible. Furthermore, by optimizing not only the fixed sequences but also the insert sequences, it becomes possible to significantly reduce the variability in the ionization efficiency of the peptide barcode library compared to cases where the present invention is not applied.

[0100] Furthermore, when insert sequences are used, it is possible to further reduce the dispersion of ionization efficiency. Reflecting this, in the converter, when optimizing for both low dispersion and high average value of the ionization efficiency of the peptide barcode group, it becomes possible to achieve a higher level of balance between low dispersion and high average value of ionization efficiency by prioritizing the high average value.

[0101] [Fifth Embodiment] A fifth embodiment of the present invention will be described below. In the preparation process, the random sequence contains the amino acids used in the random sequence with equal probability, or contains D, F, W, and / or P with a probability of 3% or less each.

[0102] When random sequences contain amino acids used in the random sequence with equal probability, the size of the peptide barcode library can be maximized, enabling high throughput. Conversely, when D, F, W, and / or P are each present with a probability of 3% or less, the size of the peptide barcode library is not maximized, resulting in decreased throughput. However, this significantly reduces the variability in the ionization efficiency of the peptide barcode library in the mass spectrometer. In other words, highly accurate quantitative comparison becomes possible even with a small number of candidate nucleic acid sequences. [Explanation of Symbols]

[0103] 1. Nucleic acid sequence obtained by the preparation step 110 Target Protein Nucleic Acids 120 Barcode Nucleic Acids 130 Candidate Nucleic Acid Sequence Library 131 DNA plasmids 132 Target protein DNA 140 Peptide Barcode Library 141 DNA plasmids 142 Barcoded DNA 150 DNA plasmids 151 Target protein DNA 152 Barcoded DNA 160 5'UTR array 161 3'UTR sequence 2. Protein obtained from the expression process 210 Target Protein 220 peptide barcodes 221 random sequence 222 fixed array 223 Insertion Sequence 230 Recognition sequences of digestive enzymes 240-254 Examples of sequence configurations for target protein and peptide barcodes 310 The cleaved target protein 320 Severed peptide barcodes Distribution of hydrophilicity retention coefficient and mass of each peptide barcode 400-404 410 Measurement values ​​of protein analysis methods 420 Knowledge of the correspondence between barcode fragmentation patterns and peptide barcodes 430 Knowledge of the correspondence between peptide barcodes and candidate nucleic acid sequences Correspondence results of 440 candidate nucleic acid sequences and their quantitative values 5. Optimal nucleic acid sequence obtained by the selection process. Frequency distribution of measurement efficiency for peptide barcodes 510-514 and 520-524.

Claims

1. A method for optimizing nucleic acid sequences, A step of preparing a nucleic acid sequence comprising a candidate nucleic acid sequence containing a non-translated region sequence and a target protein nucleic acid sequence encoding the target protein, and a barcode nucleic acid sequence encoding a peptide barcode directly or indirectly linked to the target protein. A step of expressing a protein from the nucleic acid sequence, A step of separating the peptide barcode from the protein, A step of analyzing the separated peptide barcode, Based on the results of the above analysis, the relationship between the expression of the target protein and the candidate nucleic acid sequence is obtained, and the optimal candidate nucleic acid sequence is selected. A method comprising, wherein the peptide barcode comprises multiple types of peptide barcodes contained in a peptide barcode library, and all of the peptide barcodes contained in the peptide barcode library contain a fixed sequence which is a common amino acid sequence.

2. The method according to claim 1, wherein the peptide barcode includes a random sequence consisting of a random amino acid sequence and the fixed sequence.

3. The method according to claim 1, wherein the fixed sequence is a sequence that has low dispersion in the ionization efficiency of a group of peptides including the fixed sequence, in a converter that takes an amino acid sequence or an amino acid sequence and charge as input and outputs the ionization efficiency of the amino acid sequence in a mass spectrometer.

4. The method according to claim 3, wherein the fixed sequence includes R at the N-terminus and RP at the C-terminus, or RPR at the C-terminus.

5. The method according to claim 1, wherein the peptide barcode includes a random sequence consisting of random amino acid sequences, an insertion sequence consisting of amino acid sequences common to some of the peptide barcodes in the peptide barcode library, and the fixed sequence.

6. The method according to claim 5, wherein the inserted sequence is a sequence that has low dispersion in the ionization efficiency of a group of peptides including the inserted sequence, in a converter that takes an amino acid sequence or an amino acid sequence and a charge as input and outputs the ionization efficiency of the amino acid sequence in a mass spectrometer.

7. The method according to claim 6, wherein the insertion sequence is selected from the group consisting of L, Q, E, EQ, QL, and WP.

8. The aforementioned insertion sequence is (i) L, Q, and E; (ii) Permutations of length 2 or 3 that include one or more of Q, E, D, P, G, and A; and (iii) Permutations of length 2 or more and 4 or less that include one or more of Q, E, D, P, G and A and one or more of L, W, F, V and Y The method according to claim 6, wherein the selected member is from the group consisting of the following.

9. The method according to claim 1, wherein the peptide barcode does not contain K.

10. The method according to claim 2 or 5, wherein the random sequence contains the amino acids used for the random sequence with equal probability.

11. The method according to claim 2 or 5, wherein the random sequence contains amino acids D, F, W, and / or P, each with a probability of 3% or less.

12. An insertion site for inserting a candidate nucleic acid sequence containing the sequence encoding the target protein and the sequence of the untranslated region, A sequence encoding a peptide barcode containing two or more amino acids and A kit for nucleic acid sequence optimization comprising multiple expression cassettes, The plurality of expression cassettes each have sequences that encode different peptide barcodes, When the expression cassette is expressed, the target protein inserted into the insertion site and the peptide barcode are linked and expressed together. The kit comprises several types of peptide barcodes included in a peptide barcode library, and all of the peptide barcodes included in the peptide barcode library contain a fixed sequence which is a common amino acid sequence.

13. A kit according to claim 12 for carrying out the method described in claim 1.

14. The kit according to claim 12, wherein the peptide barcode further comprises an insertion sequence consisting of an amino acid sequence common to some of the peptide barcodes in the peptide barcode library.