Method and apparatus for codon sequence design based on deep learning model
Through a deep learning model-based method, an optimized codon sequence that takes into account the preferential context of species codon use is generated, which solves the problem of insufficient optimization of gene expression in the prior art and achieves a significant improvement in the expression level of gene proteins.
Patent Information
- Application Number
- CN202310125151.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-02-03
AI Technical Summary
The prior art fails to fully consider the contextual matching of codon use preferences of species when designing codon sequences for exogenous genes, and the verification method mainly relies on traditional biological experiments and lacks in-depth evaluation of bioinformatics.
Using a deep learning model-based method, by obtaining the full quantitative spectrum of proteomes of specified species, clustering analysis, constructing a collection of highly expressed genes and converting them into mathematical vectors, the Transformer deep learning model was input to generate optimized codon sequences and predicting protein expression abundance.
The generated optimized codon sequence takes into account the codon usage preferences of the species, significantly improves the protein expression level of exogenous genes in the host species, and can meet the requirements.
Smart Images

Figure CN116153402B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene design in synthetic biology, and particularly to a method and device for codon sequence design based on a deep learning model. Background Art
[0002] In the field of synthetic biology research, the design and optimization of gene expression vectors are indispensable, and the design of open reading frames is an important step in the design and optimization of gene expression vectors. A codon is a sequence combination formed by three adjacent nucleotides in the open reading frame of an mRNA (messenger RNA) sequence, and there are a total of 4 3 (64) forms. The open reading frame composed of these 64 codons ultimately determines the sequence of a protein. However, a protein sequence is only composed of 20 common amino acids, which means that 64 codons are redundant with respect to 20 amino acids. Among them, 18 amino acids except methionine and tryptophan are encoded by 2 to 6 codons, and this phenomenon is called codon degeneracy. Therefore, the codon sequence corresponding to an amino acid sequence is not unique. For example, for a protein with a length of 100 amino acids, the corresponding codon sequence can exist in approximately 5×10 47 combinatorial forms.
[0003] In 1987, Sharp and Li discovered the phenomenon of codon usage bias in the Escherichia coli genome and believed that the choice of codons by species is not completely random. This phenomenon was later further discovered in various model organisms such as Saccharomyces cerevisiae and Chlamydomonas reinhardtii. Moreover, the codon usage biases of different species are also different. Sharp and Li further proposed that before transferring an exogenous gene into a host species for expression, the codon sequence of this gene needs to be optimized according to the codon usage preference of the host species. Subsequently, they proposed the earliest codon optimization method, that is, using the optimal codons of the host species to modify the codon sequence of the exogenous gene. This idea is called CAI (Codon Adaptation Index) Maximization, and the genes optimized and designed based on this idea have been proven to achieve a significant improvement in expression level in the host species.
[0004] However, the single and repeated use of the same codon will cause the corresponding tRNA pool of the host species to be quickly depleted, so it is not conducive to the persistence of gene expression. Other scholars have also found that relative to the optimal codons, rare codons also play an active and important role in the process of gene translation and expression, and further proposed the idea of using optimal codons and rare codons together, which is called Codon Harmonization. Many methods developed based on this idea have been proven to be relatively better than CAI maximization.
[0005] Currently, many methods developed based on either CAI maximization or Codon Harmonization often only refer to the known codon usage frequencies of a certain species to specifically design the codon sequences of foreign genes, without deeply understanding the context collocations of codon usage preferences. At the same time, existing methods generally only use traditional biological experiments for verification, and rarely conduct prior evaluation and analysis at the bioinformatics level. Summary of the Invention
[0006] (I) Technical Problems to be Solved
[0007] In view of the problems existing in the above technology, the present invention solves them to at least a certain extent. For this reason, the first object of the present invention is to propose a method for designing codon sequences based on a deep learning model, which can design the codon sequences of target foreign genes by considering the context collocations of the codon usage preferences of a species, and can predict the protein expression level of the target foreign gene at the bioinformatics level.
[0008] The second object of the present invention is to propose a device for designing codon sequences based on a deep learning model, which can implement the method for designing codon sequences based on the deep learning model.
[0009] (II) Technical Solutions
[0010] To achieve the above object, the main technical solutions adopted by the present invention include:
[0011] In the first aspect, the present invention provides a method for designing codon sequences based on a deep learning model, including the following steps:
[0012] S1. Obtain the full quantitative spectrum of the proteome of a specified species;
[0013] S2. According to the protein expression abundance of each gene in the full quantitative spectrum of the proteome, all genes in the full quantitative spectrum of the proteome are clustered and analyzed to obtain K gene clusters; according to the protein expression abundance of the genes, the K gene clusters are arranged in descending order, and starting from the gene cluster ranked first, one or more gene clusters are selected to construct a set of highly expressed genes;
[0014] S3, converting the biological sequence text in the highly expressed gene set into a mathematical vector, and converting the biological sequence text in the full quantitative spectrum of the proteome into a mathematical vector;
[0015] S4. According to the highly expressed gene set converted into a mathematical vector, the amino acid sequence is input into the Transformer deep learning model, the codon sequence corresponding to the amino acid sequence is output, and the codon sequence generator is obtained by training; according to the full quantitative spectrum of the proteome converted into a mathematical vector, the codon sequence is input into the Transformer deep learning model, the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence belongs is output, and the protein expression abundance predictor is obtained by training;
[0016] S5. Obtain an exogenous target amino acid sequence and generate multiple random number seeds, and convert the target amino acid sequence into a mathematical vector; splice different random number seeds at the starting end of the target amino acid sequence converted into a mathematical vector to obtain multiple target amino acid sequences starting with different random numbers; input the multiple target amino acid sequences into a codon sequence generator, and output multiple optimized codon sequences; input the multiple optimized codon sequences into a protein expression abundance predictor, and output the protein expression abundance ranking corresponding to the gene cluster to which each optimized codon sequence belongs.
[0017] As an improvement of the method of the present invention, S1 comprises:
[0018] S11, performing a quantitative proteome search process, performing quantitative analysis on the proteome sequencing results according to the protein sequence annotation database of the specified species, and obtaining protein quantitative results of more than one sample;
[0019] S12. Based on the protein quantification results of more than one sample, calculate the geometric mean of the protein expression abundance of each gene in all samples; remove genes whose geometric mean of protein expression abundance is lower than the preset threshold to obtain the full quantitative spectrum of the proteome.
[0020] As an improvement of the method of the present invention, the preset threshold is 1; the geometric mean of the protein expression abundance of the gene in all samples is:
[0021]
[0022] In the formula, PSM iPSMi is the protein expression abundance of the gene in sample i; n is the total number of protein quantification samples; GEO(PSM) is the geometric mean of the protein expression abundances of the gene in all samples.
[0023] As an improvement of the method of the present invention, the K-Means clustering method is used to perform clustering analysis on all genes in the whole protein quantification spectrum;
[0024] According to the protein expression abundances of the genes, the K gene clusters are arranged in descending order, including: calculating the arithmetic mean of the protein expression abundances of each gene cluster according to the protein expression abundances of the genes;
[0025]
[0026] In the formula, PSM j is the protein expression abundance of gene j; m is the total number of protein quantification samples of all genes in a gene cluster; Mean is the arithmetic mean of the protein expression abundances of all genes in a gene cluster.
[0027] According to the arithmetic mean of the protein expression abundances of each gene cluster, the K gene clusters are arranged in descending order.
[0028] As an improvement of the method of the present invention, in S3, according to the corresponding rule between the pre-designed biological sequence unit and the number, the biological sequence text in the high-expression gene set is converted into a mathematical vector, and the biological sequence text in the whole proteome quantification spectrum is converted into a mathematical vector; the biological sequence unit includes codons and amino acids.
[0029] As an improvement of the method of the present invention, the Transformer deep learning model is the T5 model.
[0030] As an improvement of the method of the present invention, in S4, when the Transformer deep learning model constructs the codon sequence generator, the loss function used in the training process is:
[0031]
[0032] In the formula, y is the predicted codon sequence, is the codon sequence in the high-expression gene set;
[0033] When the Transformer deep learning model constructs the protein expression abundance predictor, the loss function used in the training process is:
[0034]
[0035] In the formula, x is the protein expression abundance ranking corresponding to the gene cluster to which the predicted codon sequence belongs, The protein expression abundance ranking corresponding to the gene cluster to which the codon sequence for training belongs.
[0036] As an improvement of the method of the present invention, the method for designing a codon sequence based on a deep learning model further includes: Step S6. According to the protein expression abundance ranking corresponding to the gene cluster to which the optimized codon sequence belongs, arrange all the optimized codon sequences in ascending or descending order, and select the optimal codon sequence.
[0037] In a second aspect, the present invention provides a device for designing a codon sequence based on a deep learning model, including:
[0038] The first acquisition module is used to acquire the full quantitative spectrum of the proteome of a specified species;
[0039] The clustering analysis module is used to perform clustering analysis on all genes in the full protein quantitative spectrum according to the protein expression abundance of each gene in the full quantitative spectrum of the proteome of the specified species, and obtain K gene clusters;
[0040] The high-expression gene set construction module is used to arrange the K gene clusters in descending order according to the protein expression abundance of the genes, and starting from the gene cluster ranked first, select one or more than two gene clusters to construct a high-expression gene set;
[0041] The transformation module is used to transform the biological sequence text in the high-expression gene set into a mathematical vector, transform the biological sequence text in the full quantitative spectrum of the proteome into a mathematical vector, and transform the target amino acid sequence into a mathematical vector;
[0042] The codon sequence generator module is used to input the amino acid sequence into the Transformer deep learning model according to the high-expression gene set transformed into a mathematical vector, output the codon sequence corresponding to the amino acid sequence, and train to obtain a codon sequence generator; and is used to input multiple random target amino acid sequences into the codon sequence generator and output multiple optimized codon sequences;
[0043] The protein expression abundance predictor module is used to input the codon sequence into the Transformer deep learning model according to the full quantitative spectrum of the proteome transformed into a mathematical vector, output the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence belongs, and train to obtain a protein expression abundance predictor; and is used to input multiple optimized codon sequences into the protein expression abundance predictor and output the protein expression abundance ranking corresponding to the gene cluster to which each optimized codon sequence belongs;
[0044] The second acquisition module is used to acquire a random number seed and an exogenous target amino acid sequence;
[0045] A random sequence generation module, configured to splice different random number seeds at the start of a target amino acid sequence converted into a mathematical vector to obtain multiple random target amino acid sequences.
[0046] Optionally, the first acquisition module is further configured to perform a quantitative proteomics database search process, perform quantitative analysis on the proteomics sequencing results according to the protein sequence annotation database of a specified species, and obtain protein quantitative results of one or more samples; according to the protein quantitative results of one or more samples, calculate the geometric mean of the protein expression abundances of each gene in all samples, and eliminate genes with a geometric mean of protein expression abundances lower than a preset threshold to obtain a full quantitative spectrum of the proteome.
[0047] (III) Beneficial effects
[0048] The beneficial effects of the present invention are as follows:
[0049] The method for codon sequence design based on a deep learning model provided by the present invention uses a codon sequence generator to generate an optimized codon sequence with context matching considering the codon usage preference of a species according to an exogenous target amino acid sequence (i.e., a target protein). By splicing different random number seeds at the start of the target amino acid sequence to form multiple target amino acid sequences as the input of the codon sequence generator, multiple optimized codon sequences can be generated without modifying the context matching of the exogenous target amino acid sequence. Then, the multiple optimized codon sequences are input into a protein expression abundance predictor to obtain the protein expression abundance of each optimized codon sequence. Thus, according to the protein expression abundance of each optimized codon sequence, the optimal codon sequence is selected as the exogenous target gene, and the protein expression level of the obtained exogenous target gene is significantly improved and can meet the requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The present invention is described with the aid of the following drawings:
[0051] Figure 1 is a schematic structural diagram of a device for codon sequence design based on a deep learning model according to an embodiment of the present invention;
[0052] Figure 2 is a schematic diagram of an expression vector according to Embodiment 1 of the present invention.
[0053]
Description of the reference numerals
[0054] 1: First acquisition module; 2: Clustering analysis module; 3: High-expression gene set construction module; 4: Conversion module; 5: Codon sequence generator module; 6: Protein expression abundance predictor module; 7: Second acquisition module; 8: Random sequence generation module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] To better explain the present invention for easy understanding, the present invention will be described in detail below in conjunction with the accompanying drawings and through specific embodiments.
[0056] The present invention provides a method for codon sequence design based on a deep learning model, comprising the following steps:
[0057] Step S1, obtaining the proteome-wide quantitative spectrum of a specified species.
[0058] Preferably, step S1 includes:
[0059] S11, performing a quantitative proteome database search process, quantitatively analyzing the proteome sequencing results according to the protein sequence annotation database of the specified species, and obtaining the protein quantitative results of more than one sample of the specified species.
[0060] It should be noted that the proteome refers to all the proteins expressed by a species, a tissue sample, or a cell. The quantitative proteome database search process refers to the process of automatically comparing the proteome sequencing results through a computer and an analysis program with reference to the database of theoretical mass spectrometry spectra. The proteome sequencing results refer to the data set of the mass spectrometry spectra of the peptide segments obtained by protein sequencing of the protein degradation peptide segments in a biological sample based on mass spectrometry technology. The protein sequence annotation database is the data set composed of the theoretical mass spectrometry spectra of proteins referred to in the quantitative proteome database search process.
[0061] Specifically, as an example, the MaxQuant software is used to quantitatively analyze the proteome sequencing results of multiple samples of the specified species. This step requires providing a reference configuration file for the operation of the MaxQuant software, which is generated by parsing the proteome sequencing data packet of the sample and the protein sequence annotation database of the specified species. Extract the MS / MS counts item in the proteome quantitative results of MaxQuant and regard it as PSM (Peptide-spectrum matches) to represent the absolute protein expression abundance of a certain gene in one or more samples.
[0062] S12, calculating the geometric mean of the protein expression abundance of each gene in all samples according to the protein quantitative results of more than one sample of the specified species; removing the genes with the geometric mean of the protein expression abundance lower than the preset threshold to obtain the proteome-wide quantitative spectrum.
[0063] Further preferably, the geometric mean of the protein expression abundance of a gene in all samples is:
[0064]
[0065] In the formula, PSMi PSMi is the protein expression abundance of the gene in sample i; n is the total number of protein quantification samples; GEO(PSM) is the geometric mean of the protein expression abundances of the gene in all samples.
[0066] Furthermore, the preset threshold is 1. In this way, genes with protein expression levels due to systematic errors can be excluded.
[0067] The full quantitative proteomic spectrum contains codon sequences and their corresponding amino acid sequences.
[0068] Step S2: Perform clustering analysis on all genes in the full quantitative proteomic spectrum according to the protein expression abundance of each gene in the full quantitative proteomic spectrum to obtain K gene clusters; according to the protein expression abundance of the genes, sort the K gene clusters in descending order, and starting from the gene cluster ranked first, select one or more than two gene clusters to construct a high-expression gene set.
[0069] Preferably, the K-Means clustering method is used to perform clustering analysis on all genes in the full quantitative proteomic spectrum.
[0070] Preferably, sorting the K gene clusters in descending order according to the protein expression abundance of the genes includes: calculating the arithmetic mean of the protein expression abundances of each gene cluster according to the protein expression abundance of the genes;
[0071]
[0072] In the formula, PSM j is the protein expression abundance of gene j; m is the total number of protein quantification samples of all genes in a gene cluster; Mean is the arithmetic mean of the protein expression abundances of all genes in a gene cluster.
[0073] Sort the K gene clusters in descending order according to the arithmetic mean of the protein expression abundances of each gene cluster.
[0074] Step S3: Convert the biological sequence text in the high-expression gene set into a mathematical vector, and convert the biological sequence text in the full quantitative proteomic spectrum into a mathematical vector.
[0075] Specifically, the biological sequence text includes amino acid sequence text and codon sequence text.
[0076] Since biological sequences represented by characters cannot be directly input into a mathematical model for calculation, they need to be first converted into mathematical vector representations. Preferably, according to the pre-designed correspondence rules between biological sequence units and numbers, the biological sequence texts in the highly expressed gene set are converted into mathematical vectors, and the biological sequence texts in the proteome-wide quantitative spectrum are converted into mathematical vectors. In the present invention, two types of biological sequences need to be considered, namely amino acid sequences and codon sequences. Therefore, the biological sequence units include 20 amino acids and 64 codons.
[0077] Specifically, as an example, the correspondence rules between the biological sequence units designed in the present invention and numbers are as follows:
[0078]
[0079]
[0080]
[0081] Furthermore, the pre-designed correspondence rules between biological sequence units and numbers also include: adding the symbol "TER" at the termination end of the amino acid sequence converted into a mathematical vector to represent the termination site of the amino acid sequence.
[0082] Step S4: According to the highly expressed gene set converted into a mathematical vector, input the amino acid sequence into the Transformer deep learning model, output the codon sequence corresponding to the amino acid sequence, and train to obtain a codon sequence generator; according to the proteome-wide quantitative spectrum converted into a mathematical vector, input the codon sequence into the Transformer deep learning model, output the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence belongs, and train to obtain a protein expression abundance predictor.
[0083] In order to guide the codon sequence designed by the codon sequence generator to enable a higher expression level of the target foreign gene, the codon usage frequency pattern of the sequence optimized and designed by the codon sequence generator should be close to that of the highly expressed genes of the target host species. Therefore, the selected highly expressed gene set is used as the training set for the codon sequence generator.
[0084] In order to guide the protein expression abundance predictor to correctly predict the protein expression levels of different codon sequences, the proteome-wide quantitative spectrum and the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence belongs are selected as the training set for the protein expression abundance predictor. In this way, the trained protein expression abundance predictor can predict the protein expression level of the target foreign gene at the bioinformatics (i.e., computational) level, facilitating users to select the optimal codon sequence.
[0085] Preferably, when constructing a codon sequence generator using a Transformer deep learning model, the loss function used during training is cross-entropy, specifically:
[0086]
[0087] In the formula, y is the predicted codon sequence, and it is the codon sequence in the high-expression gene set.
[0088] When constructing a protein expression abundance predictor using a Transformer deep learning model, the loss function used during training is cross-entropy, specifically:
[0089]
[0090] In the formula, x is the protein expression abundance ranking corresponding to the gene cluster to which the predicted codon sequence belongs, and it is the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence used for training belongs.
[0091] More preferably, in this embodiment, the Transformer deep learning model is a T5 model (Text-To-Text Transfer Transformer text transfer model).
[0092] Step S5: Obtain an exogenous target amino acid sequence and generate multiple random number seeds, and convert the target amino acid sequence into a mathematical vector; splice different random number seeds as guiding labels at the starting end of the target amino acid sequence converted into a mathematical vector to obtain multiple target amino acid sequences with different random number starts; input the multiple target amino acid sequences into the codon sequence generator to output multiple optimized codon sequences; input the multiple optimized codon sequences into the protein expression abundance predictor to output the protein expression abundance ranking corresponding to the gene cluster to which each optimized codon sequence belongs.
[0093] Step S6: Arrange all the optimized codon sequences in ascending or descending order according to the protein expression abundance ranking corresponding to the gene cluster to which the optimized codon sequence belongs, and select the optimal codon sequence.
[0094] In summary, the method for designing a codon sequence based on a deep learning model provided by the present invention uses a codon sequence generator to generate an optimized codon sequence with context collocations considering the codon usage preferences of species according to an exogenous target amino acid sequence (i.e., the target protein). By splicing different random number seeds at the start end of the target amino acid sequence to form multiple target amino acid sequences as the input of the codon sequence generator, multiple optimized codon sequences can be generated without modifying the context collocations of the exogenous target amino acid sequence. Then, the multiple optimized codon sequences are input into a protein expression abundance predictor to obtain the protein expression abundance of each optimized codon sequence. Thus, according to the protein expression abundance of each optimized codon sequence, the optimal codon sequence is selected as the exogenous target gene. The protein expression level of the obtained exogenous target gene is significantly improved and can meet the requirements. If the random number seeds are not used to splice the target amino acid sequence, then only one target amino acid sequence is used as the input of the codon generator, and the codon generator can only generate one optimized codon sequence, and the protein expression level of this optimized codon sequence often cannot meet the requirements.
[0095] Figure 1 It is a schematic structural diagram of the device for designing a codon sequence based on a deep learning model provided by the present invention.
[0096] As Figure 1 shown, the device for designing a codon sequence based on a deep learning model includes a first acquisition module 1, a clustering analysis module 2, a high-expression gene set construction module 3, a transformation module 4, a codon sequence generator module 5, a protein expression abundance predictor module 6, a second acquisition module 7, and a random sequence generation module 8.
[0097] Among them, the first acquisition module 1 is used to obtain the full quantitative spectrum of the proteome of the specified species. The cluster analysis module 2 is used to perform cluster analysis on all genes in the full quantitative spectrum of the proteome of the specified species according to the protein expression abundance of each gene in the full quantitative spectrum of the proteome, and obtain K gene clusters. The high expression gene set construction module 3 is used to sort the K gene clusters in descending order according to the protein expression abundance of the genes, and start from the gene cluster ranked first, and select one or more gene clusters to construct a high expression gene set. The conversion module 4 is used to convert the biological sequence text in the high expression gene set into a mathematical vector, convert the biological sequence text in the full quantitative spectrum of the proteome into a mathematical vector, and convert the target amino acid sequence into a mathematical vector. The codon sequence generator module 5 is used to input the amino acid sequence into the Transformer deep learning model according to the high expression gene set converted into a mathematical vector, output the codon sequence corresponding to the amino acid sequence, and train to obtain a codon sequence generator; and to input multiple random target amino acid sequences into the codon sequence generator, and output multiple optimized codon sequences. The protein expression abundance predictor module 6 is used to input the codon sequence into the Transformer deep learning model according to the full quantitative spectrum of the proteome converted into a mathematical vector, output the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence belongs, and train to obtain the protein expression abundance predictor; and to input multiple optimized codon sequences into the protein expression abundance predictor, and output the protein expression abundance ranking corresponding to the gene cluster to which each optimized codon sequence belongs. The second acquisition module 7 is used to obtain a random number seed and an exogenous target amino acid sequence. The random sequence generation module 8 is used to splice different random number seeds at the starting end of the target amino acid sequence converted into a mathematical vector to obtain multiple random target amino acid sequences.
[0098] As an example, the first acquisition module 1 is also used to perform a quantitative proteome search process, quantitatively analyze the proteome sequencing results according to the protein sequence annotation database of the specified species, and obtain the protein quantitative results of more than one sample; based on the protein quantitative results of more than one sample, calculate the geometric mean of the protein expression abundance of each gene in all samples, eliminate genes whose protein expression abundance geometric mean is lower than a preset threshold, and obtain a full quantitative spectrum of the proteome.
[0099] It should be noted that the specific functions of each module in the device for codon sequence design based on deep learning model provided by the present invention, and the processing flow of codon sequence design based on deep learning model, can be referred to the detailed description of the codon sequence design method based on deep learning model provided above, and will not be repeated here.
[0100] In summary, in the device for codon sequence design based on a deep learning model provided by the present invention, the codon sequence generator module 5 can generate an optimized codon sequence with context collocations considering the codon usage preferences of species according to an exogenous target amino acid sequence (i.e., the target protein); the random sequence generator module 8 can form multiple target amino acid sequences as the input of the codon sequence generator by splicing different random number seeds at the starting end of the target amino acid sequence, so as to generate multiple optimized codon sequences without modifying the context collocations of the exogenous target amino acid sequence. The protein expression abundance predictor module 6 is used to input multiple optimized codon sequences into the protein expression abundance predictor to obtain the protein expression abundance of each optimized codon sequence, and thus select the optimal codon sequence as the exogenous target gene according to the protein expression abundance of each optimized codon sequence. The protein expression level of the obtained exogenous target gene is significantly improved and can meet the requirements.
[0101] Example 1
[0102] Step S1: Obtain the full quantitative spectrum of the proteome of the Yarrowia lipolytica W29 strain.
[0103] Step S2: Perform cluster analysis on all genes in the full quantitative spectrum of the proteome according to the protein expression abundance of each gene in the full quantitative spectrum of the proteome to obtain 5 gene clusters; arrange the 5 gene clusters in descending order according to the protein expression abundance of the genes, and start from the gene cluster ranked first, and select two gene clusters to construct a high-expression gene set. There are 617 genes in the high-expression gene set.
[0104] Step S3: Convert the biological sequence text in the high-expression gene set into a mathematical vector, and convert the biological sequence text in the full quantitative spectrum of the proteome into a mathematical vector.
[0105] Step S4: According to the high-expression gene set converted into a mathematical vector, input the amino acid sequence into the Transformer deep learning model to output the codon sequence corresponding to the amino acid sequence, and train to obtain a codon sequence generator; according to the full quantitative spectrum of the proteome converted into a mathematical vector, input the codon sequence into the Transformer deep learning model to output the ranking of the protein expression abundance corresponding to the gene cluster to which the codon sequence belongs, and train to obtain a protein expression abundance predictor.
[0106] Step S5: Obtain the exogenous target amino acid sequence and generate 32 random number seeds from 0 to 31, and convert the target amino acid sequence into a mathematical vector; splice different random number seeds as guiding tags at the starting end of the target amino acid sequence converted into a mathematical vector to obtain multiple target amino acid sequences with different random number starts; input the multiple target amino acid sequences into a codon sequence generator to output multiple optimized codon sequences; input the multiple optimized codon sequences into a protein expression abundance predictor to output the protein expression abundance ranking corresponding to the gene cluster to which each optimized codon sequence belongs.
[0107] Specifically, the obtained exogenous target amino acid sequences are the CrtYB and CrtI genes from Phaffia rhodozyma, and these two genes can convert the upstream substrate into β-carotene.
[0108]
[0109]
[0110] Step S6: According to the protein expression abundance ranking corresponding to the gene cluster to which the optimized codon sequence belongs, arrange all the optimized codon sequences in descending order, and select the optimized codon sequence ranked first as the optimal codon sequence.
[0111] Specifically, the obtained optimal codon sequences corresponding to CrtYB and CrtI are:
[0112]
[0113]
[0114]
[0115]
[0116] Construct the optimal codon sequences of the CrtYB and CrtI genes into an expression vector and express them using the GPD1 and FBA1 promoters respectively. The schematic diagram of this expression vector is as Figure 2 shown.
[0117] Among them, the promoter of the CrtYB gene is FBA1, and the terminator is LIP2; the promoter of the CrtI gene is GPD1, and the terminator is PEX20. The resistance screening gene of this expression vector is hph. The transformants successfully transfected with this expression vector can grow monoclonal strains in the medium supplemented with hygromycin (Hygromycin B) because they possess the hph resistance gene. IntE_1 is the site-specific insertion site of this vector, and IntE_1-right arm and IntE_1-left arm are the homologous arms of the site-specific insertion site, with a length of about 500bp.
[0118] This expression vector was transfected into the Yarrowia lipolytica W29 strain and spread on the YPD solid medium containing hygromycin at a final concentration of 200 μg / mL. After culturing for 2 days at 30°C, there were monoclonal strains with an orange-yellow phenotype, indicating that these monoclonal strains were successfully transfected with the codon-optimized CrtYB and CrtI genes and accumulated β-carotene. The monoclonal strains with an orange-yellow phenotype were picked out and preserved.
[0119] The results of this example show that the optimized codon sequences of the CrtYB and CrtI genes optimized by the model of the present invention can be successfully expressed in the host species Yarrowia lipolytica.
[0120] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0121] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions.
[0122] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a claim listing several means, several of these means can be embodied by the same hardware. The use of the terms first, second, third, etc. is for convenience only and does not denote any order. These terms can be construed as part of the name of the element.
[0123] In addition, it should be noted that in the description of this specification, the descriptions of the terms "an embodiment", "some embodiments", "embodiments", "examples", "specific examples" or "some examples", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0124] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications after learning the basic creative concept. Therefore, the claims should be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0125] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention should also include these modifications and variations.
Claims
1. A method for codon sequence design based on deep learning model, It is characterized in that The following steps are involved: S1. Obtain the full quantitative proteome profile of a specified species; S2. According to the protein expression abundance of each gene in the full quantitative spectrum of the proteome, all genes in the full quantitative spectrum of the proteome are clustered and analyzed to obtain K gene clusters; according to the protein expression abundance of the genes, the K gene clusters are arranged in descending order, and starting from the gene cluster ranked first, one or more gene clusters are selected to construct a set of highly expressed genes; S3, converting the biological sequence text in the highly expressed gene set into a mathematical vector, and converting the biological sequence text in the full quantitative spectrum of the proteome into a mathematical vector; S4. According to the highly expressed gene set converted into mathematical vectors, the amino acid sequence is input into the Transformer deep learning model, the codon sequence corresponding to the amino acid sequence is output, and the codon sequence generator is obtained by training; According to the full quantitative spectrum of the proteome converted into a mathematical vector, the codon sequence is input into the Transformer deep learning model, the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence belongs is output, and the protein expression abundance predictor is obtained by training; S5. Obtain an exogenous target amino acid sequence and generate multiple random number seeds, and convert the target amino acid sequence into a mathematical vector; splice different random number seeds at the starting end of the target amino acid sequence converted into a mathematical vector to obtain multiple target amino acid sequences starting with different random numbers; input the multiple target amino acid sequences into a codon sequence generator, and output multiple optimized codon sequences; input the multiple optimized codon sequences into a protein expression abundance predictor, and output the protein expression abundance ranking corresponding to the gene cluster to which each optimized codon sequence belongs.
2. The method for codon sequence design based on a deep learning model according to claim 1, It is characterized in that S1 includes: S11, performing a quantitative proteome search process, performing quantitative analysis on the proteome sequencing results according to the protein sequence annotation database of the specified species, and obtaining protein quantitative results of more than one sample; S12. Based on the protein quantification results of more than one sample, calculate the geometric mean of the protein expression abundance of each gene in all samples; remove genes whose geometric mean of protein expression abundance is lower than the preset threshold to obtain the full quantitative spectrum of the proteome.
3. The method for codon sequence design based on a deep learning model according to claim 2, It is characterized in that The preset threshold is 1; The geometric mean of the protein expression abundance of a gene in all samples is: In the formula, PSM i is the protein expression abundance of the gene in sample i; n is the total number of protein quantification samples; GEO(PSM) is the geometric mean of the protein expression abundances of the gene in all samples.
4. The method for codon sequence design based on a deep learning model according to claim 1, It is characterized in that The K-Means clustering method was used to perform cluster analysis on all genes in the whole protein quantitative profile; Arranging the K gene clusters in descending order according to the protein expression abundance of the genes, including: calculating the arithmetic mean of the protein expression abundance of each gene cluster according to the protein expression abundance of the genes; In the formula, PSM j is the protein expression abundance of gene j; m is the total number of protein quantification samples of all genes in a gene cluster; Mean is the arithmetic mean of the protein expression abundances of all genes in a gene cluster; According to the arithmetic mean of the protein expression abundance of each gene cluster, the K gene clusters are sorted in descending order.
5. The method for codon sequence design based on a deep learning model according to claim 1, characterized in that, in S3, according to the pre-designed correspondence rules between biological sequence units and numbers, the biological sequence texts in the highly expressed gene set are converted into mathematical vectors, and the biological sequence texts in the proteome-wide quantitative spectrum are converted into mathematical vectors; the biological sequence units include codons and amino acids.
6. The method for codon sequence design based on a deep learning model according to claim 1, characterized in that, the Transformer deep learning model is the T5 model.
7. The method for codon sequence design based on a deep learning model according to claim 1, characterized in that, in S4, when the Transformer deep learning model constructs a codon sequence generator, the loss function used during the training process is: where y is the predicted codon sequence, which is the codon sequence in the high-expression gene set; when the Transformer deep learning model constructs a protein expression abundance predictor, the loss function used during the training process is: Where x is the ranking of the protein expression abundance corresponding to the gene cluster to which the predicted codon sequence belongs, and is the ranking of the protein expression abundance corresponding to the gene cluster to which the codon sequence used for training belongs.
8. The method for codon sequence design based on a deep learning model according to claim 1, characterized in that, it further includes: Step S6, according to the protein expression abundance ranking corresponding to the gene cluster to which the optimized codon sequence belongs, arranging all the optimized codon sequences in ascending or descending order, and selecting the optimal codon sequence.
9. An apparatus for codon sequence design based on a deep learning model, characterized in that, it includes: A first acquisition module (1) for acquiring the proteome-wide quantitative spectrum of a specified species; A clustering analysis module (2) for performing clustering analysis on all genes in the whole protein quantitative spectrum according to the protein expression abundance of each gene in the proteome-wide quantitative spectrum of the specified species to obtain K gene clusters; A highly expressed gene set construction module (3) for arranging the K gene clusters in descending order according to the protein expression abundance of the genes, and starting from the gene cluster ranked first, selecting one or more than two gene clusters to construct a highly expressed gene set; A conversion module (4) for converting the biological sequence texts in the highly expressed gene set into mathematical vectors, converting the biological sequence texts in the proteome-wide quantitative spectrum into mathematical vectors, and converting the target amino acid sequence into a mathematical vector; A codon sequence generator module (5) for inputting the amino acid sequence into the Transformer deep learning model according to the highly expressed gene set converted into a mathematical vector, outputting the codon sequence corresponding to the amino acid sequence, and training to obtain a codon sequence generator; and for inputting multiple random target amino acid sequences into the codon sequence generator and outputting multiple optimized codon sequences; A protein expression abundance predictor module (6) for inputting the codon sequence into the Transformer deep learning model according to the proteome-wide quantitative spectrum converted into a mathematical vector, outputting the protein expression abundance ranking corresponding to the gene cluster to which the codon sequence belongs, and training to obtain a protein expression abundance predictor; and for inputting multiple optimized codon sequences into the protein expression abundance predictor and outputting the protein expression abundance ranking corresponding to the gene cluster to which each optimized codon sequence belongs; A second acquisition module (7) is used to obtain a random number seed and an exogenous target amino acid sequence; The random sequence generation module (8) is used to splice different random number seeds at the starting end of the target amino acid sequence converted into a mathematical vector to obtain multiple random target amino acid sequences.
10. The device for codon sequence design based on deep learning model according to claim 9, It is characterized in that The first acquisition module (1) is also used to perform a quantitative proteome search process, quantitatively analyze the proteome sequencing results according to the protein sequence annotation database of the specified species, and obtain the protein quantitative results of more than one sample; based on the protein quantitative results of more than one sample, calculate the geometric mean of the protein expression abundance of each gene in all samples, eliminate genes whose protein expression abundance geometric mean is lower than a preset threshold, and obtain a full quantitative spectrum of the proteome.
Citation Information
Patent Citations
Codon optimization
US20190325989A1
RNA replicon for versatile and efficient gene expression
WO2017162460A1