Sequence optimization method for balancing function and yield

By collaboratively optimizing nucleic acid and amino acid sequences through a deep learning model based on Ribo-seq data, the problem of balancing function and yield in protein design is solved, adapting to various production environments and improving optimization efficiency and adaptability.

CN120708709APending Publication Date: 2025-09-26ZHONGSHAN OPHTHALMIC CENT SUN YAT SEN UNIV

Patent Information

Application Number
CN202510635436.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing protein design methods make it difficult to improve production while optimizing function, fail to adapt to various production environments, ignore the linkage between nucleic acid sequences and amino acid sequences, lack multi-objective optimization capabilities, and cannot achieve balance in different environments.

Method used

A multi-objective optimization method based on machine learning is adopted, combined with Ribo-seq data to train a deep learning model, to collaboratively optimize nucleic acid sequences and amino acid sequences, taking into account different cellular environments, and balancing function and yield through a multi-objective optimization algorithm to adapt to various production environments.

Benefits of technology

It achieves the ability to simultaneously improve protein function and yield under different production environments, improves optimization efficiency and adaptability, and is suitable for laboratory and industrial-scale production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708709A_ABST
    Figure CN120708709A_ABST
Patent Text Reader

Abstract

The invention discloses a sequence optimization method for balancing functions and yield. The sequence optimization method comprises the following steps: acquiring an initial amino acid sequence, a cell environment and a nucleic acid expression quantity; obtaining an initial nucleic acid sequence according to the initial amino acid sequence; obtaining a yield fraction through a preset yield module according to the initial nucleic acid sequence, the cell environment and the nucleic acid expression quantity; obtaining at least one function score through at least one preset function module according to the initial amino acid sequence; and simultaneously optimizing the initial amino acid sequence and the initial nucleic acid sequence through an optimization module according to the yield fraction and the at least one functional fraction to obtain an optimized amino acid sequence and an optimized nucleic acid sequence. According to the invention, the functional module and the yield module are integrated, and multi-objective optimization of protein sequences and corresponding nucleic acid sequences is realized through the optimization module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sequence optimization, and more particularly to a sequence optimization method for balancing function and yield. Background Art

[0002] Traditional methods for modifying protein functionality mainly rely on large-scale library construction and high-throughput screening technologies, which are usually time-consuming and resource-intensive. Protein research and development and production face many challenges, including insufficient protein functionality (such as low thermal stability, affinity, and specificity), low expression levels, poor expression stability, and long production cycles. Current protein design methods usually focus on optimizing protein functionality or yield, but it is difficult to achieve a balance between the two. As a result, protein yields are often low during functional optimization, while protein functionality may be significantly affected when pursuing high yields. In recent years, artificial intelligence technology has made significant progress in the field of sequence and molecular design. Among them, protein design has become one of the most challenging research directions in this field due to its high complexity and specificity. We use artificial intelligence technology to achieve coordinated optimization of nucleic acid sequences and amino acid sequences, multi-objective design to simultaneously improve function and yield, and adapt to the needs of various production environments.

[0003] In the prior art, there are methods for fine-tuning initial protein characterization models based on text-protein pair datasets to improve the functional prediction performance of unknown proteins and mutants and achieve efficient protein modification effects (CN119626323A); there are large protein language models based on antibody structure fine-tuning, which use model fine-tuning modules, antibody design modules, and 3D structure modeling modules to design new antibodies with improved affinity and specificity for specific antigens (CN118658515A); there are methods for using sequence annotation models to predict antibody amino acid sequences by mutation, and training sequence annotation models based on score data combined with proximal strategy optimization algorithms to obtain optimized antibodies (CN117894373A); there are methods for optimizing their coding sequences at the CAI level and the MFE level to obtain stable and highly expressed CDS sequences (CN115814074A); there are methods for using sequence-based analysis and rational strategies to modify and improve the structure and biophysical properties (including stability, solubility, and antigen-binding affinity) of single-chain antibodies (scFv) (CN101688200A).

[0004] The above-mentioned prior art still has the following problems to be solved:

[0005] (1) Existing methods do not consider the optimization of antibody production while optimizing function (catalytic efficiency, stability, affinity, and specificity, etc.);

[0006] (2) Most methods are based on empirical knowledge or indicators to optimize protein function or yield, which has limited optimization efficiency and is prone to falling into local optimality;

[0007] (3) The disclosed method is optimized in a single environment and cannot be adapted to different recombinant protein production platforms;

[0008] (4) Existing multi-objective sequence optimization algorithms do not consider the linkage between nucleic acid sequences and protein sequences;

[0009] (5) Existing methods do not provide a sequence optimization method that can be connected to any biological sequence (DNA / RNA / amino acid) function prediction model. Summary of the Invention

[0010] The present invention provides a sequence optimization method that balances function and output, solving the technical problems of low efficiency, single target and poor adaptability to unknown scenarios in the prior art optimization methods.

[0011] In order to solve the above technical problems, the technical solutions of the present invention are as follows:

[0012] The present invention provides a sequence optimization method for balancing function and yield, comprising the following steps:

[0013] Obtain the initial amino acid sequence, cell environment, and nucleic acid expression level;

[0014] Obtaining an initial nucleic acid sequence according to the initial amino acid sequence;

[0015] Calculating a yield score based on the initial nucleic acid sequence, the cell environment, and the nucleic acid expression level using a machine learning-based yield prediction module, wherein the yield score corresponds to the amount of protein produced by translation of the nucleic acid sequence in a specific environment;

[0016] Calculating at least one functional score corresponding to a specific biological function of the protein by at least one function prediction module based on machine learning;

[0017] Based on the yield score and at least one function score, the initial amino acid sequence and the initial nucleic acid sequence are collaboratively optimized using a multi-objective optimization algorithm in an optimization module to obtain an optimized amino acid sequence and an optimized nucleic acid sequence, wherein the multi-objective optimization algorithm aims to maximize the yield score and the function score.

[0018] Furthermore, the cell environment includes transcriptome data of different tissues or cell types, and the transcriptome data is used to construct a cell environment vector to improve the accuracy of the yield prediction module and the function prediction module.

[0019] Furthermore, based on the initial amino acid sequence, an initial nucleic acid sequence is obtained by codon optimization, wherein the codon optimization comprises:

[0020] For each amino acid in the amino acid sequence, obtain all possible codons;

[0021] According to the target cell environment and the nucleic acid expression level, the yield score of each codon is evaluated by the yield prediction module;

[0022] The codon with the highest yield score is selected to construct the initial nucleic acid sequence.

[0023] Furthermore, the training process of the preset yield module includes:

[0024] Collect transcriptome and translation efficiency data in the target cell environment;

[0025] Based on the transcriptome data and translation efficiency data, a training data set is constructed, wherein each sample includes a nucleic acid coding sequence, a cell environment vector, and a nucleic acid expression level;

[0026] Training a neural network model, including a convolutional neural network or a recurrent neural network, using the nucleic acid coding sequence, the cell environment vector, and the nucleic acid expression level as input to predict translation efficiency;

[0027] The model performance is evaluated through cross-validation to obtain the preset yield module.

[0028] Furthermore, the transcriptome data is RNA-seq data, and the translation efficiency data is Ribo-seq data;

[0029] The construction of the training data set includes:

[0030] Quantify and normalize RNA-seq and Ribo-seq data and filter low-expression genes;

[0031] Convert the nucleic acid coding sequence into a one-hot coding matrix;

[0032] RNA-seq expression quantification was sorted by fixed genes to construct a cell environment vector.

[0033] Furthermore, if there are two or more functional modules, each functional prediction module predicts different protein functions based on a machine learning algorithm to obtain different functional scores, wherein the different functions include antibody affinity, hydrophobicity, enzyme catalytic activity, substrate specificity, signal peptide, immunogenicity and subcellular localization.

[0034] Furthermore, the optimization module is constructed through a multi-objective optimization method of weighted addition, constrained optimization or reinforcement learning + multi-objective reward, and combines the yield score and at least one functional score to perform multi-objective collaborative optimization of amino acid sequences and nucleic acid sequences through dynamic weight adjustment. It receives optimization weight, number of iterations, yield score, at least one functional score and maximum number of mutations as input, and outputs an optimized amino acid sequence and an optimized initial nucleic acid sequence.

[0035] Furthermore, the optimized amino acid sequence and the optimized nucleic acid sequence are output, including:

[0036] The optimization module obtains an optimized initial amino acid sequence according to the optimization weight, the number of iterations, the yield score, the at least one functional score and the maximum number of mutations;

[0037] According to the optimized amino acid sequence, an optimized nucleic acid sequence is obtained through codon optimization.

[0038] Furthermore, when only the yield prediction module or one function prediction module is used, the multi-objective optimization algorithm is simplified to a single-objective optimization, taking the yield score or one function score as input and outputting an optimized amino acid sequence or an optimized nucleic acid sequence.

[0039] Furthermore, the optimization module is constructed by activation maximization, saturation mutation, adversarial generative network (GAN) / diffusion model generation or reinforcement learning method.

[0040] Furthermore, the optimization module uses a simulated annealing algorithm to collaboratively optimize the initial amino acid sequence and the initial nucleic acid sequence.

[0041] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0042] This paper proposes a deep learning model trained on environmental Ribo-seq data. It leverages computable key sequence properties to collaboratively optimize nucleic acid and amino acid sequences and supports multi-objective design to improve function and yield. This method has the following advantages in terms of innovation and application value in the field of biological sequence optimization:

[0043] 1. Deep learning model training based on Ribo-seq data

[0044] Traditional biological sequence optimization methods rely primarily on experimental data or simple mathematical models to assess the impact of sequence on function. This paper proposes using Ribo-seq data to train deep learning models. This enables the optimization model to better capture the complex biological information related to translation efficiency, thereby providing more accurate optimization results. Compared to traditional models based on genomic or transcriptomic data, this approach can take into account the actual conditions of protein translation, thereby improving the effectiveness and adaptability of sequence optimization.

[0045] 2. Collaborative optimization of nucleic acid and amino acid sequences

[0046] This invention enables simultaneous optimization of nucleic acid and amino acid sequences. Nucleic acid and amino acid sequences are both independent and closely linked. Nucleic acid properties such as codon usage and mRNA structure directly influence protein expression, while protein functional requirements in turn constrain nucleic acid design. Traditional methods of isolated optimization result in a "one-sided" optimization, while others are inferior. Only collaborative modeling can achieve global optimization.

[0047] 3. Multi-objective optimization design

[0048] The present invention supports multi-objective design, simultaneously improving functionality and yield. Existing optimization methods often focus on improving only one aspect, such as translation efficiency or protein function. The present invention utilizes multi-objective optimization to simultaneously optimize across multiple dimensions based on actual needs, meeting the demands of diverse application scenarios and enhancing the method's versatility and adaptability.

[0049] 4. Adapt to various production environments

[0050] Traditional optimization methods often struggle with poor performance in diverse production environments. However, the present invention is designed with full environmental adaptability in mind, enabling flexible optimization based on diverse production environments and demonstrating strong adaptability. This capability allows the present invention to be applied not only in laboratory research but also in industrial-scale production, demonstrating its potential for commercialization.

[0051] 5. Wide range of application scenarios

[0052] The technological innovation of this invention not only has significant application prospects in optimizing exogenous DNA sequences, mRNA vaccines, enzyme and antibody engineering, but also can provide effective optimization tools for multiple fields such as synthetic biology, gene editing, and metabolic engineering. Its wide range of application scenarios further demonstrates the innovativeness and efficiency of this technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A schematic flow chart of a sequence optimization method for balancing function and yield provided by an embodiment of the present invention;

[0054] Figure 2 A schematic diagram of a framework of a sequence optimization method for balancing function and yield provided by an embodiment of the present invention;

[0055] Figure 3 Schematic diagram of input and output of the production module provided in an embodiment of the present invention;

[0056] Figure 4 A schematic diagram of a feasible network model structure of a production module provided by an embodiment of the present invention;

[0057] Figure 5 The embodiment of the present invention provides Figure 4 Schematic diagram of the prediction accuracy of the yield module on the test set;

[0058] Figure 6 A schematic diagram of the input and output of the functional modules provided in an embodiment of the present invention;

[0059] Figure 7 A schematic diagram of the convolutional neural network structure of the affinity prediction model provided in an embodiment of the present invention;

[0060] Figure 8 A schematic diagram of the input and output of a single-objective optimization module provided by an embodiment of the present invention;

[0061] Figure 9 A schematic diagram of the input and output of a multi-objective optimization module provided by an embodiment of the present invention;

[0062] Figure 10 A schematic diagram of the input and output of a multi-objective optimization module using a weighted sum method provided in an embodiment of the present invention;

[0063] Figure 11 A schematic diagram of the workflow of an optimization module constructed based on an improved simulated annealing algorithm provided in an embodiment of the present invention;

[0064] Figure 12 The use provided by the embodiment of the present invention Figure 11 Schematic diagram of the optimization results of the antibody variable region sequence of the optimization module. DETAILED DESCRIPTION

[0065] The accompanying drawings are for illustrative purposes only and are not to be construed as limiting this patent;

[0066] In order to better illustrate this embodiment, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product size;

[0067] It is understandable to those skilled in the art that some well-known structures and descriptions thereof may be omitted in the drawings.

[0068] The technical solution of the present invention is further described below with reference to the accompanying drawings and embodiments.

[0069] Example 1

[0070] The embodiment of the present invention provides a sequence optimization method that balances function and yield, such as Figure 1 and Figure 2 As shown, the following steps are included:

[0071] Obtain the initial amino acid sequence, cell environment, and nucleic acid expression level;

[0072] Obtaining an initial nucleic acid sequence according to the initial amino acid sequence;

[0073] Calculating a yield score based on the initial nucleic acid sequence, the cell environment, and the nucleic acid expression level using a machine learning-based yield prediction module, wherein the yield score corresponds to the amount of protein produced by translation of the nucleic acid sequence in a specific environment;

[0074] Calculating at least one functional score corresponding to a specific biological function of the protein by at least one function prediction module based on machine learning;

[0075] Based on the yield score and at least one function score, the initial amino acid sequence and the initial nucleic acid sequence are collaboratively optimized using a multi-objective optimization algorithm in an optimization module to obtain an optimized amino acid sequence and an optimized nucleic acid sequence, wherein the multi-objective optimization algorithm aims to maximize the yield score and the function score.

[0076] The technical problem to be solved by the embodiments of the present invention is that traditional protein sequence optimization has low efficiency, a single target, poor adaptability to unknown scenarios, and no balance between function and yield. A protein sequence multi-objective optimization system and method based on deep learning are proposed. The system integrates functional modules and yield modules, and realizes multi-objective optimization of protein sequences through optimization modules.

[0077] In a further embodiment, the cellular environment includes transcriptome data of different tissues or cell types.

[0078] Example 2

[0079] Based on Example 1, the present invention specifically describes the yield module.

[0080] like Figure 3The following is a schematic diagram of the input and output of the yield module, which uses deep learning methods to predict the expression level of a given mRNA coding sequence, also known as the yield score. The yield module receives the nucleic acid sequence, cellular environment, and nucleic acid expression level as input and outputs the yield score, which is a normalized RPF (ribosome footprint) value. This value refers to the mRNA fragment protected when ribosomes bind to and translate the mRNA, reflecting the amount of active ribosome translation of the gene.

[0081] In a further embodiment, the training process of the preset yield module includes:

[0082] Collect transcriptome and translation efficiency data in the target cell environment;

[0083] Based on the transcriptome data and translation efficiency data, a training data set is constructed, wherein each sample includes a nucleic acid coding sequence, a cell environment vector, and a nucleic acid expression level;

[0084] Training a neural network model, including a convolutional neural network or a recurrent neural network, using the nucleic acid coding sequence, the cell environment vector, and the nucleic acid expression level as input to predict translation efficiency;

[0085] The model performance is evaluated through cross-validation to obtain the preset yield module.

[0086] In a further embodiment, the transcriptome data is RNA-seq data, and the translation efficiency data is Ribo-seq data;

[0087] The construction of the training data set includes:

[0088] Quantify and normalize RNA-seq and Ribo-seq data and filter low-expression genes;

[0089] Convert the nucleic acid coding sequence into a one-hot coding matrix;

[0090] RNA-seq expression quantification was sorted by fixed genes to construct a cell environment vector.

[0091] In a specific embodiment, the Ribo-seq ribosome expression profiles under different cell types and experimental treatment conditions are collected, and the RNA-seq gene expression profiles paired with Ribo-seq are collected; Ribo-seq and RNA-seq are used as training data for the method; different cell types and experimental treatment conditions are regarded as different cellular environments; the Ribo-seq ribosome expression profiles under different cell types and experimental treatment conditions include but are not limited to databases such as RPFdb and RiboSeq.Org, the Ribo-Seq data in the RPFdb database is used as the main training data, and the RiboSeq.Org database is used as a supplement; proteins include antibodies and enzymes. The present invention inputs the mRNA coding sequence, cell / tissue environment characterization and mRNA expression level into the deep learning model for training, wherein the cell / tissue environment can be expanded to the transcriptome, proteome and genomic data of the cell / tissue. This allows the deep learning model to capture environmental-specific information and improve the generalization ability of the model;

[0092] Among all gene sets in the data, randomly select genes that account for the first preset number of total genes as the new gene data set;

[0093] In a specific embodiment, in the gene set of Ribo-seq data, genes accounting for 1 / 10 of the total genes are randomly selected, which are not included in the training dataset and are used as the "new gene" dataset;

[0094] Among the different cell environments from the data, cell types accounting for a second preset number of the total number of environments are randomly selected as new environment data sets;

[0095] In a specific embodiment, in the Ribo-seq dataset of different environments, cell types accounting for 3 / 10 of the total number of environments were randomly selected. These cell types were not included in the training dataset and used as the "new environment" dataset. In addition, the cell environments and genes not included in the training set were also used as the "new genes in new environments" dataset to evaluate the performance of the mRNA translation prediction model.

[0096] The data were processed to obtain expression quantification, and RNA-seq quantitative results and Ribo-seq quantitative results were obtained, representing the mRNA transcription level and mRNA translation level, respectively. To eliminate the influence of technical factors such as sequencing depth, the RNA-seq quantitative results and Ribo-seq quantitative results were converted to RPKM (Reads Per Kilobase per Million mapped reads) format. To make the data close to the standard normal distribution, log1p transformation was performed, and finally, genes with low expression (median RPKM <1) were filtered out.

[0097] The species genome sequence and annotation information were obtained from Genecode (https: / / www.gencodegenes.org / ). Transcript quantification was performed using RNA-seq. The mRNA coding sequence of each gene was represented by the transcript with the largest proportion, and the mRNA coding sequence was converted into a one-hot coding matrix. In this example, the mRNA coding sequence was converted into a one-hot coding matrix of 4×4500, where 4 corresponds to the four nucleotides A, C, G, and T.

[0098] For each sample in the data, the RNA-seq quantitative results are constructed into a one-dimensional vector with a fixed gene order to obtain a cell environment vector;

[0099] Constructing a first model based on a CNN encoder structure, wherein the input of the first model includes an mRNA coding sequence, a cell environment vector, and an mRNA expression level. The transcriptome data of different tissues or cell types are represented as the cell environment vector, and a pre-trained yield module is obtained through multiple rounds of training;

[0100] Based on the new gene dataset and the new environment dataset, the pre-trained yield module is evaluated to obtain a preset yield module.

[0101] In a further embodiment, the first model based on the CNN encoder structure, such as Figure 4 As shown in the figure, the model receives one-hot encodings of mRNA sequences, cell environment vectors, and mRNA expression data as input, and outputs a yield score through multiple convolutional layers, an attention mechanism, and a fully connected layer. Key features of the model include multi-level sequence feature extraction: four parallel convolutional detectors are used to extract different features of the DNA sequence; an attention mechanism: the weights of DNA sequence features are adjusted based on mRNA expression characteristics; and multimodal fusion: the simultaneous consideration of DNA sequence features and mRNA expression characteristics.

[0102] In a further embodiment, evaluating the pre-trained yield module specifically includes:

[0103] In the gene set of Ribo-seq data, 1 / 10 of the total number of genes were randomly selected. These genes were not included in the training dataset and used as the "new gene" dataset. In the Ribo-seq dataset of different environments, 3 / 10 of the total number of cell types were randomly selected. These cell types were not included in the training dataset and used as the "new environment" dataset. In addition, a "new gene in new environment" dataset was defined to evaluate the performance of the mRNA translation prediction model. The prediction performance on the test set was 0.81R2 goodness of fit, as shown in the figure. Figure 5 shown.

[0104] Example 3

[0105] This embodiment further illustrates the functional modules based on Embodiment 1 and Embodiment 2.

[0106] like Figure 6 Shown is a functional module that uses a neural network structure to predict the functional score of a given amino acid sequence.

[0107] The functional modules of the present invention receive an amino acid sequence as input and output a functionality score. Functionality herein includes, but is not limited to, antibody affinity, hydrophobicity, enzyme catalytic activity, substrate specificity, signal peptide, immunogenicity, subcellular localization, and the like.

[0108] Proteins have a variety of functions. For example, modifying enzymes can improve their catalytic efficiency and stability, and modifying antibodies can enhance their specificity and affinity. The present invention uses antibody affinity optimization as an example to illustrate the following:

[0109] Step 1: Model construction. The affinity prediction model uses a convolutional neural network structure, such as Figure 7 The model receives one-hot encoded amino acid sequences as input and outputs affinity scores through multiple layers of convolution, pooling, and fully connected layers.

[0110] Step 2: Model training. The amino acid sequence of the AB1101 dataset from the literature was encoded and converted into a 21×200 one-hot encoding matrix (21 corresponds to 20 common amino acids plus a terminator, and 200 is the maximum sequence length). This dataset contains 645 single-point mutations and 456 multi-point mutations in 32 different antigen-antibody complexes. The model consists of three convolutional layers, each followed by a ReLU activation function and a max pooling layer, and finally outputs the affinity score through a fully connected layer. The model architecture parameters are as follows: First convolutional layer: 21 input channels, 64 output channels, kernel size 3×1, stride 1; Second convolutional layer: 64 input channels, 128 output channels, kernel size 3×1, stride 1; Third convolutional layer: 128 input channels, 256 output channels, kernel size 3×1, stride 1; Each convolutional layer is followed by a max pooling layer with kernel size 2×1 and stride 2; Fully connected layer 1: 256×100 input, 512 output; Fully connected layer 2: 512 input, 1 output. To improve model robustness, a dropout strategy of 0.5 was applied during training.

[0111] Step 3: Model evaluation. The experimental results of the affinity prediction model compared with the other three models on three public antigen-antibody affinity prediction datasets are shown in Table 1.

[0112] Table 1

[0113]

[0114] Example 4

[0115] This embodiment further explains the optimization module based on Embodiments 1 to 3.

[0116] like Figure 8 As shown, this is a single-objective optimization module that accepts the score of the initial sequence fed into a single prediction model as the optimization module input and outputs the target optimized sequence. Its methods include but are not limited to activation maximization (calculating the gradient of the input (i.e., the partial derivative of the predicted output with respect to the input), adjusting the input along the gradient direction to improve the target output), saturation mutation (initializing a set of random sequences, using the prediction model to evaluate the expression level of each sequence, selecting sequences with high expression levels for mutation, and iteratively updating until the optimal solution is found), GAN / diffusion generation (using a generator network to generate new sequences, using the prediction model as a discriminator to determine whether the sequence expression level is high, and training the generator so that the generated sequence can deceive the prediction model and obtain the target output), reinforcement learning (sequence optimization problems can be viewed as reinforcement learning tasks, where: the state is the current sequence; the action is an operation such as mutation, deletion, and insertion; the reward is the output of the prediction model, and deep reinforcement learning can be used to train the agent to find the sequence with the highest expression level).

[0117] Within the framework of the genetic code, there is a defined correspondence between nucleic acid sequences and amino acid sequences: each specific codon (nucleic acid triplet) strictly corresponds to one amino acid, while most amino acids can be encoded by multiple synonymous codons. This characteristic is called codon degeneracy. This asymmetric correspondence provides unique challenges and opportunities for sequence optimization. While keeping the target amino acid sequence unchanged, this embodiment can optimize the expression characteristics of the nucleic acid sequence (such as codon usage preference, mRNA stability, etc.) by selecting different synonymous codons. Conversely, any modification to the nucleic acid sequence must ensure that the amino acid sequence it encodes meets the functional requirements of the protein. It is this mutually restrictive and interdependent relationship that makes it difficult for traditional single sequence optimization methods to achieve global optimization. The core value of collaborative optimization technology is that it can simultaneously consider the functional constraints of the amino acid sequence and the expression characteristics of the nucleic acid sequence. Through a multi-objective balance strategy, it maximizes its expression efficiency and yield while ensuring the functional integrity of the protein. Conventional multi-objective optimization techniques do not take into account the above-mentioned codon degeneracy. Therefore, in the technical framework of this embodiment, the corresponding relationship between amino acid sequence and nucleic acid sequence is integrated into the technical framework, and the nucleic acid sequence is optimized while optimizing the amino acid sequence, thereby solving the sequence dependency relationship and connecting to conventional multi-objective optimization methods, such as Figure 9The figure shows a multi-objective optimization module, including but not limited to the weighted sum method (directly converting the weighted sum of multiple objectives into a single-objective optimization), constrained optimization (selecting one objective to optimize while setting constraints on other objectives), reinforcement learning + multi-objective rewards (defining a reinforcement learning environment to enable the intelligent agent to explore different optimization schemes and combining multiple objectives in the reward function), etc. Figure 10 The figure shows the weighted sum method in the multi-objective optimization of this patent. Combining the prediction results of the two modules, the multi-objective collaborative optimization of protein sequences and nucleic acid sequences is achieved through dynamic weight adjustment. It receives the optimization weight, the number of iterations (used to determine the total number of iteration cycles of the algorithm), the predicted value of the initial sequence protein expression, the predicted value of the functional score and the maximum number of mutations (used to limit the upper limit of the number of amino acid / nucleotide sites allowed to change in a single iteration, determined by the annealing temperature) as input, and outputs the optimized amino acid and nucleic acid sequences.

[0118] This embodiment provides an optimization module based on an improved simulated annealing algorithm that can search for the optimal solution in the amino acid sequence space while balancing the two objectives of affinity and expression. The main steps of the optimization process are as follows: Figure 11 Shown, including:

[0119] 1. Input the initial amino acid sequence and set parameters such as affinity weight, number of iterations, initial temperature, and end temperature;

[0120] 2. Calculate the affinity score and expression score of the initial sequence to obtain a weighted comprehensive score;

[0121] 3. Enter the simulated annealing cycle: a. Calculate the number of mutations based on the current temperature and randomly mutate several amino acids in the current sequence; b. Calculate the affinity score of the mutated sequence; c. Perform codon optimization on the mutated amino acid sequence to find the mRNA encoding with the highest expression; d. Calculate the expression score of the mutated sequence; e. Combine the affinity score and expression score to calculate a weighted composite score; f. Decide whether to accept the mutation based on the change in the composite score; g. Lower the temperature and proceed to the next iteration.

[0122] 4. Output the optimal amino acid sequence and the corresponding optimal mRNA coding sequence.

[0123] In this embodiment, the temperature scheduling adopts a linear reduction strategy, and the number of mutations decreases as the temperature decreases, so as to achieve a transition from global search to local refinement. In this embodiment, temperature is a concept borrowed from the physical annealing process, which is used to control the search range of the algorithm and the ability to jump out of the local optimum. The initial temperature is high → the system is "active", and the probability of accepting inferior solutions (poor solutions) is high, which can jump out of the local optimum. The temperature gradually decreases → the system is "calm", and is more inclined to accept better solutions, so as to achieve refinement of local search. This mechanism allows the algorithm to: search widely in the early stage to explore the entire solution space; converge to the vicinity of the optimal solution in the later stage and perform detailed optimization. The calculation of the acceptance probability adopts a deterministic strategy, that is, only accepting better solutions, thereby improving the optimization efficiency.

[0124] For yield optimization, we use a codon optimization method based on greedy search. Since the same amino acid can be encoded by multiple codons, the mRNA sequence can be optimized to increase expression without changing the protein sequence.

[0125] The specific steps for codon optimization are as follows:

[0126] 1. For each amino acid in the amino acid sequence, obtain all possible codons;

[0127] 2. For amino acids with only one possible codon, use that codon directly;

[0128] 3. For amino acids with multiple possible codons, test each codon in turn: a. Add the current codon to the constructed DNA sequence; b. Use the expression prediction model to evaluate the expression score of the current mRNA sequence; c. Select the codon with the highest score;

[0129] 4. Finally, the mRNA sequence with the optimal expression level is obtained.

[0130] Optimization effect verification: This example shows the optimization results of an antibody variable region sequence. The initial sequence is the 1DVF complex from the PDB database:

[0131] DIVLTQSPASLSSASVGETVTITCRASGNIHNALAWYQQKQGKSPQLLVYYTTT

[0132] LADGVPSRF

[0133] SGSGSGTQYSLKINSLQPEDFGSYYCQHFWSTPRTFGGGTKLEIKR

[0134] The optimization parameters were set as follows: affinity weight (0.5), number of iterations (500), maximum number of mutations (5), initial temperature (2.0), and termination temperature (0.1).

[0135] The changes in affinity score, expression score and comprehensive score during the optimization process are as follows: Figure 12 After 500 iterations, the optimal amino acid sequence and mRNA sequence obtained significantly improved the comprehensive performance of affinity and expression. Compared with the initial sequence, the affinity score increased by 15%, the expression score increased by 22%, and the comprehensive score increased by 20%.

[0136] The same or similar reference numerals correspond to the same or similar components;

[0137] The terms used in the drawings to describe positional relationships are for illustrative purposes only and should not be construed as limiting this patent;

[0138] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A sequence optimization method for balancing function and yield, characterized in that: The following steps are involved: Obtain the initial amino acid sequence, cell environment, and nucleic acid expression level; Obtaining an initial nucleic acid sequence according to the initial amino acid sequence; Calculating a yield score based on the initial nucleic acid sequence, the cell environment, and the nucleic acid expression level using a machine learning-based yield prediction module, wherein the yield score corresponds to the amount of protein produced by translation of the nucleic acid sequence in a specific environment; Calculating at least one functional score corresponding to a specific biological function of the protein by at least one function prediction module based on machine learning; Based on the yield score and at least one function score, the initial amino acid sequence and the initial nucleic acid sequence are collaboratively optimized using a multi-objective optimization algorithm in an optimization module to obtain an optimized amino acid sequence and an optimized nucleic acid sequence, wherein the multi-objective optimization algorithm aims to maximize the yield score and the function score.

2. The method for optimizing the sequence for balancing function and yield according to claim 1, characterized in that: The cell environment includes transcriptome data of different tissues or cell types, and the transcriptome data is used to construct a cell environment vector to improve the accuracy of the yield prediction module and the function prediction module.

3. The method for optimizing the sequence for balancing function and yield according to claim 1, characterized in that: According to the initial amino acid sequence, an initial nucleic acid sequence is obtained by codon optimization, wherein the codon optimization comprises: For each amino acid in the amino acid sequence, obtain all possible codons; According to the target cell environment and the nucleic acid expression level, the yield score of each codon is evaluated by the yield prediction module; The codon with the highest yield score is selected to construct the initial nucleic acid sequence.

4. The method for optimizing the sequence for balancing function and yield according to claim 1, wherein: The training process of the preset yield module includes: Collect transcriptome and translation efficiency data in the target cell environment; Based on the transcriptome data and translation efficiency data, a training data set is constructed, wherein each sample includes a nucleic acid coding sequence, a cell environment vector, and a nucleic acid expression level; Training a neural network model, including a convolutional neural network or a recurrent neural network, using the nucleic acid coding sequence, the cell environment vector, and the nucleic acid expression level as input to predict translation efficiency; The model performance is evaluated through cross-validation to obtain the preset yield module.

5. The sequence optimization method for balancing function and yield according to claim 4, characterized in that: The transcriptome data is RNA-seq data, and the translation efficiency data is Ribo-seq data; The construction of the training data set includes: Quantify and normalize RNA-seq and Ribo-seq data and filter low-expression genes; Convert the nucleic acid coding sequence into a one-hot coding matrix; RNA-seq expression quantification was sorted by fixed genes to construct a cell environment vector.

6. The sequence optimization method for balancing function and yield according to claim 1, characterized in that: If there are two or more functional modules, each functional prediction module predicts different protein functions based on a machine learning algorithm to obtain different functional scores, wherein the different functions include antibody affinity, hydrophobicity, enzyme catalytic activity, substrate specificity, signal peptide, immunogenicity and subcellular localization.

7. The method for sequence optimization for balancing function and yield according to claim 1, characterized in that: The optimization module is constructed through multi-objective optimization methods such as weighted addition, constrained optimization, or reinforcement learning + multi-objective rewards.

8. The method for sequence optimization for balancing function and yield according to claim 1, characterized in that: When only the yield prediction module or one function prediction module is used, the multi-objective optimization algorithm is simplified to a single-objective optimization, taking the yield score or one function score as input and outputting an optimized amino acid sequence or an optimized nucleic acid sequence.

9. The method for sequence optimization for balancing function and yield according to claim 8, characterized in that: The optimization module is constructed through activation maximization, saturation mutation, adversarial generative network / diffusion model generation or reinforcement learning methods.

10. The sequence optimization method for balancing function and yield according to claim 7 or 8, characterized in that: The optimization module uses a simulated annealing algorithm to collaboratively optimize the initial amino acid sequence and the initial nucleic acid sequence.

Citation Information

Patent Citations

  • Sequence based engineering and optimization of single chain antibodies

    CN101688200A

  • Codon optimized mRNA vaccine against novel coronavirus

    CN115814074A

  • Antibody optimization method and system

    CN117894373A

  • System for designing new antibody for specific antigen based on antibody structure fine-tuning protein large language model

    CN118658515A

  • Unified protein modification method and device and computer equipment

    CN119626323A

Cited By

  • Nucleic acid sequence optimization method, system and equipment, storage medium and program product

    CN121747703A