Viscosity prediction method for antibody drug development
By extracting the atomic spatial charge distribution characteristics based on the three-dimensional structure of the antibody and using a machine learning model, the antibody viscosity can be predicted quickly and accurately, solving the problems of low viscosity prediction accuracy and high computational complexity in existing technologies and reducing the risk of antibody drug development.
Patent Information
- Application Number
- CN202480002202.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-04
- Filing Date
- 2024-04-03
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-04-03
AI Technical Summary
Existing viscosity prediction methods for antibody drug development have low accuracy, high computational complexity, and are time-consuming, making it difficult to effectively evaluate the viscosity of high-concentration antibodies in the early stages, increasing development risks.
By extracting the atomic spatial charge distribution features based on the three-dimensional structure of the antibody and using machine learning models such as random forests, the antibody viscosity can be quickly and accurately predicted. AlphaFold2 is used to construct a static 3D structural model of the antibody, and the spatial charge map (SCM) of amino acid residues is extracted as features and input into the random forest model for prediction.
It enables rapid and accurate viscosity prediction in early antibody drug development, reduces trial and error costs, and improves the accuracy and efficiency of viscosity prediction.
Smart Images

Figure CN119032398B_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims priority to Chinese patent application No. 202310354522.0 filed on April 4, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to antibody drug development, and more particularly to an antibody viscosity prediction method based on a spatial charge map and a machine learning model for antibody drug development. Background Art
[0004] Monoclonal antibodies (mAbs) are an important class of drugs for treating a wide range of diseases and are becoming increasingly popular with over 100 approved and commercially available therapeutic antibody products. Although antibodies have become the fastest-growing class of therapeutic drugs on the market, developing them for therapeutic applications remains challenging.
[0005] Besides achieving the desired biological activity, there are numerous other hurdles on the path of a candidate antibody toward clinical reality.
[0006] These include:
[0007] (1) Manufacturability risks, such as poor expression, instability during viral inactivation or elution, and adhesion to purification columns;
[0008] (2) Formulation risks such as chemical and conformational instability, self-association, high viscosity, and aggregation;
[0009] (3) In vivo problems such as lack of specificity, immunogenicity, precipitation after administration, and rapid clearance.
[0010] Among the aforementioned developability characteristics, low viscosity is one of the most important factors in new drug development. However, due to the limited number of molecules with available sequences, biophysical property data, and sufficient materials, it is difficult to evaluate the viscosity profile of high-concentration antibodies in the early stages of development. Therefore, computational tools that predict antibody viscosity using only the designed / discovered antibody sequence are important for reducing the risk of therapeutic antibody development, especially in the early stages.
[0011] To predict the viscosity of antibodies, different researchers have developed methods ranging from statistical methods based on physicochemical properties to machine learning or deep learning methods.
[0012] For example, Sharma et al. proposed a linear equation for calculating the viscosity of an antibody at 180 mg / ml, which includes three factors: the net charge of the antibody variable region (Fv), charge symmetry, and hydrophobicity. See VK Sharma, TW Patapoff, B. Kabakoff, S. Pai, E. Hilario, B. Zhang, et al., "In silico selection of therapeutic antibodies for development: viscosity, clearance, and chemical stability", Proc Natl Acad Sci, 111(52) (2014), pp. 18601-18606; PMID: 25512516; http: / / dx.doi.org / 10.1073 / pnas.1421779112. The entire contents of this document are hereby incorporated into the present disclosure by reference, making it a part of the content of the present disclosure.
[0013] The Spatial Charge Map (SCM) method calculates the exposed surface negative charge distribution of the Fv region through molecular dynamics simulation to obtain a corresponding score to evaluate viscosity. See U.S. Patent Application Publication No. US2016 / 0132634A1 (granted U.S. Patent No. US10,762,980B2) by the Massachusetts Institute of Technology (MIT) and Novartis Pharma AG. The entire contents of this patent document are hereby incorporated into the present disclosure by reference, making it a part of the contents of the present disclosure.
[0014] DeepSCM, based on deep learning, trains a convolutional neural network surrogate model on 6596 antibody Fv regions to fit SCM values. See Pin-Kuang Lai, “DeepSCM: An efficient convolutional neural network surrogate model for the screening of therapeutic antibody viscosity,” Computational and Structural Biotechnology Journal, Volume 20, 2022, pages 2143-2152; ISSN 2001-0370; https: / / doi.org / 10.1016 / j.csbj.2022.04.035. The entire contents of this document are hereby incorporated by reference into this disclosure, making it a part of the present disclosure.
[0015] Despite the success of these methods, they still have some drawbacks. Low prediction accuracy (ACC) is an area in which these methods urgently need improvement. Computational complexity is also a concern. To accurately capture the dynamic three-dimensional structure of antibody molecules, molecular dynamics simulations are the only option. However, molecular dynamics simulations are extremely computationally demanding, time-consuming, and expensive. Summary of the Invention
[0016] To address the low accuracy, time-consuming, and expensive calculations of existing viscosity prediction methods, this paper proposes a viscosity prediction method for antibody drug development. Simply by providing the variable region amino acid sequence of an antibody, the viscosity value can be quickly, accurately, and cost-effectively calculated. This allows for targeted viscosity screening in the early stages of antibody drug development, reducing trial-and-error costs.
[0017] According to a first aspect of the present disclosure, a method for predicting antibody viscosity is provided. The method may include the following steps: extracting antibody features based on the spatial charge distribution of atoms in the antibody's three-dimensional structure; and inputting the extracted features into a machine learning-based antibody viscosity prediction model to obtain a predicted value for the antibody viscosity.
[0018] In the method according to the first aspect of the present disclosure, the feature extraction of the antibody based on the atomic spatial charge distribution in the three-dimensional structure of the antibody may include: using a protein structure prediction tool to model the input antibody sequence to obtain the three-dimensional structure of the antibody, wherein the antibody sequence is the variable region amino acid sequence of the antibody; based on the three-dimensional structure of the antibody, calculating the SCM value of the antibody amino acid residues as the extracted feature.
[0019] Preferably, the protein structure prediction tool is AlphaFold2, HelixFold, SWISS-MODEL or Uni-Fold.
[0020] In the method according to the first aspect of the present disclosure, the three-dimensional structure is described by a PDB file and a PQR file.
[0021] Preferably, the three-dimensional structure of the antibody obtained by the modeling is described using a PDB file. The PDB file is converted into a PQR file using the PDB2PQR tool. The charge value of each atom is extracted based on the PQR file.
[0022] In the method according to the first aspect of the present disclosure, calculating the SCM value of the antibody amino acid residues based on the three-dimensional structure of the antibody as the extracted feature further includes: extracting the charge value of each atom in the amino acid residues of the antibody based on the three-dimensional structure of the antibody; calculating the SCM value of each atom; and calculating the SCM value of each amino acid residue based on the SCM value of each atom.
[0023] Preferably, the calculating of the SCM value of each atom comprises: for each atom i, calculating its SCM value by the following formula:
[0024]
[0025] Where, represents the charge value of the jth atom in the statistical range.
[0026] Preferably, the calculating the SCM value of each amino acid residue based on the SCM value of each atom comprises: calculating the SCM value of each amino acid residue by the following formula:
[0027]
[0028] in, 原子, represents the SCM value of the i-th atom in the k-th amino acid residue.
[0029] Preferably, when the SCM value of the amino acid residue is greater than 0, its SCM value is replaced by 0.
[0030] In the method according to the first aspect of the present disclosure, the machine learning-based antibody viscosity prediction model is trained by the following steps: obtaining a data set, the data set including multiple training samples, each of the training samples including the variable region amino acid sequence and viscosity value of the sample antibody; performing feature extraction on the variable region amino acid sequence of each training sample based on the atomic spatial charge distribution in the antibody three-dimensional structure; using a machine learning method, cross-validating on the data set to determine the machine learning model parameters; using the machine learning model parameters, training the machine learning model on the data set to obtain the final machine learning-based antibody viscosity prediction model. The training uses the extracted features as the input data of the machine learning model and the antibody viscosity value as the output data of the machine learning model.
[0031] Preferably, the machine learning method includes at least one of the following: random forest, support vector machine regression, Gaussian regression, linear regression; more preferably, the machine learning method is random forest.
[0032] Preferably, cross validation is performed on the dataset using a leave-one-out method.
[0033] According to the second aspect of the present disclosure, a machine learning-based antibody viscosity prediction model is provided, wherein the prediction model is trained by the following steps: obtaining a data set, wherein the data set includes multiple training samples, each of the training samples includes the variable region amino acid sequence and viscosity value of the sample antibody; performing feature extraction on the variable region amino acid sequence of each training sample based on the atomic spatial charge distribution in the antibody three-dimensional structure; using a machine learning method, cross-validating on the data set to determine the machine learning model parameters; using the machine learning model parameters, training the machine learning model on the data set to obtain the final machine learning-based antibody viscosity prediction model. Wherein, the training uses the extracted features as the input data of the machine learning model and uses the antibody viscosity value as the output data of the machine learning model.
[0034] Preferably, the machine learning method includes at least one of the following: random forest, support vector machine regression, Gaussian regression, linear regression; more preferably, the machine learning method is random forest.
[0035] Preferably, cross validation is performed on the dataset using a leave-one-out method.
[0036] Preferably, extracting features of the variable region amino acid sequence of each training sample based on the spatial charge distribution of atoms in the three-dimensional structure of the antibody includes: using a protein structure prediction tool to model the variable region amino acid sequence of each training sample to obtain the three-dimensional structure of the sample antibody; and calculating the SCM value of the amino acid residues of the sample antibody based on the three-dimensional structure of the sample antibody as the extracted feature.
[0037] Preferably, the protein structure prediction tool is AlphaFold2, HelixFold, SWISS-MODEL or Uni-Fold.
[0038] Preferably, the three-dimensional structure is described by a PDB file and a PQR file.
[0039] Preferably, the three-dimensional structure of the sample antibody obtained by the modeling is described by a PDB file, the PDB file is converted into a PQR file using the PDB2PQR tool, and the charge value of each atom is extracted based on the PQR file.
[0040] Preferably, the calculating of the SCM value of the amino acid residues of the sample antibody based on the three-dimensional structure of the sample antibody as the extracted feature further includes: extracting the charge value of each atom in the amino acid residues of the sample antibody based on the three-dimensional structure of the sample antibody; calculating the SCM value of each atom; and calculating the SCM value of each amino acid residue based on the SCM value of each atom.
[0041] Preferably, the calculating of the SCM value of each atom comprises: for each atom i, calculating its SCM value by the following formula:
[0042]
[0043] Where, represents the charge value of the jth atom in the statistical range.
[0044] Preferably, the calculating the SCM value of each amino acid residue based on the SCM value of each atom comprises: calculating the SCM value of each amino acid residue by the following formula:
[0045]
[0046] in, 原子, represents the SCM value of the i-th atom in the k-th amino acid residue.
[0047] Preferably, when the SCM value of the amino acid residue is greater than 0, its SCM value is replaced by 0.
[0048] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is provided for storing a computer program. The computer program includes instructions. When executed by a processor of an electronic device, the instructions cause the electronic device to implement the antibody viscosity prediction method according to the first aspect of the present disclosure.
[0049] According to a fourth aspect of the present disclosure, a system for predicting antibody viscosity is provided. The system comprises a processor, a memory, and a computer program. The computer program is stored in the memory and configured to be executed by the processor. The computer program includes instructions for implementing the antibody viscosity prediction method according to the first aspect of the present disclosure.
[0050] According to the fifth aspect of the present disclosure, a method for constructing an antibody viscosity prediction model based on machine learning is provided, comprising: obtaining a data set, the data set comprising multiple training samples, each of the training samples comprising a variable region amino acid sequence and viscosity value of a sample antibody; performing feature extraction on the variable region amino acid sequence of each training sample based on the atomic spatial charge distribution in the antibody's three-dimensional structure; employing a machine learning method to perform cross-validation on the data set to determine machine learning model parameters; and using the machine learning model parameters to train the machine learning model on the data set to obtain a final machine learning-based antibody viscosity prediction model. The training method for the antibody viscosity prediction model of the second aspect of the present disclosure is also applicable to the fifth aspect of the present disclosure.
[0051] The viscosity prediction method proposed in this disclosure uses machine learning methods, such as using a random forest (RF) model to rapidly construct a static 3D structural model of the antibody using AlphaFold2. The spatial charge map (SCM) of the amino acid residues is extracted as features for training the RF model. After designing a feature filter, the extracted features are aligned with the Kabat numbering system, and positive values in the extracted features are set to zero. The obtained features are then input into the constructed model to provide a predicted antibody viscosity value. The model constructed in this disclosure was based on 40 publicly available antibody drug samples and was constructed and validated using a leave-one-out cross-validation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The present disclosure will be more fully understood from the following detailed description taken in conjunction with the accompanying drawings, in which like elements are numbered in a similar manner, and in which:
[0053] Figure 1 4 is a flow chart of a method for predicting antibody viscosity according to an embodiment of the present disclosure.
[0054] Figure 2 3 is a schematic diagram of the association between the antibody viscosity prediction method and the prediction model training method according to an embodiment of the present disclosure.
[0055] Figure 3 is a more detailed flowchart of the antibody feature extraction method according to an embodiment of the present disclosure.
[0056] Figure 4 Detailed flowchart of the method for training the machine learning-based antibody viscosity prediction model used in the present disclosure.
[0057] Figure 5 It is a scatter plot comparing the prediction results of the random forest model and the experimental measurement results.
[0058] Figure 6 It is a scatter plot comparing the prediction results of the support vector machine regression model and the experimental measurement results.
[0059] Figure 7 It is the ROC curve of the classification performance using the random forest model. DETAILED DESCRIPTION
[0060] Unless otherwise defined, technical and scientific terms used in the present disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs.
[0061] The technical solutions of the present disclosure are further described in detail below through examples and in conjunction with the accompanying drawings. Unless otherwise stated, the methods and materials of the embodiments described below are conventional products that can be purchased on the market. Those skilled in the art will understand that the methods and materials described below are merely exemplary and should not be construed as limiting the scope of the present disclosure.
[0062] From the overall perspective of forecasting methods, Figure 1 A flowchart of a method for predicting antibody viscosity according to an embodiment of the present disclosure is shown.
[0063] like Figure 1 As described in, the antibody viscosity prediction method 100 according to an embodiment of the present disclosure begins at step S110, in which features of the antibody are extracted based on the atomic spatial charge distribution in the three-dimensional structure of the antibody.
[0064] Next, in step S120 , the features extracted in step S110 are input into an antibody viscosity prediction model based on machine learning to obtain a predicted value of antibody viscosity.
[0065] Spatial Charge Map (SCM) is a technique used to analyze computer-generated information about the predicted structure of a protein. As previously mentioned, in methods such as US Patent Application Publication No. US 2016 / 0132634A1, molecular dynamics simulations are used to determine the distribution of negative surface charges on the Fv region, which is then used to generate a corresponding score to estimate viscosity.
[0066] More specifically, in prior art such as U.S. patent application publication US2016 / 0132634A1, SCM can be used to implement it as a structure-based phenomenological molecular modeling tool to identify patches of charged (amino acid) residues on the surface of proteins. For example, in such a tool, the structure of a protein (such as an antibody) is input. Viscosity predictions are directly generated based on computer-generated information about the antibody structure. The specific viscosity prediction uses the SCM score obtained by combining the SCM of each molecule with the Heaviside function, and the viscosity is determined based on the SCM score. DeepSCM can be seen as an improvement on this method, which uses deep learning to implement the above process.
[0067] The present disclosure differs from the prior art described above in that, after extracting antibody features using SCM, the present disclosure does not use the combinatorial operations employed in the aforementioned methods to calculate the SCM score and thereby determine viscosity. Instead, the present disclosure predicts antibody viscosity using a machine learning-based prediction model. The machine learning-based prediction model described herein is trained by extracting features based on SCM values from training samples and can therefore be used to predict antibody viscosity based on the SCM feature information of the antibody to be predicted.
[0068] Figure 2 A schematic diagram 200 is provided showing the association between the antibody viscosity prediction method and the prediction model training method according to an embodiment of the present disclosure. Figure 2 In the schematic diagram 200 shown, the left side of the dotted line shows the model training method, and the right side of the dotted line shows the method of using the model to make predictions.
[0069] Figure 2 The process on the right side of the dotted line is actually Figure 1 The antibody viscosity prediction method according to an embodiment of the present disclosure is shown. Figure 2 The process on the left side of the dotted line will be described in detail below (for example Figure 4 and its corresponding text description). Figure 2 It can be seen that both the training and actual prediction processes require the use of a feature extraction method based on the atomic spatial charge distribution in the three-dimensional structure of the antibody (in a preferred embodiment, based on the SCM value), which will be described in detail below (e.g. Figure 3and its corresponding text description). For the training samples, it includes both the variable region amino acid sequence of the antibody and the viscosity value of the antibody. Therefore, during the model training process, the sequence of the antibody in the training sample is used to extract features (such as SCM values). The machine learning model is fully trained by using the extracted features (such as SCM values) as input data, the initial model parameters of the machine learning model (such as the random forest model or the support vector machine regression model), and the antibody viscosity value data in the training sample as output data, thereby obtaining the final antibody viscosity prediction model based on machine learning. For the trained prediction model, it can be put into the actual work of antibody viscosity prediction, that is, Figure 2 The process to the right of the middle dashed line.
[0070] The following describes in detail the detailed processes of the antibody feature extraction method based on the atomic spatial charge distribution in the three-dimensional structure of the antibody and the training method of the antibody viscosity prediction model based on machine learning according to the embodiments of the present disclosure.
[0071] Figure 3 is a more detailed flowchart of the antibody feature extraction method according to an embodiment of the present disclosure.
[0072] Generally, Figure 1 Step S110 in the prediction method, i.e., extracting features of the antibody based on the spatial charge distribution of atoms in the three-dimensional structure of the antibody, can be achieved by the following steps: First, the input antibody sequence is modeled using a protein structure prediction tool to obtain the three-dimensional structure of the antibody. Here, the antibody sequence refers to the amino acid sequence of the variable region of the antibody. Then, based on the three-dimensional structure of the antibody, the SCM values of the antibody amino acid residues are calculated as the extracted features. In a preferred embodiment of the present disclosure, the protein structure prediction tool described herein for modeling the input antibody sequence can be an artificial intelligence tool, such as AlphaFold2. In this preferred embodiment, a static 3D structural model of the antibody is quickly constructed by AlphaFold2.
[0073] AlphaFold2 is an artificial intelligence program developed by DeepMind that can be used to predict most protein structures with extremely small errors from the actual structure.
[0074] Those skilled in the art will appreciate that, in addition to AlphaFold2, those skilled in the art may also use various other protein structure prediction tools to obtain the three-dimensional structure of the antibody. For example, the protein structure prediction tool may be HelixFold (see Fang X, Wang F, Liu L, et al. Helixfold-single: Msa-free protein structure prediction by using protein language model as an alternative [J]. arXiv preprint arXiv: 2207.13921, 2022.), SWISS-MODEL (see https: / / swissmodel.expasy.org / ), or Uni-Fold (see https: / / github.com / dptech-corp / Uni-Fold).
[0075] like Figure 3 As shown in , the antibody feature extraction method 300 according to an embodiment of the present disclosure can start at step S310, that is, as described above, according to a preferred embodiment of the present disclosure, in step S310, AlphaFold2 is used to model the input antibody sequence to obtain its three-dimensional structure.
[0076] Those skilled in the art will appreciate that the three-dimensional structure of an antibody can be expressed and recorded using specialized file formats. For example, in a preferred embodiment of the present disclosure, the three-dimensional structure is described using a PDB file and a PQR file. More specifically, in a preferred embodiment of the present disclosure, the three-dimensional structure of the modeled antibody is described using a PDB file. In this embodiment, the PDB file is converted to a PQR file using the PDB2PQR tool. Those skilled in the art will appreciate that the charge values of each atom can be extracted based on the PQR file.
[0077] like Figure 3 As shown in FIG, in step S320 of the antibody feature extraction method 300 according to an embodiment of the present disclosure, the three-dimensional structure PDB obtained in step S310 is converted into PQR using PDB2PQR.
[0078] The Protein Data Bank (PDB) is an international protein structure database. The PDB file format is used to store structural information about proteins and other biomolecules. PDB files are stored in text format and consist of a header and atomic structure data records, which contain structural information about the biomolecule, such as atomic coordinates, connectivity, structural features, and temperature factors.
[0079] PQR is a special file format that adds Q (charge) and R (radius) information to the PDB format. Specifically, the PQR format provides a simple way to add parameter information: the occupancy and temperature columns in the PDB structure file are replaced with charge (Q) and radius (R).
[0080] PDB2PQR is a tool used to convert PDB to PQR.
[0081] In the embodiments of the present disclosure, the specific process of calculating the SCM value of the antibody amino acid residues may further include: extracting the charge value of each atom in the antibody amino acid residues based on the three-dimensional structure of the antibody. Figure 3 As shown in , for the three-dimensional structure converted from PDB to PQR, in step S330, the charge value of each atom can be extracted according to the PQR file obtained in step S320.
[0082] Next, we can calculate the SCM value of each atom. Figure 3 As shown in FIG, in step S340, for each atom, its SCM value is calculated. More specifically, for each atom i, its SCM value is calculated by the following formula:
[0083]
[0084] Where, represents the charge value of the jth atom in the statistical range.
[0085] In addition, the operator <> in the above formula represents the ensemble average in a statistical sense. The "exposed residue" mentioned here can be defined as the solvent accessible surface area of the side chain atoms being greater than or equal to For a more specific definition and explanation of this calculation method, please refer to the description in US 2016 / 0132634 A1 and DeepSCM introduced above.
[0086] Next, the SCM value for each amino acid residue can be calculated. Figure 3 As shown in , in step S350, the SCM value of each amino acid is calculated based on the SCM value of each atom. More specifically, the SCM value of each amino acid residue is calculated by the following formula:
[0087]
[0088] in, 原子, represents the SCM value of the i-th atom in the k-th amino acid residue.
[0089] The calculated SCM value for each amino acid residue can be mapped to the numbered coordinate system of the antibody sequence.
[0090] For example, the coordinate system numbering for the antibody sequence can be set as follows. Figure 3 As shown in step S360, the light chain and heavy chain of the antibody sequence are numbered using the ANARCI tool, the numbering scheme adopts Kabat numbering, and the numbering coordinates take the complete coordinates of the complete numbering system.
[0091] ANARCI is a tool for classifying and numbering antibody and T-cell receptor amino acid variable region sequences. It can annotate sequences using, for example, the Kabat numbering scheme.
[0092] The Kabat numbering scheme is a numbering scheme developed based on the location of regions of high sequence variation among sequences of the same domain type.
[0093] like Figure 3 As shown in step S370, the SCM value calculated in step S350 is mapped to the coordinate system numbered in step S360. It should be understood by those skilled in the art that for coordinates without corresponding amino acids, 0 is used to fill in the gaps.
[0094] In addition, if the calculated SCM value of an amino acid residue is greater than 0, its SCM value is replaced by 0.
[0095] Through the above series of steps, according to the method disclosed in the present invention, the SCM value of the variable region amino acid sequence of the antibody can be extracted as the extracted feature, and then the feature can be used as input to predict the viscosity of the antibody through a machine learning model to output the viscosity value of the antibody.
[0096] On the other hand, Figure 2 As described in, the above series of steps can also be used to extract features of the training samples based on the spatial charge distribution of atoms in the three-dimensional structure of the antibody, so that the extracted features (such as SCM values) are used as input data and input into a machine learning model with initial model parameters for training, and the viscosity values of the corresponding antibodies in the training samples are used as output data to train the training samples, so as to finally determine the model parameters and train an antibody viscosity prediction model with relatively satisfactory accuracy.
[0097] Figure 4 Detailed flowchart of the method for training the machine learning-based antibody viscosity prediction model used in the present disclosure.
[0098] Those skilled in the art should know that Figure 4Before starting the training method, a data set including multiple training samples should be obtained. Specifically, each of the multiple training samples in the data set includes the variable region amino acid sequence of the sample antibody and the viscosity value of the sample antibody.
[0099] Then, as mentioned above, it is necessary to extract features of the variable region amino acid sequence in each training sample based on the atomic space charge distribution in the antibody three-dimensional structure. Figure 4 As shown in step S410, according to Figure 3 The method is used to extract features of the variable region amino acid sequence of each training sample based on the atomic space charge distribution in the antibody three-dimensional structure.
[0100] Next, a machine learning method is used to perform cross-validation on all data sets to determine the machine learning model parameters. The machine learning method that can be used in the present disclosure may include at least one of the following: Random Forest (RF), Support Vector Regression (SVR), Gaussian Regression (GR), Linear Regression (LR), etc. Correspondingly, the machine learning model can also be an RF model, an SVR model, a GR model, an LR model, etc. Those skilled in the art should understand that, as mentioned above, when training a machine learning model, a set of initial model parameters can be set according to the type of machine learning model, and the model parameters are continuously adjusted as the training proceeds. As mentioned above, during the model training process, the extracted features (such as SCM values) are used as the input data of the machine learning model, and the viscosity values of the antibodies are used as the output data of the machine learning model to train the training samples. As the model parameters are continuously optimized, the output data of the machine learning model will match the viscosity values of the sample antibodies.
[0101] In the preferred embodiment of the present disclosure, the random forest (RF) machine learning method is adopted. In addition, during the parameter training process, the leave-one-out method can be used to perform cross-validation on all data sets. Therefore, Figure 4 As shown in , in step S420, a random forest machine learning method is adopted, and cross-validation is performed on all data sets using the leave-one-out method to determine the random forest model parameters.
[0102] Finally, the machine learning model parameters are used to train the machine learning model on all data sets to obtain the final machine learning-based antibody viscosity prediction model. Figure 4 As shown in , in step S430, the random forest model is trained on all data sets using the model parameters determined in step S420 to obtain a final prediction model.
[0103] It will be understood by those skilled in the art that by, for example Figure 4 The prediction model obtained by this method can be used to Figure 1 The prediction method shown in .
[0104] In practice, the present disclosure uses 40 commercially available antibodies and their viscosity data at a concentration of 150 mg / ml to construct and validate the model.
[0105] In order to verify the performance difference between the prediction method provided by the present disclosure and the existing prediction methods described in patent documents and non-patent documents in the prior art, several antibody viscosity prediction methods were selected for comparison. Among them, the prior art methods include two prediction methods based on spatial charge maps (SCM and DeepSCM) and Sharma Model (see VK Sharma, TW Patapoff, B. Kabakoff, S. Pai, E. Hilario, B. Zhang, et al., "In silico selection of therapeutic antibodies for development: viscosity, clearance, and chemical stability", Proc Natl Acad Sci, 111 (52) (2014), pp. 18601-18606; PMID: 25512516; http: / / dx.doi.org / 10.1073 / pnas.1421779112), and the random forest (RF) and support vector machine regression (SVR) machine learning methods disclosed in the present disclosure. Antibodies were classified based on high (>30 cP) or low (<30 cP) viscosity. Classification performance evaluation metrics are shown in Table 1. The experimental results show that the random forest (RF) and support vector machine regression (SVR) methods provided in this disclosure are effective for predicting antibody drug viscosity, and their accuracy is superior to existing methods. In particular, the random forest (RF) method outperformed existing methods in all metrics, demonstrating the effectiveness and superiority of the methods disclosed herein.
[0106] Table 1 Comparison between the methods disclosed herein and prior art methods
[0107]
[0108] The meanings or calculation formulas of the comparison indicators listed in the table are as follows:
[0109] PCC is the Pearson correlation coefficient, which ranges from -1 to 1. The closer the PCC is to 1, the higher the positive correlation between the predicted value and the true value. The closer the PCC is to -1, the higher the negative correlation between the predicted value and the true value. The closer the PCC is to 0, the weaker the correlation between the predicted value and the true value.
[0110] Accuracy (ACC) = (TP + TN) / (P + N), the number of matched samples divided by the total number of samples. Generally speaking, the higher the accuracy, the better the classifier.
[0111] Precision = TP / (TP+FP), which refers to the proportion of correctly predicted results among all the results predicted as positive samples. The closer the value is to 1, the better the performance.
[0112] Recall (Recall) = TP / (TP+FN), which refers to the ratio of the number of correctly predicted positive samples to the total number of positive samples. The closer the value is to 1, the better the performance.
[0113] Matthews Correlation Coefficient (MCC) = (TP*TN-FP*FN) / (sqrt((TP+FP)*(TP+FN)*(TN+FP)*(TN+FN))). MCC is essentially a correlation coefficient that describes the actual classification and the predicted classification. Its value range is [-1,1]. A value of 1 indicates a perfect prediction of the subject. A value of 0 indicates that the predicted result is worse than the random prediction result. -1 means that the predicted classification is completely inconsistent with the actual classification.
[0114] F1 = 2 / ((1 / Precision) + (1 / Recall)) is a statistical metric used to measure the accuracy of binary classification models. It is often used in situations where there is an imbalance between positive and negative samples. It takes into account both the precision and recall of a classification model. The F1 value can be considered a weighted average of the model's precision and recall, with a maximum value of 1 and a minimum value of 0.
[0115] Among them, TP (True Positive) and TN (True Negative) represent the number of correctly classified positive samples and negative samples, FP (False Positive) and FN (False Negative) represent the number of incorrectly classified positive samples and negative samples, and P and N represent the number of positive samples and negative samples.
[0116] Figure 5A scatter plot shows the viscosity predictions for 40 samples using the RF method described in this disclosure and the actual experimentally measured values. The horizontal axis represents the predicted values using the method described in this disclosure, while the vertical axis represents the experimentally measured viscosity values. The results show a high positive correlation between the RF predicted values and the experimentally measured values (R² = 0.44, R² = 0.666).
[0117] Figure 6 A scatter plot showing the viscosity predictions of 40 samples using SVR using the method disclosed herein and the actual experimentally measured values is shown. The horizontal axis of the figure shows the predicted values using the method disclosed herein, and the vertical axis shows the experimentally measured viscosity values.
[0118] Figure 7 This is a receiver operating characteristic (ROC) curve for the classification performance of the method disclosed herein using random forest machine learning. The horizontal axis represents the false positive rate, the vertical axis represents the true positive rate, and the area under the ROC curve corresponds to the AUC (Area Under Curve) value. The results show that the classification performance indicator AUC value is 0.92, indicating that this method has good classification performance.
[0119] In addition, those skilled in the art will recognize that the methods disclosed herein can be implemented as computer programs. As described above in conjunction with the accompanying drawings, the methods of the above embodiments are executed by one or more programs, and the instructions in the programs cause a computer or processor to execute the algorithms described in conjunction with the accompanying drawings. These programs can be stored and provided to a computer or processor using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (such as floppy disks, magnetic tapes, and hard disk drives), magneto-optical recording media (such as magneto-optical disks), CD-ROMs (compact disk read-only memories), CD-Rs, CD-R / Ws, and semiconductor memories (such as ROMs, PROMs (programmable ROMs), EPROMs (erasable PROMs), flash ROMs, and RAMs (random access memories)). Furthermore, these programs can be provided to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transient computer-readable media can be used to provide programs to a computer via wired communication paths such as wires and optical fibers or wireless communication paths.
[0120] For example, according to one embodiment of the present disclosure, a non-transitory computer-readable storage medium may be provided for storing a computer program, wherein the computer program includes instructions that, when executed by a processor of an electronic device, enable the electronic device to implement the antibody viscosity prediction method as described above.
[0121] In addition, according to the present disclosure, a system for predicting antibody viscosity may be proposed, comprising: a processor; a memory; and a computer program. The computer program is stored in the memory and configured to be executed by the processor. The computer program includes instructions for implementing the above-described method for predicting antibody viscosity.
[0122] The implementation methods of the present disclosure are not limited to those described in the above embodiments. Without departing from the spirit and scope of the present disclosure, ordinary technicians in this field can make various changes and improvements to the present disclosure in form and details, and these are all considered to fall within the scope of protection of the present disclosure.
Claims
1. A method for predicting antibody viscosity, characterized in that: The method comprises: Extract antibody features based on the atomic spatial charge distribution in the antibody's three-dimensional structure; The extracted features are input into the antibody viscosity prediction model based on machine learning to obtain the predicted value of antibody viscosity. The machine learning-based antibody viscosity prediction model uses at least one machine learning method of random forest and support vector machine regression, and The feature extraction of the antibody based on the atomic spatial charge distribution in the antibody three-dimensional structure includes: Using a protein structure prediction tool to model the input antibody sequence to obtain the three-dimensional structure of the antibody, wherein the antibody sequence is the variable region amino acid sequence of the antibody; Based on the three-dimensional structure of the antibody, the SCM value of the antibody amino acid residues is calculated as the extracted feature, Wherein, the calculation of the SCM value of the antibody amino acid residues based on the three-dimensional structure of the antibody as the extracted feature further includes: extracting the charge value of each atom in the amino acid residues of the antibody based on the three-dimensional structure of the antibody; Calculate the SCM value of each atom; The SCM value of each amino acid residue was calculated based on the SCM value of each atom. Wherein, the calculation of the SCM value of each atom further includes: For each atom i, its SCM value is calculated by the following formula: in, Indicates the first j The charge value of an atom, Wherein, the calculation of the SCM value of each amino acid residue based on the SCM value of each atom further comprises: Based on the three-dimensional structure of the antibody, the SCM value of each amino acid residue of the antibody was calculated by the following formula as the extracted feature: in, Indicates the k The amino acid residues i The SCM value of atoms; The light and heavy chains of the antibody sequences were numbered using the ANARCI tool, respectively. The numbering scheme adopted Kabat numbering, and the numbering coordinates took the complete coordinates of the complete numbering system; The calculated SCM value of each amino acid residue is mapped to the numbered coordinate system of the antibody sequence. If there is no corresponding amino acid in the coordinate, it is padded with 0; if the calculated SCM value of the amino acid residue is greater than 0, its SCM value is replaced by 0.
2. The method according to claim 1, characterized in that The protein structure prediction tool is AlphaFold2, HelixFold, SWISS-MODEL or Uni-Fold.
3. The method according to claim 1, characterized in that The three-dimensional structure is described by PDB file and PQR file.
4. The method according to claim 3, characterized in that The three-dimensional structure of the antibody obtained by the modeling is described using a PDB file. The PDB file is converted into a PQR file using the PDB2PQR tool, and the charge value of each atom is extracted based on the PQR file.
5. The method according to claim 1, wherein The machine learning-based antibody viscosity prediction model is trained through the following steps: Acquiring a data set, the data set comprising a plurality of training samples, each of the training samples comprising a variable region amino acid sequence and a viscosity value of a sample antibody; Performing feature extraction on the variable region amino acid sequence of each training sample based on the atomic spatial charge distribution in the three-dimensional structure of the antibody; Using a machine learning method, cross-validation is performed on the data set to determine the machine learning model parameters; Using the machine learning model parameters, the machine learning model is trained on the data set to obtain a final antibody viscosity prediction model based on machine learning. The training uses the extracted features as input data of the machine learning model and uses the viscosity value of the antibody as output data of the machine learning model.
6. The method according to claim 5, characterized in that The cross validation on the data set includes: Cross validation was performed on the dataset using the leave-one-out method.
7. A machine learning-based antibody viscosity prediction model, wherein the prediction model is trained by the following steps: Acquiring a data set, the data set comprising a plurality of training samples, each of the training samples comprising a variable region amino acid sequence and a viscosity value of a sample antibody; Performing feature extraction on the variable region amino acid sequence of each training sample based on the atomic spatial charge distribution in the three-dimensional structure of the antibody; Using a machine learning method, cross-validation is performed on the data set to determine the machine learning model parameters; Using the machine learning model parameters, the machine learning model is trained on the data set to obtain a final antibody viscosity prediction model based on machine learning. The training uses the extracted features as input data of the machine learning model and the viscosity value of the antibody as output data of the machine learning model. Wherein, the machine learning method includes at least one of random forest and support vector machine regression, and The feature extraction based on the atomic space charge distribution in the three-dimensional structure of the antibody for the variable region amino acid sequence of each training sample includes: Using a protein structure prediction tool to model the variable region amino acid sequence of each training sample to obtain the three-dimensional structure of the sample antibody; Based on the three-dimensional structure of the sample antibody, the SCM value of the amino acid residues of the sample antibody is calculated as the extracted feature, Wherein, the calculation of the SCM value of the amino acid residues of the sample antibody based on the three-dimensional structure of the sample antibody as the extracted feature further includes: Extracting the charge value of each atom in the amino acid residues of the sample antibody based on the three-dimensional structure of the sample antibody; Calculate the SCM value of each atom; The SCM value of each amino acid residue was calculated based on the SCM value of each atom. Wherein, the calculation of the SCM value of each atom includes: For each atom i, its SCM value is calculated by the following formula: in, Indicates the first j The charge value of an atom, The step of extracting features of the variable region amino acid sequence of each training sample based on the atomic spatial charge distribution in the three-dimensional structure of the antibody further includes: Based on the three-dimensional structure of the sample antibody, the SCM value of each amino acid residue of the sample antibody was calculated by the following formula as the extracted feature: in, Indicates the k The amino acid residues i The SCM value of atoms; The light and heavy chains of the variable region amino acid sequences of the sample antibodies were numbered using the ANARCI tool, using the Kabat numbering scheme and the complete coordinates of the complete numbering system. The calculated SCM value of each amino acid residue is mapped to the numbered coordinate system of the sample antibody sequence. If there is no corresponding amino acid in the coordinate, it is padded with 0; if the calculated SCM value of the amino acid residue is greater than 0, its SCM value is replaced by 0.
8. The prediction model according to claim 7, characterized in that The cross validation on the data set includes: Cross validation was performed on the dataset using the leave-one-out method.
9. The prediction model according to claim 7, characterized in that The protein structure prediction tool is AlphaFold2, HelixFold, SWISS-MODEL or Uni-Fold.
10. The prediction model according to claim 7, characterized in that The three-dimensional structure is described by PDB file and PQR file.
11. The prediction model according to claim 10, characterized in that The three-dimensional structure of the sample antibody obtained by the modeling is described by a PDB file. The PDB file is converted into a PQR file using the PDB2PQR tool, and the charge value of each atom is extracted based on the PQR file.
12. A non-transitory computer-readable storage medium for storing a computer program, wherein the computer program comprises instructions, which, when executed by a processor of an electronic device, cause the electronic device to implement the antibody viscosity prediction method according to any one of claims 1 to 6.
13. A system for predicting antibody viscosity, comprising: processor; Memory; and A computer program, wherein the computer program is stored in the memory and configured to be executed by the processor, the computer program comprising instructions for implementing the antibody viscosity prediction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Computer-implemented methods of determining protein viscosity
US10762980B2
Computer-implemented methods of determining protein viscosity
US20160132634A1