Protein and ligand binding affinity prediction method, system, equipment and medium
By using a method of pre-trained prediction model based on data samples, combining olfactory protein-odor molecule experimental data and protein-ligand binding biological experimental database data, the problem of insufficient accuracy in the binding affinity prediction of olfactory receptors and odor small molecules in the prior art is solved, and a more efficient prediction effect is achieved.
Patent Information
- Application Number
- CN202510113368.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The existing virtual screening methods have insufficient accuracy and reliability in the prediction of binding affinity between olfactory receptors and odor small molecules, and cannot meet the needs of efficient screening.
A method of pre-trained prediction model based on data samples is provided, and binding affinity prediction is carried out by obtaining olfactory protein and odor molecules, combining olfactory protein-odor molecule experimental data, protein-ligand binding affinity biological experiment database data, and generated sludge data.
It improves the accuracy and efficiency of the prediction of binding affinity between proteins and ligands, and meets the needs of efficient screening, especially in the library of extremely high-diversity odor molecules.
Smart Images

Figure CN120032706A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of biological technology, and in particular to a method, system, device and medium for predicting the binding affinity between a protein and a ligand. Background Art
[0002] The combination of olfactory receptors and odor small molecules is the core mechanism for organisms to perceive external odors. The key lies in the complex and precise interaction between specific proteins and small molecules. At present, olfactory recognition research mainly relies on in vitro experiments, which analyze the characteristics of their interactions by measuring the binding affinity and functional effects of olfactory receptors and odor small molecules. Although these experiments have high accuracy, they are usually time-consuming and labor-intensive, and have huge resource requirements. They cannot meet the actual needs of large-scale and efficient screening, especially when exploring in a highly diverse odor molecule library.
[0003] In order to improve screening efficiency, virtual screening technology has gradually been applied. Through computational simulation, virtual screening can quickly screen potential binding molecules, significantly improve enrichment multiples and reduce the cost and difficulty of experimental screening. However, most of the existing virtual screening methods are general models designed for the universal binding rules between proteins and small molecules, and are not specifically optimized for the unique binding characteristics of olfactory receptors and odor small molecules. This general model has limited performance in the field of olfaction, and the accuracy and reliability of the prediction results are difficult to fully meet the needs of actual applications. Summary of the invention
[0004] Based on this, it is necessary to provide a method, system, device and medium for predicting the binding affinity between a protein and a ligand in response to the above-mentioned technical problems, which effectively improves the prediction accuracy and efficiency of the binding affinity between a protein and a ligand.
[0005] In a first aspect, a method for predicting the binding affinity between a protein and a ligand is provided, comprising:
[0006] Obtain olfactory proteins and odor molecules;
[0007] The olfactory protein and the odor molecule are input into a prediction model to obtain a prediction result of the combination of the olfactory protein and the odor molecule,
[0008] Among them, the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from the data source, protein-ligand binding biological experimental data that meets the similarity requirements with the olfactory protein sequence screened out from the protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data.
[0009] In some examples, the method for obtaining the olfactory protein-odor molecule experimental data in the data sample includes:
[0010] Obtaining an olfactory protein-odor molecule experimental data set from a data source, wherein the data source includes a public database and a non-public database;
[0011] Performing data standardization on each data in the olfactory protein-odor molecule experimental data set;
[0012] After the data standardization process, duplicate data and abnormal data in the olfactory protein-odor molecule experimental data set are deleted to obtain the olfactory protein-odor molecule experimental data.
[0013] In some examples, the method for obtaining protein-ligand binding biological experiment data that meets the similarity requirement with the olfactory protein sequence screened from the protein-ligand binding affinity biological experiment database includes:
[0014] Obtain protein-ligand binding affinity biological experiment database;
[0015] Calculating the similarity between the sequence of each data in the protein-ligand binding affinity biological experiment database and the olfactory protein sequence;
[0016] Data with a similarity greater than a preset value are screened out from the protein-ligand binding affinity biological experiment database to obtain the protein-ligand binding biological experiment data.
[0017] In some examples, the bait generated based on the positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data includes:
[0018] Based on the molecular fingerprint, preliminary screening molecules whose structural similarity with the active molecules meets the predetermined requirements are screened from the public database;
[0019] Eliminate molecules that are active in the target protein from the preliminary screening molecules, and for the remaining preliminary screening molecules, screen out candidate molecules whose molecular properties differ from those of the active molecules by more than a preset difference;
[0020] The candidate molecules are optimized, the optimized candidate molecules are ranked, and the decoys are screened out according to the ranking results.
[0021] In some examples, the step of optimizing the candidate molecules, ranking the optimized candidate molecules, and screening out the decoys according to the ranking results includes:
[0022] Optimizing the candidate molecules by introducing the following fixed-rule local structural modifications into the candidate molecules: replacing non-critical groups and modifying the positions of branch chains, wherein the treatment process maintains the chemical rationality of the perturbed molecules;
[0023] Normalizing each molecular term of the optimized candidate molecules;
[0024] Calculate the score of each molecular term of the candidate molecule according to the normalized result of each molecular term;
[0025] Determine the score of the candidate molecule according to the score of each molecular item of the candidate molecule;
[0026] Filter out the N candidate scores with the highest scores, where N is a positive integer;
[0027] Affinity assignment is performed on the N candidate molecules to obtain baits.
[0028] In some examples, this also includes:
[0029] Construct a protein map of the data sample, specifically:
[0030] Extracting protein features, wherein the protein features include node features, global features, local features and edge features, wherein the global features include high-dimensional global features extracted using the BERT model, and the local features include local spatial structural features and residue contact maps extracted using the ESM2 model;
[0031] Screening the physicochemical properties of amino acids to enhance the node characteristics of protein graphs;
[0032] Transmembrane region features enhance node properties of protein graphs.
[0033] In some examples, this also includes:
[0034] Construct a graph of molecules in the data sample, specifically:
[0035] Extracting molecular features, wherein the molecular features include atomic features and overall molecular characteristics, and fusing the atomic features and overall molecular characteristics.
[0036] In a second aspect, a system for predicting the binding affinity between a protein and a ligand is provided, comprising:
[0037] Acquisition module, used to obtain olfactory proteins and odor molecules;
[0038] A prediction module is used to input the olfactory protein and the odor molecule into a prediction model to obtain a prediction result of the binding of the olfactory protein and the odor molecule, wherein the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from a data source, protein-ligand binding biological experimental data that meets the similarity requirements with the olfactory protein sequence screened from a protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data.
[0039] In a third aspect, a computing device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for predicting the binding affinity between a protein and a ligand according to the first aspect above is implemented.
[0040] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for predicting the binding affinity between a protein and a ligand according to the first aspect above is implemented.
[0041] Using the embodiments of the present application, olfactory proteins and odor molecules are input into a prediction model to obtain a prediction result of the binding of the olfactory protein and the odor molecule. Since the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from the data source, protein-ligand binding biological experimental data that meet the similarity requirements with the olfactory protein sequence screened out from the protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data, the prediction accuracy and reliability of the prediction model after training are higher, thereby effectively improving the prediction accuracy and efficiency of the binding affinity between the protein and the ligand. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0043] Figure 1 A flowchart of a method for predicting the binding affinity between a protein and a ligand provided in an embodiment of the present application;
[0044] Figure 2 A diagram showing the effect of the binary classification model of an embodiment of the present application on the prediction of olfactory protein-small molecule binding probability in a test set;
[0045] Figure 3A structural block diagram of a protein-ligand binding affinity prediction system provided in an embodiment of the present application;
[0046] Figure 4 A structural block diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] The present application is further described in detail below in conjunction with the embodiments and drawings. It is to be understood that the specific embodiments described herein are only used to explain the relevant application, rather than to limit the application. It is also necessary to explain that, for ease of description, only the parts related to the application are shown in the drawings.
[0048] It should be noted that, in the absence of conflict, the embodiments of the present application, that is, the features of the embodiments, can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0049] The following is a detailed description of the protein-ligand binding affinity prediction method, system, device and medium according to the embodiments of the present application in conjunction with the accompanying drawings.
[0050] The implementation environment of the application embodiment can obtain olfactory proteins and odor molecules by personal computing devices, such as computers, mobile terminals, etc.; the olfactory proteins and the odor molecules are input into a prediction model to obtain a binding prediction result of the olfactory protein and the odor molecules, wherein the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from a data source, protein-ligand binding biological experimental data that meets the similarity requirements with the olfactory protein sequence screened from a protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data.
[0051] Alternatively, it can also be implemented by a server, for example: an individual's computing device sends a request to the server, and the server obtains the olfactory protein and the odor molecule; the olfactory protein and the odor molecule are input into a prediction model to obtain a binding prediction result of the olfactory protein and the odor molecule, wherein the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from a data source, protein-ligand binding biological experimental data that meets the similarity requirement with the olfactory protein sequence screened out from a protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data. Finally, the results are returned to the individual's computing device.
[0052] Among them, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.
[0053] Figure 1 FIG. 1 is a flow chart of a method for predicting the binding affinity between a protein and a ligand according to an embodiment of the present application. Figure 1 As shown, according to an embodiment of the present application, a method for predicting the binding affinity between a protein and a ligand comprises the following steps:
[0054] S101: Obtain olfactory proteins and odor molecules.
[0055] S102: Inputting the olfactory protein and the odor molecule into a prediction model to obtain a prediction result of the binding of the olfactory protein and the odor molecule, wherein the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from a data source, protein-ligand binding biological experimental data that meets the similarity requirement with the olfactory protein sequence screened out from a protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data.
[0056] Decoys refer to negative samples, for example, negative sample data with high similarity to positive samples, that is, data with similar structure or characteristics to positive samples but different properties or functions in model training and verification. For olfactory protein-odor molecule experimental data from data sources, the data source can be some public databases and non-public databases, such as internal databases.
[0057] Taking public databases and non-public databases as examples, the method for obtaining the olfactory protein-odor molecule experimental data in the data sample includes: obtaining the olfactory protein-odor molecule experimental data set in the data source, and the data source includes the public database and the non-public database; performing data standardization processing on each data in the olfactory protein-odor molecule experimental data set; after the data standardization processing, deleting the duplicate data and abnormal data in the olfactory protein-odor molecule experimental data set to obtain the olfactory protein-odor molecule experimental data.
[0058] Specifically, olfactory protein-odor molecule experimental data, namely olfactory protein experimental data, includes olfactory experimental data in public databases (such as Pubchem, Chembl, GPCRdb, M2OR, FlavorDB, etc.) and data in the company's internal database. Among them, the types of olfactory protein experimental data include: (a) Affinity quantitative data: Affinity quantitative data is used to predict binding affinity (regression model). (b) Active / inactive binary classification data: Active / inactive binary classification data is used to predict binding probability (binary classification model). The prediction model in the embodiments of the present application usually refers to a regression model.
[0059] After searching the experimental data of olfactory proteins and odor molecules from the data source, data preprocessing is required, that is, standardization of the data, experimentally measured activity values such as IC 50 or EC 50 , uniformly converted to -logIC 50 or -logEC 50 In addition, for entries that cannot provide experimental activity data, only their classification (active or inactive) information is retained.
[0060] Duplicate data removal: Delete duplicate olfactory receptor-odor molecule pairs, retain data with activity value error within ±0.5, and take the average value as the final experimental activity value.
[0061] Abnormal data processing: Delete experimental data with obvious measurement errors (such as negative values or affinity values exceeding the reasonable range or multiple experimental results with result deviations exceeding ±0.5), and filter invalid small molecule structures (such as isolated atoms, chemical bond errors).
[0062] For the expansion and enhancement of data, including GPCR data screening based on sequence similarity, specifically, the method for obtaining protein-ligand binding biological experiment data that meets the similarity requirements with the olfactory protein sequence screened out from the protein-ligand binding affinity biological experiment database includes: obtaining the protein-ligand binding affinity biological experiment database; calculating the similarity between the sequence of each data in the protein-ligand binding affinity biological experiment database and the olfactory protein sequence; screening out data with similarity greater than a preset value from the protein-ligand binding affinity biological experiment database to obtain the protein-ligand binding biological experiment data.
[0063] Specifically, since there are relatively few experimental data on olfactory receptors (i.e., experimental data on olfactory proteins), in the embodiments of the present application, the training data is expanded by a GPCR data screening strategy based on sequence similarity, for example, calculating the similarity between each sequence in the GPCR data set and the olfactory receptor sequence. In a specific example, the similarity threshold is set to 0.3, and then the GPCR protein sequence similar to the olfactory receptor sequence can be screened out according to the similarity threshold. The binding data of the screened GPCR protein and its corresponding small molecule are supplemented to ensure the biological relevance of the expanded data.
[0064] The expansion and enhancement of data also includes the supplement of bait data. Specifically, the bait generated according to the positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data includes: based on molecular fingerprints, preliminary screening molecules whose structural similarity with the active molecules meets predetermined requirements are screened from a public database; molecules that are active in the target protein are eliminated from the preliminary screening molecules, and for the remaining preliminary screening molecules, candidate molecules whose molecular properties differ from those of the active molecules by more than a preset difference are screened; the candidate molecules are optimized, the optimized candidate molecules are sorted, and the baits are screened out according to the sorting results.
[0065] In this example, the candidate molecules are optimized, the optimized candidate molecules are ranked, and decoys are screened out based on the ranking results, including: optimizing the candidate molecules, specifically introducing the following fixed-rule local structural modifications into the candidate molecules: replacing non-critical groups and modifying branch positions, wherein the processing maintains the chemical rationality of the disturbed molecules; normalizing each molecular item of the optimized candidate molecules; calculating the score of each molecular item of the candidate molecules based on the normalized result of each molecular item; determining the score of the candidate molecules based on the score of each molecular item of the candidate molecules; screening out the N candidate scores with the highest scores, wherein N is a positive integer; assigning affinity to the N candidate molecules to obtain decoys.
[0066] Specifically, the purpose of the design of the lure is to enhance the model's ability to distinguish complex situations and improve the robustness of the model. In the embodiment of the present application, the lure of negative samples is generated based on positive samples to further balance the distribution of active and inactive samples, and further enhance the model's ability to distinguish positive and negative samples. The generation process of the lure specifically adopts the following process:
[0067] Database screening: Molecules with similar structures to active molecules (for example, structural Tanimoto similarity ≥ 0.7) are preliminarily screened from public databases (such as PubChem, ChEMBL, and ZINC). The calculation is based on molecular fingerprints. When the target number, such as 50 molecules, is found, the calculation is stopped to save computing resources.
[0068] The molecular fingerprint may be a Morgan molecular fingerprint. Of course, in other examples, it may also be a Pubchem fingerprint, a Rdkit fingerprint, etc.
[0069] Functional property screening, based on known data, excludes small molecules that are active in the target protein. If there is activity data related to the target protein in the database (there is an IC 50 or EC 50 The remaining molecules were further screened, and only candidates whose molecular properties differed from the active molecules by more than the preset difference (absolute value of hydrophobicity (LogP) deviation>1, absolute value of polar surface area (PSA) deviation>20) were retained.
[0070] Data enhancement and optimization, in order to further improve the diversity and scientificity of the lures, the selected lure molecules are optimized: the following fixed rules of local structural modifications are introduced into the candidate lure molecules: replacing non-critical groups (such as replacing hydroxyl groups in active molecules with fluorine atoms), modifying the position of the side chain (such as changing the position of the atom connected to the side chain), and the processing process must ensure that the perturbed molecules are still chemically reasonable (checked by RDKit). If the optimized molecules still maintain high structural similarity (Tanimoto similarity ≥ 0.7) and meet the differences in physical and chemical properties (LogP, PSA).
[0071] Ranking of bait molecules and final output, all small molecules are scored and ranked, the value of each indicator is normalized to the interval [0,1], and then calculated according to the following formula. DiversityPenalty is a diversity penalty term, which encourages the model to output more diverse bait molecules. ij represents the Tanimoto similarity between bait molecules i and j (calculated by molecular fingerprint), and N represents the total number of bait molecules. Tanimoto represents structural similarity, ΔLogP represents the difference in hydrophobicity, and ΔPSA represents the difference in polar surface area. Finally, the molecules are sorted according to their scores, and the 20 molecules with the highest scores are retained as the final bait set.
[0072]
[0073]
[0074] Affinity value generation, based on the normal distribution method, randomly generates a floating point number between [1.5-3] as the affinity assignment of the bait data. The embodiment of the present application uses bait and GPCR supplementary data sets for prediction in the regression model (ie, prediction model) of affinity prediction.
[0075] In one embodiment of the present application, the method for predicting the binding affinity of a protein and a ligand also includes: constructing a protein graph in a data sample, specifically: extracting protein features, wherein the protein features include node features, global features, local features, and edge features, wherein the global features include high-dimensional global features extracted using the BERT model, and the local features include local spatial structural features and residue contact maps extracted using the ESM2 model; screening the physicochemical properties of amino acids to enhance the node characteristics of the protein graph; and enhancing the node characteristics of the protein graph using transmembrane region features.
[0076] Specifically, protein features are extracted: the protein sequence is processed into an amino acid graph, each amino acid is regarded as a node of the graph, and the edge is composed of the contact probability between amino acids. The following is a description of the node and edge characteristics of the protein graph:
[0077] (1) Node characteristics:
[0078] Amino acid type: Each amino acid node contains its type information, which can be represented by one-hot encoding.
[0079] Physicochemical properties: Physicochemical properties (including molecular weight, hydrophobicity, polarity, etc.) screened based on the random forest method. These properties are the core characteristics of nodes when constructing protein graphs and are used to capture important information related to olfactory receptor binding.
[0080] Transmembrane region characteristics: amino acids in the transmembrane region are predicted by prediction tools, such as TMHMM (TransMembrane prediction with Hidden Markov Models), TMHMM2.0, and Phobius. Transmembrane regions are usually composed of highly hydrophobic amino acids, so the hydrophobicity information of these amino acids and their positions in the transmembrane region are embedded in the node characteristics to further enhance the expressiveness of the features.
[0081] (2) Global features:
[0082] BERT features: The high-dimensional global features extracted by the BERT model are used to capture long-range dependencies in protein sequences.
[0083] ESM2 features: Provides detailed local structural representation using local spatial structural features and residue contact maps extracted by the ESM2 model.
[0084] (3) Edge characteristics:
[0085] Edges are generated based on the contact probabilities between amino acids, which are predicted using the contact graph from the ESM2 model.
[0086] Edge construction rules:
[0087] Edges are generated by amino acid pairs with a contact probability greater than a certain threshold (for example, set to 0.5), and the weight is the contact probability value. To prevent the appearance of isolated nodes, edges between adjacent amino acids in the sequence (directly connected 1st and 2nd order adjacent edges) are added to the initial graph (1st order adjacent edges: add edges for directly adjacent amino acids in the sequence (i.e., between the i-th and i+1th amino acids); 2nd order adjacent edges: add edges for nodes separated by one amino acid (i.e., between the i-th and i+2th amino acids), and the weights of these edges are set to the default value (e.g., 0.5).
[0088] After generating the initial edges, the edges in the graph are optimized. Edges that appear multiple times are merged into one edge, and their maximum weight is taken as the final edge weight. At the same time, the edges are deduplicated to ensure that the graph is undirected, that is, if there is an edge between node A and node B, the same edge also exists between node B and node A. Subsequently, node self-loops (edges between a node and itself) are removed from the graph to prevent redundant features. Finally, self-loops are added to all nodes (the weight of the self-loop is fixed to 1) to ensure that each node can connect to itself.
[0089] Enhance the protein graph node characteristics by screening the physicochemical properties of amino acids through machine learning:
[0090] (a) Preparation of amino acid physicochemical property dataset.
[0091] Data collection and feature selection, sorting out the known physical and chemical properties of amino acids, including but not limited to hydrophobicity, hydrophilicity, polarity, molecular weight, pKa value, charge distribution, etc., a total of more than 500 physical and chemical properties.
[0092] Data normalization,To ensure the robustness of the model and the comparability of the characteristic data, all physical and chemical properties are normalized.,The normalization method is implemented by the following formula,This processing ensures that all characteristics are compressed into the range of [0,1],which is convenient for subsequent analysis.
[0093]
[0094] To supplement unknown amino acid data, for unknown or unidentified amino acids (X), the average of the minimum and maximum values of known characteristics is used to fill them in, so as to avoid the influence of missing values on model training and prediction effects.
[0095] (b) Extract the amino acid characteristics of each protein.
[0096] Traverse each amino acid in the protein sequence (including olfactory receptor sequences and a certain number of non-GPCR protein sequences) and extract more than 500 corresponding physicochemical property values to form a property matrix. Map the corresponding physicochemical properties for each amino acid in each protein sequence. In order to eliminate the impact of differences in sequence length and amino acid distribution on classification, this study uses the weighted mean method to summarize the physicochemical properties of amino acids in each protein sequence:
[0097]
[0098] Where f(AA i ) is the frequency of amino acid i, p ij is the value of amino acid i at the jth property, and n is the number of amino acids in the sequence.
[0099] Extract protein sequences from the olfactory receptor protein dataset and ensure the uniqueness of the data after deduplication. Traverse the protein sequence and map more than 500 physicochemical property values for each amino acid in the sequence according to the physicochemical property table of amino acids. Average the property values of all amino acids in each sequence to generate the property vector of the entire sequence.
[0100] (c) Target variable generation.
[0101] Generate a target variable Y associated with the protein sequence for training the classification model. Olfactory receptor proteins are labeled as 1. Non-GPCR proteins are labeled as 0.
[0102] (d) Random forest model construction and training.
[0103] A random forest model was constructed using RandomForestRegressor, with the weighted average of the amino acid physicochemical properties (property matrix X) as input features. The model was trained with the protein type target variable Y as the target value to quantify the importance of each physicochemical property.
[0104] (e) Feature importance assessment.
[0105] After the model training is completed, the influence of each physicochemical property on the target variable is calculated and output through the feature importance mechanism of random forest. The feature screening threshold (0.002) is set to screen out the key feature set with the greatest influence on the target variable, and the screened key features are stored as independent files for subsequent virtual screening simulation of amino acid physicochemical property feature construction.
[0106] Transmembrane region features enhance protein graph node properties:
[0107] Based on the hidden Markov model, the amino acid sequence is used to predict whether the protein contains transmembrane regions and the location of these regions. After predicting the transmembrane region of the olfactory receptor using prediction tools (such as TMHMM, TMHMM2.0, and Phobius, etc.), the characteristic information of the transmembrane region can be added to each amino acid. If an amino acid is located in the transmembrane region, it is assigned a value of 1, otherwise it is 0.
[0108] Transmembrane regions are usually composed of highly hydrophobic amino acids, so the hydrophobicity of these amino acids and their transmembrane region positions can be used as node characteristics and added to the protein map. By combining the characteristics of the transmembrane region, the node characteristics of the protein map are further enriched, so that it can not only express the basic physical and chemical properties of amino acids, but also reflect their functional region information in proteins. In this way, the node characteristics in the protein map will be more comprehensive, and can simultaneously reflect the physical and chemical properties of amino acids and their role in the functional region of proteins, especially the relevant information of the key functional region of the transmembrane region.
[0109] This enhanced feature representation can help virtual screening models better capture the structure-function relationship of proteins, thereby improving the predictive performance of the models and providing more meaningful biological explanations for understanding protein function.
[0110] The global features of protein graphs are jointly extracted based on the BERT model and the ESM2 model.
[0111] In the embodiments of the present application, a combination of the BERT model and the ESM2 model is used to extract features from the protein graph. The BERT model and the ESM2 model are two deep learning models that focus on understanding protein sequences from different perspectives, giving full play to their respective advantages to generate a richer and more comprehensive protein feature representation.
[0112] (a) BERT model: The BERT (Bidirectional Encoder Representations from Transformers) model is a pre-trained language model based on the Transformer architecture, which can capture the complex dependencies between amino acid residues in protein sequences through the contextual self-attention mechanism. Through large-scale pre-training, the BERT model can effectively extract the long-range dependency information between amino acids in protein sequences. Using the BERT model to process protein sequences can obtain high-dimensional protein sequence embeddings that cover the global information and contextual features of the sequence, providing effective input for subsequent tasks such as protein function prediction and interaction modeling.
[0113] (b) ESM2 model: The ESM2 (Evolutionary Scale Modeling 2) model is a language model designed specifically for protein sequences. It can capture the local structural information of protein sequences through convolutional neural networks (CNN). Unlike the global contextual information focused on by the BERT model, ESM2 can deeply explore the local features of amino acid sequences, especially the spatial structure and residue contact relationships of proteins. By extracting structural information such as contact maps, ESM2 can predict the folding structure of proteins and provide structured representations for tasks such as protein-ligand interactions and protein function analysis.
[0114] (c) The global features of the BERT model are then combined with the local features of the ESM2 model, and the two features are merged along the feature dimension through feature concatenation (torch.cat).
[0115] Combined use of BERT and ESM2 models: In this application, the combined use of BERT and ESM2 models can comprehensively utilize the advantages of both in protein sequence feature extraction. The BERT model mainly provides in-depth semantic representation for protein sequences from the perspective of global context, while the ESM2 model provides more detailed spatial information from the perspective of local structure and residue contact. The features extracted by the two complement each other to a certain extent. The contextual information processed by BERT can be combined with the contact map and local structural information extracted by ESM2. The BERT model is used jointly to extract global features, and the local spatial structure features and residue contact map extracted by the ESM2 model are used. A more comprehensive and efficient protein graph feature representation is formed.
[0116] In one embodiment of the present application, the method for predicting the binding affinity between a protein and a ligand further includes: constructing a molecular graph in a data sample, specifically: extracting molecular features, wherein the molecular features include atomic features and overall molecular characteristics, and fusing the atomic features and overall molecular characteristics.
[0117] Specifically, extract small molecule features: The structure and properties of small molecules play a vital role in molecular function research and drug development. Traditional molecular graphs are usually composed of nodes (representing atoms) and edges (representing bonds), and only rely on local properties at the atomic level. This representation method ignores the important influence of the overall properties of the molecule (such as molecular weight, hydrophobicity, polarity, etc.) on the molecular function, resulting in limited accuracy in molecular function prediction and molecule-target affinity assessment. In order to improve the characterization capability of molecular graphs, the method of this application, through the use of RDKit and Torch related libraries, integrates the local atomic properties and overall properties of molecules into molecular graphs, which not only retains the details of the molecular structure, but also makes full use of the global information of the molecule, providing a more accurate graph data structure for downstream molecular function prediction.
[0118] A molecular graph construction method based on the fusion of small molecule atoms and overall properties.
[0119] (a) Atomic property extraction:
[0120] Atomic type, atomic number, chirality, bond connection degree (number of adjacent atoms), number of participating rings, formal charge, number of free radical electrons, electron orbital hybridization type, number of valence electrons, whether it is aromatic, whether it is in a ring, contribution of lipid-water partition coefficient and TPSA contribution, a total of 14 properties.
[0121] These properties reflect the local structure and function of atoms and provide the basis for the node properties of molecular graphs.
[0122] (b) Extraction of overall molecular characteristics:
[0123] Based on RDKIT, 210 properties were generated, including average molecular weight, absolute value of maximum partial charge, lipid-water partition coefficient, number of free radical electrons, etc. All global properties were normalized to ensure the comparability of property values and embedded in the molecular graph as shared properties.
[0124] (c) Fusion of atomic and global properties:
[0125] The innovation lies in assigning the overall characteristics of the molecule to each atomic node, so that each node in the molecular graph contains both the local characteristics of the atom and the shared characteristics of the overall properties of the molecule. This method makes up for the deficiency of the separation of local characteristics and global information in traditional molecular graphs, allowing the molecular graph to express the characteristics of the molecule more comprehensively.
[0126] (d) Edge characteristics and topological structure:
[0127] Edge characteristics: The edge characteristics of the molecular graph are constructed based on the bond type, including single bond, double bond, triple bond and aromatic bond information. Topological structure: The adjacency matrix of the molecular graph is generated according to the chemical bond relationship between atoms to ensure the accuracy of the topological relationship of the molecular graph.
[0128] The training and application of the prediction model are as follows:
[0129] Feature extraction and embedding:
[0130] The PNA-GNN (Principal Neighbourhood Aggregation Graph Neural Network) machine learning method is used for information transmission and aggregation, so that the characteristics of each node can be updated and optimized, thereby learning a more accurate embedding representation. This is crucial for subsequent protein-small molecule binding prediction.
[0131] The physicochemical graph convolution layer models the interaction between ligands and proteins through a graph neural network (GNN). The information transmission and node update involved in this process are based on physicochemical properties and structural properties. The specific process is as follows:
[0132] Characteristic transfer of ligands and proteins:
[0133] (a) Ligand atom groups and protein residue groups:
[0134] For small molecules, use "Junction Tree Decomposition" to group the atoms of the ligand into functional groups. Functional groups are grouped based on the similarity of chemical structure and chemical properties, and usually include similar chemical groups (such as benzene rings, amino acid residues, etc.). This method helps the model better understand the biological role of each part of the small molecule.
[0135] For proteins, the minCUT algorithm is used to cluster protein residues into protein regions. The residues in each protein region are connected by physicochemical properties (such as hydrophilicity, hydrophobicity, etc.) to avoid the invalid propagation of interaction information in proteins.
[0136] The core purpose of these steps is to pass the local and regional features of the ligand and protein to the next layer so that the interaction relationship can be captured more accurately. This method helps to reduce the computational complexity of the model and improve the accuracy of the model by constraining the interaction to occur only in functional areas.
[0137] (b) Convolutional layer operation:
[0138] After grouping the molecular graph and the residue graph, the interaction between the ligand and the protein is processed through graph convolution operations, and the interaction strength between the ligand and the protein region is calculated at the end of each layer. The use of the cross-attention mechanism in the process helps the model focus on the relationship between the ligand and the protein region, thereby further strengthening their interaction. Specifically, an attention weight is calculated between each atom of the ligand and each region of the protein, which reflects the strength of the interaction between the ligand atom and the protein region. In the calculation of each layer, the atomic features of the ligand and the regional features of the protein are transferred through attention weighting to obtain a more accurate node representation. A total of three layers of convolution operations are performed, as follows:
[0139] Step 1: In each layer, the interaction between the internal atoms of the ligand and the internal residues of the protein is first processed through message passing. Each layer updates the features of the nodes through a graph neural network (GNN).
[0140] Step 2: Physicochemical constraints were applied to the group features of ligands and proteins to optimize the model by focusing on the key contributions of each functional group and protein region.
[0141] Step 3: At the end of each layer, the interaction strength between the ligand and the protein region is calculated, and the interaction between the ligand sphere and the protein region is strengthened through a cross-attention model.
[0142] Residue and atom importance scores accumulate importance scores for each ligand atom and protein residue through three convolutional layers, which are calculated based on their contribution to the ligand-protein interaction. These importance scores provide interpretive information for downstream tasks. The contribution of each ligand atom and protein residue is reflected in the final interpretive fingerprint, through which key regions and functional groups of the interaction can be identified.
[0143] Calculation of importance scores,At each layer of the model, the importance scores are calculated as follows:
[0144] The importance of each node is evaluated based on the interaction relationship between the node and its neighboring nodes. If a node contributes more to the prediction of ligand-protein interaction, the weight of the node will be higher and the importance score will be higher.
[0145] The calculation of residue and atom importance scores usually relies on the similarity between node features and the features of neighboring nodes. By calculating the similarity of each node, the model can identify the most critical atoms and residues. Based on the importance score, an interpretative fingerprint can be generated to help researchers identify key protein regions and molecular functional regions.
[0146] Prediction,Using embedded features obtained through graph convolution and importance scoring, we can accurately predict the interaction properties between ligands and proteins.
[0147] Functional testing, olfactory protein-small molecule binding probability prediction (binary classification model):
[0148] The model of the embodiment of the present application is tested based on a self-organized olfactory experiment database (binary classification database), such as Figure 2 As shown, the protein sequence specificity accounts for the top 20% of all protein sequences in the entire data set, and the small molecule SMILES The data set with the top 20% of the specificity among all small molecule SMILES in the entire data set is used as the test set, and the changes in prediction accuracy under different classification thresholds are shown. When the classification threshold is set to 0.98, the model prediction accuracy reaches 0.49, and the recall is 0.25. The proportion of active data in the test set is 3%, which means that after virtual screening, the enrichment effect can be improved by more than 16 times.
[0149] Olfactory protein-small molecule binding affinity prediction (regression model):
[0150] The prediction model of the embodiment of the present application is tested based on a self-organized olfactory experiment database (experimental value database, including GPCR data and bait data collected based on sequence similarity), and a data set that satisfies both the protein sequence specificity of the top 20% of all protein sequences in all data sets and the small molecule SMILES specificity of the top 20% of all small molecule SMILES in all data sets is selected as the test set. The final test results are shown in Table 1:
[0151] Table 1
[0152] CI MAE MSE 0.8558 0.5736 0.6401
[0153] According to the method for predicting the binding affinity of a protein and a ligand in an embodiment of the present application, the olfactory protein and the odor molecule are input into the prediction model to obtain a prediction result of the binding of the olfactory protein and the odor molecule. Since the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from the data source, protein-ligand binding biological experimental data that meet the similarity requirements with the olfactory protein sequence screened out from the protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data, the prediction accuracy and reliability of the prediction model after training are higher, thereby effectively improving the prediction accuracy and efficiency of the binding affinity of the protein and the ligand.
[0154] That is, by screening GPCR data with similar sequences to olfactory receptors, training samples can be effectively expanded in the data-scarce olfactory field. In addition, the bait data generated for the characteristics of olfactory receptors can simulate and enhance the interactions between olfactory receptors and small molecules that are uncommon or understudied in nature. It not only improves the applicability of the model to olfactory receptors and the accuracy of predictions, but also enhances the robustness of the model in dealing with rare or complex olfactory molecules. By deeply analyzing the global and local characteristics of olfactory receptors, especially the physicochemical properties related to olfactory proteins and the uniqueness of their transmembrane regions, the key factors affecting olfactory signal transduction can be accurately captured. When predicting how olfactory receptors interact with various small molecules, it can provide higher accuracy and relevance, which is of great significance for analyzing the molecular mechanism of olfactory perception and developing drugs targeting specific olfactory pathways. By combining the global characteristics of small molecules and the local characteristics at the atomic level, a more detailed and comprehensive molecular representation can be provided. This fusion enables the model to accurately capture the key characteristics of small molecules when interacting with olfactory receptors, thereby improving the accuracy of binding affinity and functional effect predictions. After training the predictive model of the embodiments of the present application, the model can show extremely high practicality in olfactory science research and related applications, and is particularly suitable for rapid screening and evaluation of a large number of potential olfactory-affecting compounds, accelerating the development process of olfactory-related drugs and products.
[0155] Figure 3 is a structural block diagram of a protein-ligand binding affinity prediction system according to one embodiment of the present application. Figure 3 As shown, according to an embodiment of the present application, the protein-ligand binding affinity prediction system includes: an acquisition module 310 and a prediction module 320, wherein:
[0156] An acquisition module 310 is used to obtain olfactory proteins and odor molecules;
[0157] The prediction module 320 is used to input the olfactory protein and the odor molecule into a prediction model to obtain a prediction result of the binding of the olfactory protein and the odor molecule, wherein the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from a data source, protein-ligand binding biological experimental data that meets the similarity requirements with the olfactory protein sequence screened from a protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data.
[0158] According to the method for predicting the binding affinity of a protein and a ligand in an embodiment of the present application, the olfactory protein and the odor molecule are input into the prediction model to obtain a prediction result of the binding of the olfactory protein and the odor molecule. Since the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from the data source, protein-ligand binding biological experimental data that meet the similarity requirements with the olfactory protein sequence screened out from the protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data, the prediction accuracy and reliability of the prediction model after training are higher, thereby effectively improving the prediction accuracy and efficiency of the binding affinity of the protein and the ligand.
[0159] For the specific definition of the protein-ligand binding affinity prediction system, please refer to the definition of the protein-ligand binding affinity prediction method above, which will not be repeated here. The various modules of the above-mentioned protein-ligand binding affinity prediction system can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned various modules can be embedded in or independent of the processor in the computing device in the form of hardware, or can be stored in the memory of the computing device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned various modules.
[0160] Reference below Figure 4 , Figure 4 A schematic diagram of the structure of a computing device suitable for implementing an embodiment of the present application is shown.
[0161] like Figure 4As shown, computing device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage part 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation instructions of the system are also stored in RAM 1003. CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.
[0162] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed, so that a computer program read therefrom is installed into the storage section 1008 as needed.
[0163] In particular, according to an embodiment of the present application, the above reference flow chart Figure 1 The described process can be implemented as a computer software program. For example, a computer program product is included in an embodiment of the present application, which includes a computer program carried on a computer readable medium, and the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the above-mentioned functions defined in the system of the present application are executed.
[0164] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium such as a computer-readable storage medium that can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0165] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operating instructions of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the aforementioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the box can also occur in a sequence different from that marked in the above-mentioned accompanying drawings. For example, the boxes represented by two connections can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operating instruction, or can be implemented with a combination of dedicated hardware and computer instructions.
[0166] The units or modules involved in the embodiments described in the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor. The names of these units or modules do not, in some cases, constitute limitations on the units or modules themselves.
[0167] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0168] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent application shall be subject to the attached claims.
Claims
1. A method for predicting the binding affinity between a protein and a ligand, characterized in that: include: Obtain olfactory proteins and odor molecules; The olfactory protein and the odor molecule are input into a prediction model to obtain a prediction result of the combination of the olfactory protein and the odor molecule, Among them, the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from the data source, protein-ligand binding biological experimental data that meets the similarity requirements with the olfactory protein sequence screened out from the protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data.
2. The method for predicting the binding affinity between a protein and a ligand according to claim 1, characterized in that: The method for obtaining the olfactory protein-odor molecule experimental data in the data sample includes: Obtaining an olfactory protein-odor molecule experimental data set from a data source, wherein the data source includes a public database and a non-public database; Performing data standardization on each data in the olfactory protein-odor molecule experimental data set; After the data standardization process, duplicate data and abnormal data in the olfactory protein-odor molecule experimental data set are deleted to obtain the olfactory protein-odor molecule experimental data.
3. The method for predicting the binding affinity between a protein and a ligand according to claim 1, characterized in that: Methods for obtaining protein-ligand binding biological experiment data that meet the similarity requirements with the olfactory protein sequence screened from the protein-ligand binding affinity biological experiment database include: Obtain protein-ligand binding affinity biological experiment database; Calculating the similarity between the sequence of each data in the protein-ligand binding affinity biological experiment database and the olfactory protein sequence; Data with a similarity greater than a preset value are screened out from the protein-ligand binding affinity biological experiment database to obtain the protein-ligand binding biological experiment data.
4. The method for predicting the binding affinity between a protein and a ligand according to claim 1, characterized in that: The bait generated according to the positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data includes: Based on the molecular fingerprint, preliminary screening molecules whose structural similarity with the active molecules meets the predetermined requirements are screened from the public database; Eliminate molecules that are active in the target protein from the preliminary screening molecules, and for the remaining preliminary screening molecules, screen out candidate molecules whose molecular properties differ from those of the active molecules by more than a preset difference; The candidate molecules are optimized, the optimized candidate molecules are ranked, and the decoys are screened out according to the ranking results.
5. The method for predicting the binding affinity between a protein and a ligand according to claim 4, characterized in that: The step of optimizing the candidate molecules, ranking the optimized candidate molecules, and selecting the decoys according to the ranking results comprises: Optimizing the candidate molecules by introducing the following fixed-rule local structural modifications into the candidate molecules: replacing non-critical groups and modifying the positions of branch chains, wherein the treatment process maintains the chemical rationality of the perturbed molecules; Normalizing each molecular term of the optimized candidate molecules; Calculate the score of each molecular term of the candidate molecule according to the normalized result of each molecular term; Determine the score of the candidate molecule according to the score of each molecular item of the candidate molecule; Filter out the N candidate scores with the highest scores, where N is a positive integer; Affinity assignment is performed on the N candidate molecules to obtain baits.
6. The method for predicting the binding affinity between a protein and a ligand according to claim 1, characterized in that: Also includes: Construct a protein map of the data sample, specifically: Extracting protein features, wherein the protein features include node features, global features, local features and edge features, wherein the global features include high-dimensional global features extracted using the BERT model, and the local features include local spatial structural features and residue contact maps extracted using the ESM2 model; Screening the physicochemical properties of amino acids to enhance the node characteristics of protein graphs; Transmembrane region features enhance node properties of protein graphs.
7. The method for predicting the binding affinity between a protein and a ligand according to claim 1, characterized in that: Also includes: Construct a graph of molecules in the data sample, specifically: Extracting molecular features, wherein the molecular features include atomic features and overall molecular characteristics, and fusing the atomic features and overall molecular characteristics.
8. A protein-ligand binding affinity prediction system, characterized in that: include: Acquisition module, used to obtain olfactory proteins and odor molecules; A prediction module is used to input the olfactory protein and the odor molecule into a prediction model to obtain a prediction result of the binding of the olfactory protein and the odor molecule, wherein the prediction model is pre-trained based on data samples, and the data samples include olfactory protein-odor molecule experimental data from a data source, protein-ligand binding biological experimental data that meets the similarity requirements with the olfactory protein sequence screened from a protein-ligand binding affinity biological experimental database, and lures generated based on positive samples in the olfactory protein-odor molecule experimental data and the protein-ligand binding biological experimental data.
9. A computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for predicting the binding affinity between a protein and a ligand according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting the binding affinity between a protein and a ligand according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method and equipment for establishing target protein acidity coefficient prediction model, medium and program product
CN120853685A
A method, device, medium and program product for establishing a target protein acidity coefficient prediction model
CN120853685B