An AI-based method for predicting drug-protein function regulation

By constructing a prior knowledge base for drug-protein functional regulation and a multimodal feature fusion graph neural network, the problem of determining the direction of drug functional regulation in drug-target interaction prediction is solved. This enables direct determination of drug-induced protein activation or inhibition effects and system-level pharmacodynamic simulation, improving the accuracy of drug screening and biological interpretation capabilities.

CN122493974APending Publication Date: 2026-07-31SHANGHAI PUDONG HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI PUDONG HOSPITAL
Filing Date
2026-04-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing drug-target interaction prediction technologies are insufficient to directly infer the direction of drug regulation of protein function and lack the ability to predict drug-induced protein activation or inhibition effects, making it difficult to meet the needs of drug screening and mechanism of action research for complex diseases.

Method used

By constructing a prior knowledge base of drug-protein functional regulation, combining drug molecular structure information, protein sequence characteristics, and protein interaction network topology, and utilizing multimodal feature fusion and graph neural networks, the direction of drug-induced protein functional regulation can be predicted, and the overall role of drugs in biological systems can be evaluated through network propagation analysis.

Benefits of technology

This technology enables direct determination of drug-induced protein activation or inhibition effects, improves the accuracy of drug screening and the ability to simulate drug action patterns in biological systems, reduces false positive rates, and optimizes experimental resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493974A_ABST
    Figure CN122493974A_ABST
Patent Text Reader

Abstract

This invention discloses an artificial intelligence-based method for predicting drug-protein functional regulation, with the following specific steps: Step 1, constructing a prior knowledge base for protein functional regulation; Step 2, multimodal biofeedback extraction and fusion; feature fusion of multimodal data, including structural features of drug nodes, structural features of protein nodes, protein interactions, and drug-protein functional interaction binary network features; Step 3, constructing a deep learning-based functional regulation prediction model; Step 4, constructing an automated prediction system based on an optimal model engine and evaluating its multidimensional reliability; by integrating drug molecular structure information, protein sequence features, and protein interaction network topology information, the method predicts the direction of drug-induced protein functional regulation, and evaluates the overall role of the drug in the biological system through network propagation analysis, thereby providing an effective technical means for screening candidate drugs for complex diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of bioinformatics and artificial intelligence, specifically involving a drug-protein function regulation prediction method based on multi-source biological information fusion. It can be used to predict the activation or inhibition of target proteins by drugs and can be applied to drug screening and mechanism of action research. Background Technology

[0002] 1. Relevant content of existing technology

[0003] Drug reuse provides an efficient pathway for therapeutic discovery through compounds with known clinical safety profiles. In this process, computational methods for the interaction between small drug molecules and their targets have become a core tool for large-scale, high-throughput screening of candidate drugs. These methods primarily determine the ability of small molecules to interact with proteins based on the binding free energy of the small molecule's structure and the functional structural regions of the protein. Currently, multimodal integration and deep learning-based methods (such as DeepDTA, DrugCLIP, and DeepPurpose) have significantly improved the predictive power of drug-target interactions (DTI), enabling more accurate prediction of binding affinity between drugs and proteins, or identification of ligand binding modes. These models typically integrate the molecular structural features of the drug (e.g., simplified linear input canonical system (SMILES)) and the sequence information of the protein (e.g., amino acid sequences), as well as biological information such as protein-protein interaction networks, for end-to-end interactive inference. Most DTI models primarily focus on predicting the binding affinity between drugs and proteins. DeepDTA Co-VAE and GraphCL-DTA Representative methods, including Drug CLIP, infer interaction strength from molecular and protein characterization. Recent studies, including DeepPurpose, have applied contrastive learning to achieve genome-wide virtual screening. However, binding affinity does not necessarily translate into functional regulation, and drug-protein pairwise interactions alone cannot explain the complex mechanisms underlying pharmacological action. and MolTrans Various multimodal models, including those for protein sequencing and molecular structure, integrate heterogeneous features. However, these methods typically treat proteins as isolated entities, rarely considering their interactions within cellular networks. .

[0004] 2. Problems and shortcomings of existing technologies

[0005] Despite the progress made in affinity prediction by existing DTI techniques, key shortcomings remain: (1) Existing models lack identification of the direction of drug action on proteins: Firstly, DTI itself cannot resolve the functional directionality of drugs. Most structure- or sequence-based methods only focus on whether the drug binds to the target, but cannot systematically resolve the functional consequences after ligand binding. In actual pharmacology, binding is not equivalent to regulation, and existing techniques cannot distinguish whether a compound has an "excitatory (activating)" or "antagonistic (inhibiting)" effect on the target. (2) Existing models lack integration of complex protein interactions and pathway transmission: Most existing models treat the function of proteins as isolated entities during drug screening, failing to fully consider the complex protein-protein interaction (PPI) network context within cells. (3) Existing models lack drug screening strategies for multi-target diseases: Due to the inability to identify the functional regulatory direction of small drug molecules, traditional DTI models have limitations in explaining multi-target pharmacology and complex systemic drug effects, resulting in limited accuracy in drug reuse.

[0006] 3. Analysis of the reasons for the above problems

[0007] The root of the above problems lies in the complexity of protein function realization, which is mainly reflected in three aspects: (1) The complexity of functional coupling: Proteins have the complexity of functional coupling, and their function originates from the subtle coupling between structure, regulatory elements and conformational dynamics. Simple binding prediction models lack simulation of the dynamic regulatory process of proteins, so they cannot infer the direction of pharmacological regulation. In addition, many proteins have contradictory functional regions, such as the self-inhibitory region of serine / threonine kinase (RAF1) protein and the activation region of receptor binding domain (RBD), which perform completely opposite functions. Blindly performing DTI interaction prediction will lead to a significant deviation in the functional direction of the drug. (2) The network characteristics of signal propagation: The signal propagation of protein function has network characteristics. Drug-induced perturbations often propagate through PPI and biological pathways, generating downstream regulatory cascade reactions. If the network topology context is not introduced, the model will lose key biological background information, resulting in poor robustness when dealing with targets embedded in complex regulatory networks. (3) Semantic gaps in training data: Most existing public databases and training sets focus on data such as energy affinity, equilibrium dissociation constant (KD), and half-inhibitory concentration (IC50), lacking the integration and utilization of large-scale, standardized evidence on the direction of drug-protein function regulation (activation / inhibition).

[0008] 4. References

[0009] 【1】ÖZTÜRK H, ÖZGÜR A, OZKIRIMLI E. DeepDTA: deep drug-target bindingaffinity prediction [J]. Bioinformatics, 2018, 34(17): i821-i9.

[0010] Translation: ÖZTÜRK H, ÖZGÜR A, OZKIRIMLI E. DeepDTA: Prediction of deep drug target binding affinity[J]. Bioinformatics, 2018, 34(17): i821-i9.

[0011] 【2】LI T, ZHAO XM, LI L. Co-VAE: Drug-Target Binding AffinityPrediction by Co-Regularized Variational Autoencoders [J]. IEEE Trans PatternAnal Mach Intell, 2022, 44(12): 8861-73.

[0012] Translation: LI T, ZHAO XM, LI L. Co-VAE: Drug-target binding affinity prediction based on co-canonical variational autoencoder [J]. Acta Physico-Chimica Sinica, 2017, 35(10): 1003-1007. IEEE Transmission Mode Analysis Mach Intelligence, 2022, 44(12): 8861-73.

[0013] 【3】YANG

[0014] Translation: Yang Xiao, Yang Gang, Chu Jie. GraphCL-DTA: Graph contrastive learning based on molecular semantics for predicting drug target binding affinity[J]. Acta Physico-Chimica Sinica, 2017, 35(1): 107-100. IEEE J Biomed HealthInform, 2024, 28(8): 4544-52.

[0015] 【4】JIA Y, GAO B, TAN J, et al. Deep contrastive learning enables genome-wide virtual screening [J]. Science, 2026, 391(6781): eads9530.

[0016] Translation: Jia Y, Gao B, Tan J, et al. Deep contrastive learning enables genome-wide virtual screening [J]. Science, 2026, 391(6781): eads9530.

[0017] 【5】HUANG K, FU T, GLASS LM, et al. DeepPurpose: a deep learning library for drug-target interaction prediction [J]. Bioinformatics, 2021, 36(22-23): 5545-7.

[0018] Translation: Huang Ke, Fu Tao, GLASS LM, et al. DeepPurpose: A deep learning library for predicting drug-target interactions [J]. Bioinformatics, 2021, 36(22-23): 5545-7.

[0019] 【6】HUANG K, XIAO C, GLASS LM, et al. MolTrans: Molecular InteractionTransformer for drug-target interaction prediction [J]. Bioinformatics, 2021,37(6): 830-6.

[0020] Translation: Huang K, Xiao C, Glass LM, et al. MolTrans: Molecular interaction transformer for predicting drug-target interactions [J]. Bioinformatics, 2021, 37(6): 830-6.

[0021] 【7】NADA H, CHOI Y, KIM S, et al. New insights into protein-protein interaction modulators in drug discovery and therapeutic advance [J]. SignalTransduct Target Ther, 2024, 9(1): 341.

[0022] Translation: NADA H, CHOI Y, KIM S, et al. New insights into protein-protein interaction modulators in drug discovery and therapeutic progress [J]. Signal Transduction Targets, 2024, 9(1): 341. Summary of the Invention

[0023] This invention aims to address the problem that existing drug-target interaction prediction techniques are insufficient to directly infer the direction of drug-induced protein functional regulation. Existing methods typically only assess the binding affinity between drugs and proteins, lacking the ability to predict drug-induced protein activation or inhibition effects, thus failing to meet the needs of drug screening and mechanism-of-action research for complex diseases. Therefore, this invention provides an artificial intelligence-based drug-protein functional regulation prediction method. This method, based on multimodal feature fusion and graph neural networks, integrates drug molecular structure information, protein sequence characteristics, and protein interaction network topology information to predict the direction of drug-induced protein functional regulation. Furthermore, network propagation analysis is used to evaluate the overall effect of the drug in the biological system, thereby providing an effective technical means for screening candidate drugs for complex diseases.

[0024] The technical solution of this invention is:

[0025] 1. Overview of Technical Solution

[0026] This invention proposes an artificial intelligence-based method for predicting drug-protein functional regulation. By constructing a prior knowledge base of drug-protein functional regulation and combining drug molecular structure information, protein sequence characteristics, and protein interaction network topology, the method can predict drug-induced protein activation or inhibition effects.

[0027] 2. Detailed operating procedures and technical details

[0028] Core Concept Layer: Artificial Intelligence Predicts Drugs → Protein Agonists / Inhibitors

[0029] Algorithm Implementation Layer: Multimodal Features + Artificial Intelligence Model

[0030] Systems pharmacology layer: protein-protein interaction network / pathway network propagation + systemic scoring

[0031] This invention proposes an artificial intelligence-based drug-protein function regulation prediction method, which deeply integrates a core concept layer (artificial intelligence prediction), an algorithm implementation layer (multimodal features + artificial intelligence model), and a systems pharmacology layer (network propagation + system scoring). Its core processing flow is as follows:

[0032] Step 1: Construct a prior knowledge base (dependent variable matrix) for protein function regulation

[0033] This step aims to transform unstructured biomedical literature descriptions into structured tags with clear direction. .

[0034] First, approximately 3 million literature records containing information on drug-protein functional interactions were collected from publicly available biomedical literature databases. The system uses a large language model to identify keywords and contextual semantic analysis to determine the drug-protein regulatory relationship, categorizing these relationships into activation and inhibition. For complex nested sentence structures (such as "Drug A affects the phosphorylation of protein C through protein B"), which exhibit multiple semantic flips, a stack data structure is used to segment semantic blocks, determining the direction of drug A's influence on protein C as the model's outcome variable. Technical details include: performing a stack push operation when encountering logical introductory words or parentheses, and popping from the stack after parsing to calculate local polarity. Subsequently, symbolic valence inference and causal modeling are performed to calculate the symbolic direction of each semantic block. :

[0035]

[0036] in For positive verbs, add 1; for negative verbs, subtract 1. Then, determine the direction using multiplication. In direction determination, there is a bidirectional interaction between the drug and protein, where the pathway of the drug's influence on the protein is assigned a value. Conversely, the pathway (such as protein-driven drug metabolism) is assigned a value. Finally, the final label is generated by combining the sign direction (positive / negative) and the position coefficient. :

[0037]

[0038] in, This is a sign function used to extract the sign of a numerical value. If the result within the parentheses is greater than 0, then... (Indicates an excitatory effect), and secondly, if the calculation result within the parentheses is less than 0, then... (Indicating an inhibitory effect). B represents the total number of semantic blocks obtained after semantic segmentation of the above complex nested sentence structure. This includes the number of local blocks (B) after the text has been separated into specific semantic blocks (e.g., semantic segmentation using a stack structure). b is the counting variable for accumulating symbols, representing the b-th semantic block currently being calculated. b = 1 in the formula indicates that the inference results of all symbols from the 1st to the Bth semantic blocks are summed. Through the above process, each (3M) unstructured description in the original literature is transformed into a drug-protein function regulation tag with a clear direction. This allows for the construction of a structured prior knowledge base.

[0039] Step 2: Multimodal biometric feature extraction and fusion (independent variable matrix)

[0040] This section mainly focuses on feature fusion of multimodal data, including structural features of protein and drug nodes, protein interactions, and drug-protein functional interaction binary network features.

[0041] (1) Drug molecule fingerprint feature encoding (obtaining the structural features of drug nodes)

[0042] In obtaining the functional regulation label ( Following the matrix, this invention further extracts multimodal feature information of drugs and proteins. This step integrates a feature matrix that provides high-dimensional functional determination. Drug molecular fingerprint characterization utilizes the molecular structure information dimensionality reduction tool RDKit (open-source cheminformatics software), which describes the molecular spatial structure according to SMILES (Simplified Linear Input Canonical System) and calculates the feature quantity of each molecular fingerprint.

[0043]

[0044] in It is the current atom. They are adjacent atoms. It is an atom exist The next iteration (atomic radius is) ) eigenvectors at that time It is with atoms The set of directly connected neighboring atoms, It is an atom and Information on the types of chemical bonds between them. It is a preset hash mapping operator used to perform a nonlinear mapping between the current atomic feature and the neighboring atomic features to obtain a feature vector of fixed length. Represents atoms In the The feature representation at the next iteration encapsulates the topological distance as The local chemical environment within the range. Finally, the fingerprint of the entire molecule. It is a "folded" mapping of the feature sets generated by all atoms at all iteration levels:

[0045]

[0046] This represents traversing each atom in the molecular graph G (each SMILES string). , Recording atoms From itself to a radius of All structural topology information in the environment. This represents the merging of all substructure features generated by all atoms in the total molecular diagram G at different radii. It involves performing a folding operation to compress the feature identifier to a fixed length. This is a condition that Fold mapping must satisfy, namely, mapping all extracted features of the molecular graph G into a 2048-dimensional vector to capture the connectivity information of molecular substructures. A complex drug molecule may generate thousands of different local substructure features, and the number of features varies from molecule to molecule. However, computer models (such as GNN or MLP) require that the length of the input vector must be fixed (e.g., 2048).

[0047] (2) Protein sequence deep embedding characterization (obtaining the structural features of protein nodes)

[0048] Protein sequence deep embedding utilizes a pre-trained protein language model (ESM2) to extract semantic representations of the sequence and constructs fixed-length deep embedding vectors for the protein sequence. :

[0049]

[0050] Where L is the sequence length and i is the position index of the residue in the sequence. This represents the specific amino acid residue at the i-th position in the sequence. It is a pre-trained ESM2 (Transformer) encoder. The global average pooling operation maps protein sequences of different lengths to a fixed-dimensional feature space (640 dimensions) by averaging the hidden states of all residues.

[0051] (3) Message-passing-based protein system embedding (acquiring the characteristics of protein-protein interaction and drug-protein functional interaction binary network)

[0052] Then the generated sequence features are... As nodes, protein system embedding vectors are constructed using a protein-protein interaction network (PPI) graph (the graph is based on interaction data with moderate confidence (Score > 400) extracted from the Protein Interaction Functional Network Database (STRING)). The graph neural network operator is used to encode the topological location of nodes (proteins) in a biological system, and the initial state features of nodes in the PPI network are also discussed. ,Right now = Construct protein system embedding vectors that incorporate network topology information. :

[0053]

[0054] Here, A represents the adjacency matrix of the protein-protein interaction (PPI) network. Its elements represent whether there is an interaction between protein nodes. GNN stands for Graph Neural Network Operator. In this invention, it specifically refers to GATConv (graph attention operator with multi-head attention) or SAGEConv (neighborhood sampling-based operator), used to perform feature aggregation between nodes. Finally, the heterogeneous features of the above three dimensions are mapped to a unified representation space to form a multimodal feature matrix.

[0055] (4) Attention deep fusion to form a multimodal heterogeneous feature matrix (multimodal data feature fusion)

[0056] Finally, the heterogeneous features across the three dimensions are deeply interacted to output the final predicted probability. First, to input the attention mechanism, the system treats the aligned three types of features as a "feature sequence":

[0057]

[0058] in It's a tensor splicing operation. It is the generated input feature matrix, with 3 rows (corresponding to the three modalities) and the number of columns being the hidden layer dimension. It is a vector after mapping, which maps all features to the same feature dimension. This means mapping the three features to the same real number space. ,in This represents the hidden layer dimension, which is set to 512 here. Then, the association weights between the hidden layer and each dimension of the protein are calculated based on multimodal self-attention:

[0059]

[0060] in, It is a randomly initialized 512 The system employs a 512-parameter learnable matrix. Its core function is to map the input modalities (drug, protein, network) to the query space. During model training, the system utilizes backpropagation to automatically adjust the model based on the deviation between the predicted label Y and the true label. The weight values ​​in the equation. This iterative optimization makes... It has gradually evolved from an initial random state into a mature feature extractor capable of accurately identifying cross-spatial correlations between different modalities. When Multiply After that, the generated Matrix (3) 512) represents the retrieval vector for three modalities. For example, The representative drug is based on the "matching requirement" between its chemical structure and protein modality, such as whether the presence of a benzene ring in the drug is conducive to protein binding; This represents a "recognition request" made by a protein sequence to a drug based on its functional domains. Based on the protein's location, what kind of surrounding network information is needed to increase the success rate of Y inference? It is a randomly initialized 512 The system uses a 512-learnable parameter matrix. Its core function is to map input modalities (drugs, proteins, networks) to a key space. During model training, the system automatically adjusts the weights based on the prediction error using the backpropagation algorithm. It gradually evolved into an indexer capable of extracting core recognition features. When Multiply After that, the generated Matrix (3) 512) represents the identification label vectors for the three modalities. It determines what information the modality can provide for retrieval by other modalities. (Drug label): Represents the physical properties that a drug exhibits to proteins based on its molecular backbone, such as: "I am a hydrogen bond donor with a hydrophobic surface and specific locations." (Protein tag): Represents the matching ability of a protein sequence based on its functional domains, such as: "I am a protein with a hydrophobic cavity that can accommodate an aromatic ring." (Network tag): Represents the functional properties of the protein in the interaction network, such as: "I am a hub node for signal transduction, responsible for transmitting inhibitory signals". It is a randomly initialized 512 A 512-learnable parameter matrix. Its core function is to map the original modal features to the value space. Through training, We learned how to extract the most valuable substantive features for determining drug efficacy from massive amounts of biological information. Multiply After that, the generated Matrix (3) 512) represents the essential content vectors of the three modalities. It is the "essence of knowledge" that the model ultimately extracts and integrates. (Drug essence): Represents the most critical molecular fragment information for determining the direction of activation / inhibition. (Protein Essence): Represents the active center characteristics of receptor proteins that are actually involved in the drug response. (Network Essence): This represents key topological data that can quantify drug delivery pathways. Therefore, The representative asked, "Who can meet my biochemical needs?" (Tag): Answer "I have these characteristics, see if they match up." (Key Takeaway): The conclusion is, "If it matches, then take this core part and use it to predict drug efficacy." In this way, the model can automatically identify the key features that contribute most to the determination of drug efficacy using a self-attention mechanism.

[0061]

[0062] in It is the final comprehensive representation, which, through a self-attention mechanism, enables the model to learn to automatically focus on the most critical information sources when faced with different drug-protein pairs. The weighted average, with weights from and The degree of matching determines this. Calculate the correlation score matrix among the three types of features (e.g., calculate the matching degree between drug fingerprint and network topology location). This is a scaling factor used to prevent the gradient from vanishing due to excessively large dot products. For the dimension of attention head. It is a normalization function that converts the relevance score into weight coefficients (summing up to 1).

[0063] Finally, the system reassembles the attention-weighted sequence features into a fixed-length fusion vector.

[0064]

[0065] in, This represents global average pooling. It extracts the most representative overall pharmacodynamic characterization by averaging the attention outputs of the three modalities. The original input... With attention output Add them together and retain the underlying feature information for residual connection. Representative layer normalization ensures that the fused feature vectors The distribution is statistically stable, which helps the subsequent classifier determine the activation / inhibition direction.

[0066] Step 3: Construct a deep learning-based functional regulation prediction model

[0067] The functional control tags obtained in step 1 are used as monitoring information. And use the multimodal features obtained in step 2 as the model input ( A drug-protein function regulation prediction model was constructed. Before model training, the system first performed Z-score normalization on the feature matrix to eliminate dimensional differences between features of different modalities.

[0068]

[0069] in This represents the standardized feature vector. It is the original multimodal feature input. It is the arithmetic mean of the features in the training set. It is the standard deviation of the features in the training set.

[0070] (1) Integration model construction and hyperparameter optimization

[0071] The dataset was divided into training, test, and independent validation sets in an 80:15:15 ratio. Stratified 5-Fold Cross-Validation (CV) was employed to ensure that the ratio of "activation" (Y=1) to "inhibition" (Y=0) labels in each fold was consistent with the overall distribution, thus improving the model's robustness on imbalanced data. Parallel training included four heterogeneous classification algorithms: Random Forest (RF), XGBoost, LightGBM, and Logistic Regression (LR). The system evaluated each model's precision and recall on the validation set. Precision reflects the proportion of samples predicted as "activators" (Y=1) that are actually active.

[0072]

[0073] The models are TP (True Positive, a sample predicted as 1 that is actually 1), FP (False Positive, a sample predicted as 1 that is actually 0), TN (True Negative), and FN (False Negative). Recall reflects the proportion of all actually active samples that were successfully retrieved by the model.

[0074]

[0075] To evaluate the model's classification robustness on the independent validation set, the F1-Score is obtained based on the harmonic mean of precision and recall. This metric comprehensively measures the model's discriminative performance under uneven distribution of positive and negative samples, and is calculated as follows:

[0076]

[0077] Subsequently, ROC and AUC are the core statistical indicators for evaluating the model's "excitatory" and "inhibitory" efficacy. The ROC curve is based on... The false positive rate (FPR) is plotted by moving the classification threshold, depicting the dynamic performance of the model's discrimination. The FPR calculation formula is as follows:

[0078]

[0079] AUC, or Area Under the ROC Curve, is used to quantify a model's ability to distinguish between different classes. Its value ranges from 0.5 to 1; the closer to 1, the more accurate the model's discrimination. Generally, a value above 0.85 is considered excellent.

[0080]

[0081] It is the sum of the ranks of all positive samples (activation classes) in the predicted probability ranking. , These represent the number of positive and negative samples, respectively. The essence of AUC is the probability that the model will rank a positive sample before a negative sample when randomly selecting one positive and one negative sample. Based on AUC, F1-Score, and Recall performance, the best classifier is automatically determined or a weighted ensemble is performed (a value greater than 0.85 indicates an excellent model). Subsequently, to overcome the technical challenge of the extremely uneven distribution of positive and negative samples in biological interaction data, the Focal Loss function is introduced for model optimization:

[0082]

[0083] in, It is the target loss function value during the model training process. It is the confidence probability that the model predicts the current sample belongs to its actual class. It is a weighting balancing coefficient, which is preferably set to 0.3 in this invention to balance the differences in the number of samples. This is the focusing parameter, which is preferably set to 1.5 in this invention. Its physical meaning is to force the model to focus on difficult-to-classify functional interactions by reducing the weight contribution of easily classified samples. (Setting weight parameters) Focus parameters By reducing the weight of easily classifiable samples, the model is forced to pay more attention to hard samples (such as weak excitation or inhibition) during training, thereby effectively overcoming the prediction bias caused by the severe imbalance of positive and negative samples in biological interaction data and improving the model's accuracy in capturing rare functional regulatory terms.

[0084] (2) Automatic correction of the discrimination threshold based on Youden index

[0085] The system abandons the traditional fixed threshold of 0.5 and uses the Youden index to automatically correct the judgment boundary between "activation" (Y=1) and "inhibition" (Y=0):

[0086]

[0087] in Yoden index when the predicted probability threshold is t. It is the true rate (sensitivity / recall rate), which represents the proportion of valid functional interactions that are correctly identified. This is the false positive rate (1 - specificity), representing the proportion of incorrectly judged items. The system automatically searches for... Maximize the threshold As the final decision point, it significantly improves the accuracy of preclinical screening.

[0088] (3) Model selection and ensemble strategy based on independent validation sets

[0089] After training each base model, the system uses an independent validation set (N=749) that was not involved in training and parameter tuning to perform a final performance evaluation of the models, in order to determine the best core classification model with the highest generalization ability. The system calculates the F1-Score for the predicted probabilities P output from the validation set. Then, it performs a rank-sum test on the predicted probabilities of the validation set samples and calculates the area under the ROC curve to quantify the model's ability to distinguish between "activation" and "inhibition" categories. The system iterates through all trained classifiers (RF, XGBoost, LightGBM, LR) and ranks them according to their AUC on the independent validation set to determine the final classification model (AIPER). Subsequently, a decision enrichment analysis (EF factor) is performed based on the best model, determining that the accuracy threshold for the best model's inference is 30%, with its EF factor approaching 100% prediction accuracy.

[0090] Step 4: Construction of an automated prediction system based on the best model engine and multidimensional reliability evaluation

[0091] After determining the optimal ensemble model in step 3, this invention constructs a standardized automated prediction system for identifying drug-protein functional defenses, realizing end-to-end reasoning from raw molecular information (based on molecule name or molecule smiles) to systemic efficacy evaluation. Its core process and algorithm logic are as follows:

[0092] (1) Standardization and entity alignment of input data

[0093] The system receives the SMILES structure of the drug to be predicted and the amino acid sequence of the protein. To ensure feature consistency across databases, naming standardization is performed, using a regularization algorithm to clean the protein and drug names (e.g., unifying lowercase and removing special characters). Then, SMILES similarity matching is performed. For new drugs not in the database, the Tanimoto similarity algorithm is used to match them against the aforementioned 3M reference database.

[0094]

[0095] in, , These are the molecular fingerprint bit depths for drugs A and B, respectively. The number of bits shared by both. When At that time, the system automatically associates known pharmacological properties.

[0096] (2) Automated feature extraction and model inference

[0097] The system automatically calls the attention fusion parameters pre-trained in step 2. The input drug-protein pairs are feature-encoded. First, cross-spatial feature synthesis is performed to calculate the drug fingerprint in real time. Protein sequence embedding and the topological features of protein-protein interaction networks (PPIs) Then, the model is forward-propagated to fuse the features. The selected optimal ensemble model (AIPER) is input, and the output is the direction of drug regulation of the target protein and its original probability score P. Subsequently, to address the difficulty of quantifying "predictive power" in traditional models, this prediction system introduces two indicators: probability gap and entropy. The probability gap formula is as follows:

[0098]

[0099] in, The probability of the direction of maximum prediction. The probability of being the second-best direction. The larger the value, the higher the certainty of the prediction. Prediction uncertainty (information entropy) is measured using:

[0100]

[0101] in, These are the values ​​assigned to specific categories after the model's output layer has been normalized. The values ​​range from 0 to 1. When k=1, This could represent the probability of "activation". When k=2, This may represent the probability of "inhibition". This represents the number of categories. The lower the value, the more concentrated the model output and the more robust the prediction results.

[0102] (3) High-quality candidate drug-protein pair screening

[0103] To further improve the reliability of prediction results, when performing large-scale prediction tasks involving hundreds of millions of drugs and proteins, the probability of model prediction errors increases significantly. Therefore, EF (Enrichment Factor) analysis is performed on the independent validation set to determine the top percentage of drug-protein pairs to be included in subsequent studies, thereby improving translational efficiency. Based on the independent validation set, the original probabilities of all model predictions are... Sort the samples in descending order to obtain their rankings. The formula for calculating EF is:

[0104]

[0105] in Top of the rankings Enrichment factors in the sample. It is the original proportion of live samples in the entire dataset (random success rate). It is in the top of the model prediction scores The proportion of truly active samples in the sample. Ultimately, this determines an accuracy of EF of over 80%. The model is selected and determined based on the accuracy of the independent validation set.

[0106]

[0107] The system only retains those in the front row. High-confidence predictors of quantiles are prioritized for subsequent candidate drug screening or pharmacological validation. Representing probability rankings exceeding Drug-protein pair, This indicates that the effect of the drug on the protein is uncertain and should be excluded.

[0108] Step 5: Multi-target systemic pharmacological effect scoring based on network propagation

[0109] (1) Initial signal construction and high confidence locking

[0110] The initial signal construction is based on the predicted direction (activation / inhibition) and confidence level of the system output from step 4. Tensor quantization is performed to lock 30% of the prediction results. An initial perturbation vector is defined. Its elements The original effect value of the drug on target i:

[0111]

[0112] in It is the initial perturbation score of the drug on the i-th protein node. It is a sign function. Its physical meaning is to quantify the predicted direction of regulation: +1 for activating (Agonist) and -1 for inhibiting (Inhibitor). It is the predictive model's determination of the functional direction of the drug's action on target i. This is the confidence probability score of the prediction result; the higher the value, the more reliable the initial signal.

[0113] (2) Topological cascade propagation: random walk diffusion model

[0114] Considering that drug effects propagate along the PPI network, a diffusion model based on random walks is introduced for scoring. The signal is updated iteratively within the network:

[0115]

[0116] in This represents the initial state vector. If the protein is a direct target of the drug (labeled as...) ), its in

[0117] The value at the corresponding position is usually set to 1 if the protein is not labeled as a direct target (labeled as...). ), its in The value at the corresponding position is usually set to 0. In the The effect value state vector of all protein nodes in the network at the next iteration. It is the network state vector of the previous time step. It is the normalized adjacency matrix of the PPI network, representing the probability distribution of signal transmission between adjacent proteins. It is the restart probability, which is set to 0.6 in a preferred embodiment of the present invention to balance the direct target effect (staying in place) and the neighborhood diffusion effect (spreading outward). This is the coefficient for signal propagation to neighboring nodes. When the iteration reaches convergence (i.e., ...), ... The system obtains the global signal distribution under steady state, i.e., the perturbation score of all components. :

[0118]

[0119] in, It is the identity matrix. This represents the inverse operation of a matrix. Its mathematical essence is to find the steady-state convergent solution of a signal propagation over an infinite time span. This is a perturbation vector containing all protein nodes in the network. Its magnitude reflects the strength of the drug signal's influence after topological diffusion, while the positive or negative sign indicates the direction of the final effect triggered by the cascade reaction. Through this calculation, the system can effectively identify indirect pharmacodynamic signals generated by network compensation or cascade amplification effects.

[0120] (3) Soft penalty factor and final score

[0121] To suppress noise caused by prediction uncertainty A directional penalty factor (DPF) was introduced. This factor is based on a logistic regression function and is determined by the number of correctly targeted drug receptors. Number of target points in the wrong direction Perform dynamic weighting:

[0122]

[0123] DPF is the direction penalty factor, with a value range of (0, 1). It is a natural constant. The penalty sensitivity coefficient is set to 4 in a preferred embodiment of the present invention. The larger the value, the lower the system's tolerance for directional errors. It is the number of correct targets whose predicted direction aligns with the needs of disease treatment. It is the number of incorrect targets whose predicted direction is contrary to the needs of disease treatment.

[0124] (4) Final overall score

[0125] Final score ( It integrates system disturbance strength, path consistency, and direction determinism:

[0126]

[0127] in This represents the pathway consistency weight. Ultimately, all drugs will be ranked. The higher the score, the greater the drug's therapeutic potential for the target disease (multi-target) and the clearer its direction.

[0128] Compared with existing technologies, the drug-protein function regulation prediction method based on artificial intelligence provided by this invention has significant advantages in drug action direction determination, system-level pharmacodynamic simulation and data noise control. Its beneficial effects are mainly reflected in the following aspects.

[0129] 1. It solves the problem of difficulty in determining the direction of drug function regulation and improves the decision-making value of pharmacological prediction.

[0130] In existing technologies, most computational pharmacology methods primarily focus on the binding ability between drugs and targets, such as molecular docking methods, quantitative structure-activity relationship (QSAR)-based prediction methods, and some drug-target interaction (DTI) deep learning models. These methods typically only assess the binding strength or probability between drugs and proteins, making it difficult to further determine whether the drug exerts an activating or inhibitory effect on the target protein. Therefore, they introduce significant uncertainty when guiding subsequent functional experiments. This invention constructs a drug-protein functional regulation knowledge base based on literature semantic parsing and combines multimodal biofeedback fusion with graph neural network learning mechanisms to model and predict the functional direction of drug action, thereby achieving direct determination of drug-induced protein activation or inhibition effects. Through this technical solution, the system can not only screen potential candidate drugs but also output functional regulation prediction results with clear directions, thus providing clear hypotheses about the mechanism of action for subsequent biological experiments and significantly reducing the blind spots in experimental design.

[0131] 2. Overcoming the limitations of single-target analysis to achieve system-level pharmacodynamic simulation.

[0132] Traditional drug screening methods often treat drug action as a single-target event, assuming that drugs primarily exert their effects by regulating a single protein. However, in real biological systems, proteins coordinately regulate cellular function through complex interaction networks. The effect of a drug on a particular protein may propagate through this network, affecting multiple downstream pathways and generating cascade reactions or compensatory effects. This invention introduces a protein-protein interaction network into its technical solution and utilizes a network propagation algorithm to simulate the diffusion process of drug regulatory signals within the network. By mapping the predicted drug-protein regulatory signals onto the protein interaction network and performing iterative propagation calculations on the network nodes, a comprehensive score reflecting the system-wide drug perturbation effect is obtained. This method can identify indirect pharmacological effects induced by network topology, more realistically simulating the drug's action pattern in vivo from a systems biology perspective, thereby improving the biological interpretability and potential clinical translational value of drug prediction results.

[0133] 3. Establish a dynamic denoising mechanism for biological data to reduce the false positive rate of predictions.

[0134] In the process of integrating multi-source biological data, the original data often contains a certain degree of noise and uncertainty due to differences in literature evidence, experimental conditions, and inconsistent database updates. Without an effective data screening mechanism, a large number of false positives can easily occur during network propagation analysis, thus affecting the reliability of prediction results. To address this issue, this invention introduces a directional consistency penalty mechanism into the scoring system. This mechanism dynamically weights drug-protein regulatory relationships with conflicting prediction directions or low confidence levels, thereby suppressing potential noise signals. Through this mechanism, the system can automatically reduce the influence of uncertain signals during network propagation, resulting in candidate drugs with higher biological consistency and predictive reliability.

[0135] 4. Optimize experimental resource allocation to achieve a balance between high-throughput screening and accurate validation.

[0136] In traditional drug screening processes, high-throughput computational screening often generates a large number of candidate drugs. However, due to the lack of effective prioritization strategies, subsequent experimental validation is costly, limiting research efficiency. This invention establishes a confidence-based screening strategy based on predicted probability, statistically ranking the prediction results and prioritizing the retention of high-confidence drug-protein regulatory relationships for subsequent analysis. This strategy effectively reduces the number of candidates requiring experimental validation while ensuring that key candidate drugs are not overlooked, allowing research focus to concentrate on high-confidence prediction results, thereby significantly reducing experimental costs and improving drug discovery efficiency. Attached Figure Description

[0137] Figure 1This is a flowchart of an AI-based method for predicting drug-protein function regulation. Detailed Implementation

[0138] The present invention will now be further described with reference to the accompanying drawings.

[0139] Example 1: Construction, training, and automated hyperparameter tuning of a protein activating / inhibiting model

[0140] like Figure 1 As shown, this invention proposes a method for predicting drug-induced protein effects. The implementation schemes are tested below: Example 1 (Constructing a prediction model): Original literature data → Feature vector GNN fusion → Classification prediction → Establishment of a protein activating / inhibiting model; Example B (Prediction and drug recommendation ranking based on new drug-proteome): Model-based prediction → RWR network propagation → DPF penalty → Final drug score ranking. Its core processing flow is as follows:

[0141] This embodiment describes in detail how to construct a predictive model for drug-protein functional effects.

[0142] A1: Construction of a dataset of drug protein functional characteristics (Y feature matrix)

[0143] This step primarily relies on 3M data downloaded from the Biomedical Information Retrieval System (PubMed) and the Comparative Toxicology Genomics Database (CTD) to determine functionality. The following example uses local data; see Table 1 for details:

[0144] Table 1: Description of drug-protein interaction relationships

[0145]

[0146] The text description is first processed into a symbol matrix based on the defined set of regulatory relationships "activate, inhibit…", where words related to activation or similar words are marked Y=+1, and words related to inhibit are marked as… Then, calculations were performed according to the sign polarity formula. Subsequently, semantics defines the effect of drugs on proteins as The reverse path is Finally, the processed Y feature matrix is ​​formed (Table 2):

[0147] Table 2: Drug-protein interaction matrix

[0148]

[0149] A2: Data Preprocessing and Multimodal Feature Encoding (X Feature Matrix)

[0150] Drug feature encoding converted the drug SMILES (obtainable from the drug name via the Pubcherm website: https: / / pubchem.ncbi.nlm.nih.gov / ) of the drugs listed in Table 2 into 2048-bit molecular fingerprints (radius 2) using the RDKit tool, serving as the biophysical characterization of the drugs (encoding results are shown in Table 3). Protein feature characterization was performed based on protein sequence (the complete sequence of the protein can be obtained from the protein name via Uniprot: https: / / www.uniprot.org / ) to obtain pre-trained ESM2 embedding vectors (640 dimensions) for capturing sequence-level evolutionary information (encoding results are shown in Table 4).

[0151] Table 3: Drug Feature Coding Matrix

[0152]

[0153] Table 4: Protein Feature Coding Matrix

[0154]

[0155] Subsequently, the PPI network (Table 5) was used to extract topological embedding features such as degree centrality and betweenness centrality of nodes. Label standardization followed the semantic parsing method described in the invention, constructing a training set based on the Y (activation (+1) and inhibition (-1)) labels in Table 2. During feature fusion, a model named EnhancedGNNFeatureExtractor was constructed, where the preprocessing layer used two linear fully connected layers (hidden layer dimension 512), combined with BatchNorm1d and LeakyReLU (slope 0.1) functions for nonlinear transformation. The graph convolutional layers stacked three GNN operators; the first two layers used GATConv (multi-head attention mechanism, Heads=8), and the last layer used SAGEConv for neighborhood information aggregation to capture the deep contextual information of proteins in the PPI network. Finally, feature fusion utilized the MultiheadAttention mechanism to perform attention-weighted fusion of basic protein features, PPI topological features, and drug molecule fingerprints. To address the problem of uneven distribution of positive and negative samples in biological data, Focal Loss was introduced. The technical parameters were set to alpha=0.3 and gamma=1.5. By reducing the weights of easily classified samples, the model was forced to focus on functional interactions that were difficult to classify. The entire feature fusion process integrated the Optuna framework, using the TPE sampler to perform 20-30 rounds of iterative optimization on the hidden layer dimension (256-768), learning rate (1e-5 to 5e-3), and Dropout rate to lock in the optimal model architecture parameters. Finally, the feature fusion matrix (Table 6) and the feature learning parameters of the GNN were obtained.

[0156] Table 5: Protein-PPI Interaction Matrix

[0157]

[0158] Table 6: Feature Fusion Matrix

[0159]

[0160] A3: Ensemble learning of downstream classifiers

[0161] Based on the A2 fusion feature matrix (Table 6), the optimal model was determined using an ensemble evaluation system. Parallel training employed Random Forest (RF), XGBoost, LightGBM, and Logistic Regression (LR). Stratified KFold (5-fold hierarchical cross-validation) was used, and RandomizedSearchCV was employed to optimize tree depth and regularization parameters. The optimal model was determined during training based on a comprehensive evaluation of validation set AUC, F1-Score, and Recall (Table 7). The model with the best performance (e.g., AUC > 0.85) was defined as the core classifier. Finally, the optimal model parameters were saved (Table 8).

[0162] Table 7 Model Validation Set Parameters

[0163]

[0164] Table 8 Model Parameters

[0165]

[0166] A4: External validation of the optimal model determines the effectiveness of the model.

[0167] The optimal model based on A3 was used, and the model training parameters were saved. However, independent validation was performed on the same dataset as A1-3, using the same feature fusion and feature transformation methods. The external validation parameters of the optimal model were obtained to ultimately determine the generalization accuracy of the model. The validation results were evaluated using the following parameters, where an AUC greater than 0.8 indicates that the model has excellent classification ability (Table 9).

[0168] Table 9. Best Model Performance (These results are based on running the model on 3M data)

[0169]

[0170] A5: Optimal model predicts new datasets

[0171] Based on the model parameters saved in A4, a new dataset was used for testing (Table 10), employing the same processing methods as A1-5: drug protein feature acquisition, protein network parameter acquisition, GNN feature fusion, and optimal model classification prediction. Additionally, the protein name was entered into the STRING website (…). https: / / string-db.org / In the Multiple proteins module of the drug model, after selecting Human Organisms, PPI network data was prepared (Table 11). Subsequently, based on the model parameters provided in A4 above, the final model prediction results were obtained (Table 12). The Original_prediction column includes the functional orientation of the drug against multiple protein targets (1 represents agonist, 0 represents inhibitor); Original_probability is the inferred probability; the KD_prediction column is an auxiliary model column that determines the binding ability, with a value distribution of 0 to 5, where a larger value indicates a stronger binding ability; the IC50_prediction column is also an auxiliary model column that determines the half-inhibitory concentration of the drug small molecule against the protein, with a value distribution of 0 to 5, where a larger value indicates a larger IC50.

[0172] Table 10. New Drug-Protein Matrix (Only partial data shown here)

[0173]

[0174] Table 11: Protein-PPI Interaction Matrix (Only partial data is shown here)

[0175]

[0176] Table 12 Model prediction results of the new drug-protein matrix (only partial data is shown here)

[0177] .

[0178] Example 2: Drug Prioritization

[0179] B1: Preparation of Complex Target Disease Testing Datasets

[0180] To conduct screening tests at the systems pharmacology level, this section uses recently reported drug-protein pairs for multi-target diseases (metabolic steatohepatitis) as predictive test data (Table 10). Predictions are made based on the best model (Part A5), resulting in the prediction results shown in Table 12. In addition, PPI results are prepared (Table 11), and pathway data for significant enrichment of each protein are prepared from literature or KEGG / GO pathway enrichment (Table 13), along with data on the influence of a given protein target on the direction of disease outcome (Table 14).

[0181] Table 13 Pathways Involved by Proteins

[0182]

[0183] Table 14. The Influence of Proteins on Disease Outcomes

[0184]

[0185] B2: Perform multi-target drug prioritization analysis

[0186] Subsequently, to identify the optimal drugs for complex target diseases, the most promising lead compounds were determined based on the best model inference results (Table 12). First, all drug-protein pairs predicted by the model were sorted in descending order of their original probability scores, and only the top 30% of high-confidence predictions were included in subsequent evaluation, effectively filtering prediction noise. First, the PPI network was processed and transformed into an adjacency matrix. Taking the protein ACLY as an example, it connects multiple nodes such as DGAT2 (0.583), PPARG (0.563), and HMGCR (0.711). Therefore, to ensure signal conservation in the network, the sum of all output weights of the ACLY node needs to be 1. The total score was calculated as: 0.583 + 0.563 + 0.711 + 0.994 + 0.812 + 0.578 = 4.241. A probability was assigned to each protein pair; for example, the probability of the signal flowing from ACLY to FASN was 0.994 / 4.241. 0.234. Repeat this operation for all proteins in Table 11 to form a... The transition probability matrix W.

[0187] Subsequently, we selected one sample from Table 10 as input. For example, the drug Docetaxel acts on the target FXR, and the functional direction of Docetaxel on FXR is read as activating (Table 12: Original prediction, OP = 1), with a probability P of 0.95. Therefore, the vector of FXR... The corresponding field should be filled with: Meanwhile, all other protein positions in the network (such as ACLY, FASN, etc., which have no effect on the drug) are filled with 0, for example:

[0188]

[0189] After that, set The number of iterations T is set to 10. -6 .calculate :

[0190]

[0191] This means that in each iteration, 0.4 (40%) of the signal energy flows along the previously calculated W map to neighboring proteins, while 0.6 (60%) of the energy is forced back to PPARG. I is based on all 10 proteins (Table 10), forming a 10x10 matrix with 1s on the diagonal and 0s elsewhere. The matrix is ​​then calculated. This represents the equilibrium state reached by the entire network after countless signal transmissions, and finally, the total number of genes is calculated. :

[0192]

[0193] Docetaxel drug In the matrix, FXR achieved the highest score across the entire network (e.g., 0.60), reflecting the core efficacy. Although the model did not predict a direct interaction between Docetaxel and GLP1R, GLP1R achieved the second-highest perturbation score (e.g., 0.18) due to the strong connection between GLP1R and FXR in the PPI network (score 0.563). This demonstrates that Docetaxel can indirectly intervene in downstream lipid synthesis pathways through the PPI network compensation mechanism.

[0194] B2: Perform final drug ranking analysis for multiple targets (AIPER_UDS score)

[0195] After completion After topological diffusion calculations, the system has captured the "fluctuation intensity" of the drug throughout the biological network. The next core task is to perform self-correction based on directional consistency and ultimately output a comprehensive score for drug screening. To filter out "noise" signals that, although strong in effect, conflict with therapeutic needs in terms of regulatory direction, the system introduces a DPF (Directional Penalty Factor). This is key to the AIPER-UDS system's automatic error correction. Taking Docetaxel as an example, the system evaluates the accuracy of its regulatory direction through self-correction. Table 12 shows that Docetaxel correctly regulated 10 key targets in the disease network (N_correct=10) and had no reverse regulation terms (N_incorrect=0). If the empirical value is 4, then the penalty factor DPF of the drug is .

[0196]

[0197] Therefore, the DPF approaches 1, and the system determines that the regulatory direction of Docetaxel is extremely accurate, so no points are deducted. The DPF value is then calculated for each drug sequentially. Finally, the above dimensions are weighted and integrated to output the final evaluation index:

[0198]

[0199] Among them, the path consistency weight is set. Each medicine will receive a final score, with an experience value of 2. (Table 15) Among them, Lepirudin is the most effective for multi-target treatment of the disease.

[0200] Table 15. The Influence of Proteins on Disease Outcomes

[0201] .

[0202] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for predicting drug-protein functional regulation based on artificial intelligence, characterized in that, The specific steps are as follows: Step 1: Construct a prior knowledge base for protein function regulation; Transform each unstructured description in the original biomedical literature into a drug-protein function regulation tag with a clear direction. This allows for the construction of a structured prior knowledge base; Step 2: Multimodal biometric feature extraction and fusion; Feature fusion is performed on multimodal data, which includes structural features of drug nodes, structural features of protein nodes, protein interactions, and drug-protein functional interaction binary network features. Step 3: Construct a deep learning-based functional regulation prediction model; Using the functional regulation tags obtained in step 1 as supervision information and the multimodal biofeedback obtained in step 2 as model input, a drug-protein function regulation prediction model is constructed. Step 4: Construction of an automated prediction system based on the best model engine and multi-dimensional reliability evaluation; A standardized, automated predictive system for identifying drug-protein functional defenses will be constructed to achieve end-to-end reasoning from raw molecular information to systemic efficacy evaluation; specifically as follows: (41) Standardization and entity alignment of input data; The system receives the SMILES structure and protein amino acid sequence of the drug to be predicted; then it proceeds to SMILES similarity matching. For new drugs not in the database, it uses the Tanimoto similarity algorithm to perform matching in the prior knowledge base described in step 1. ; in, , These are the molecular fingerprint bit depths for drugs A and B, respectively. The number of bits shared by both; when At that time, the system automatically associates known pharmacological properties; (42) Automated feature extraction and model inference; The system automatically calls the pre-trained multimodal biometric feature extraction and fusion parameters from step 2. The input drug-protein pairs are feature-encoded; firstly, cross-spatial feature synthesis is performed to calculate the structural features of the drug nodes in real time. Structural features of protein nodes and the characteristics of protein-protein interaction and drug-protein functional interaction binary networks. Then, the model is forward propagated to fuse the features. Input the selected functional regulation prediction model and output the direction of drug regulation of the target protein and its original probability score P; then introduce the dual indicators of probability gap and information entropy. The probability gap formula is: ; in, The probability of the direction of maximum prediction. The probability of being the second-best direction; The larger the value, the higher the certainty of the prediction; The information entropy used to predict uncertainty is: in, It is the numerical value assigned to a specific category after the model output layer has been normalized, ranging from 0 to 1; The number of categories; The lower the value, the more concentrated the model output and the more robust the prediction results; (43) High-quality candidate drug-protein pair screening; Based on the independent validation set, the raw probabilities predicted by all models are... Sort the samples in descending order to obtain their rankings. The formula for calculating EF is: ; in, Indicates ranking at the top Enrichment factors in the sample; This represents the original proportion of live samples in the entire dataset, and the random success rate. This indicates the top score in the model prediction. The proportion of truly active samples in the sample; Ultimately, EF was determined to have an accuracy rate of over 80%. Based on the accuracy of the model, an independent verification set is used to determine the model selection criteria. ; The system only retains those in the front row. High-confidence predictors of quantiles are prioritized for subsequent candidate drug screening or pharmacological validation; among them, Representing probability rankings exceeding Drug-protein pair, This indicates that the effect of the drug on the protein is uncertain and should be excluded.

2. The drug-protein function regulation prediction method based on artificial intelligence according to claim 1, characterized in that, Step 1 is described in detail as follows: (11) Collect literature records containing information on drug-protein functional interactions from publicly available biomedical literature databases. The system identifies the regulatory relationship between drugs and proteins through keyword identification and contextual semantic analysis using a large language model, and classifies the regulatory relationship into activation and inhibition regulation categories. (12) For complex nested sentences, there are multiple semantic flips. After the semantic blocks are divided by the stack data structure, the direction is determined as the final variable of the model. When a logical guide word or parentheses is encountered, a push operation is performed. After parsing, the local polarity is calculated by popping from the stack. (13) Perform symbolic valence reasoning and causal modeling, and calculate the symbolic direction of each semantic block. : ; in, For positive verbs, add 1; for negative verbs, subtract 1. Then, determine the direction using multiplication. In direction determination, there is a bidirectional interaction between the drug and the protein, where the drug's effect on the protein pathway is assigned a value. Conversely, the path is assigned a value. ; (14) Combine the positive / negative direction of the sign with the position coefficient to generate the final label. : ; in, This is a sign function used to extract the sign of a numerical value. If the result within the parentheses is greater than 0, then... This indicates an excitation effect; secondly, if the calculation result within the parentheses is less than 0, then... , represents the inhibition effect; B represents the total number of semantic blocks obtained after semantic segmentation of the above complex nested sentence structure, including the fact that after the text is separated into specific semantic blocks, each complex sentence will be broken down into several local Bs; b is the counting variable of the accumulated symbols, representing the b-th semantic block currently being calculated; b equals 1 in the formula, which means that the inference results of all symbols from the first to the B-th semantic blocks will be accumulated and summed.

3. The method for predicting drug-protein function regulation based on artificial intelligence according to claim 1, characterized in that, Step 2 is described in detail below: (21) Obtain the structural features of drug nodes; Drug molecular fingerprint characterization uses the RDKit molecular structure information dimensionality reduction tool, which describes the molecular spatial structure based on SMILES, and calculates the feature quantities of each molecular fingerprint: ; in, It is the current atom. They are adjacent atoms. Where is the atomic radius. It is an atom exist The feature vector at the nth iteration It is with atoms The set of directly connected neighboring atoms, It is an atom and Information on the type of chemical bonds between them; It is a preset hash mapping operator used to perform a nonlinear mapping between the current atomic feature and the neighboring atomic features to obtain a feature vector of fixed length; Represents atoms In the The feature representation at the next iteration encapsulates the topological distance as The local chemical environment within the range; fingerprint of the entire molecule It is a folded mapping of the feature sets generated by all atoms at all iteration levels: ; in, This indicates traversing every atom in the molecular diagram G. ; Recording atoms From itself to a radius of All structural topology information in the environment; This represents merging all substructural features generated by all atoms of the total molecular diagram G at different radii; It performs a folding operation, compressing the feature identifier to a fixed length; This is a condition that Fold mapping needs to satisfy, namely, mapping all extracted features of the molecular graph G into a vector of a set threshold length, so as to capture the connection information of molecular substructures; (22) Obtain the structural features of protein nodes; Protein sequence deep embedding utilizes a pre-trained protein language model, ESM2, to extract semantic representations of the sequence and construct a fixed-length deep embedding vector for the protein sequence. : ; Where L is the sequence length; i is the position index of the residue in the sequence; This indicates the specific amino acid residue at the i-th position in the sequence; It is a pre-trained ESM2 encoder; Global average pooling is a process that maps protein sequences of different lengths to a fixed-dimensional feature space by averaging the hidden states of all residues. (23) Obtain the characteristics of the protein-protein interaction and drug-protein functional interaction binary network; The sequence features generated in step (22) As nodes, protein system embedding vectors are constructed using the protein-protein interaction network (PPI) graph. The protein-protein interaction network (PPI) graph is based on interaction data with a moderate confidence score > 400 extracted from the STRING protein-protein interaction functional network database. Graph neural network operators are used to encode the topological location of protein nodes within the biological system. The initial state features of nodes in the PPI network are also included. ,Right now = Construct protein system embedding vectors that incorporate network topology information. : ; Where A represents the adjacency matrix of the protein-protein interaction (PPI) network, and its elements represent whether there is an interaction between protein nodes; GNN represents the graph neural network operator. (24) Multimodal data feature fusion; The heterogeneous features from the three dimensions of steps (21), (22), and (23) are deeply interacted to output the final predicted probability; as follows: First, in order to input the attention mechanism, the system treats the aligned three types of features as a single feature sequence: ; in, It is a tensor splicing operation; It is the generated input feature matrix, with 3 rows corresponding to the three modalities and the number of columns representing the hidden layer dimension; It is a vector after mapping, which maps all features to the same feature dimension; This means mapping the three features to the same real number space. ,in Represents the hidden layer dimension; Then, the association weights between the multimodal self-attention and the various dimensions of the protein are calculated: ; in, It is a randomly initialized 512 The 512 learnable parameter matrix has the core function of mapping input modes to the query space; It can gradually evolve from an initial random state into a mature feature extractor, capable of accurately identifying cross-spatial correlations between different modalities; when Multiply After that, the generated The matrix represents the retrieval vectors for the three modalities; It is a randomly initialized 512 The 512 learnable parameter matrix has the core function of mapping input modes to the key space; It can gradually evolve into an indexer capable of extracting core recognition features, when Multiply After that, the generated The matrix represents the recognition label vectors for the three modalities; It is a randomly initialized 512 The 512 learnable parameter matrix's core function is to map the original modal features to the value space; through training, Having learned how to extract the most valuable substantive features for determining drug efficacy from massive amounts of biological information, when Multiply After that, the generated The matrix represents the content vectors of the three modalities; in this way, the model can automatically identify the key features that contribute most to the determination of drug efficacy using a self-attention mechanism. : ; in, It is the final comprehensive representation, which, through a self-attention mechanism, enables the model to learn to automatically focus on the most critical information sources when faced with different drug-protein pairs. The weighted average, with weights from and The degree of matching determines; Calculate the correlation score matrix among the three types of features; This is a scaling factor used to prevent the gradient from vanishing due to excessively large dot products. The dimension of the attention head; It is a normalization function that converts the relevance score into a weighting coefficient. Finally, the system reassembles the attention-weighted sequence features into a fixed-length fusion vector. ; ; in, This represents global average pooling, which extracts the most representative comprehensive efficacy characterization by averaging the attention outputs of the three modalities, and then converts the original input... With attention output Add them together, retaining the underlying feature information for residual connection; Representative layer normalization ensures that the fused feature vectors The distribution is statistically stable, which helps the subsequent classifier determine the activation / inhibition direction.

4. The drug-protein function regulation prediction method based on artificial intelligence according to claim 1, characterized in that, Step 3 is as follows: Before entering model training, the system first performs Z-Score normalization on the feature matrix: ; in, This represents the standardized feature vector; It is the original multimodal feature input; It is the arithmetic mean of the features in the training set; It is the standard deviation of the features in the training set; (31) Integrated model construction and hyperparameter optimization; The dataset was divided into training, testing, and independent validation sets in an 80:15:15 ratio; 5-fold hierarchical cross-validation was used to ensure that the ratio of activated Y=1 labels to suppressed Y=0 labels in each fold was consistent with the overall ratio; parallel training included four heterogeneous classification algorithms: Random Forest (RF), XGBoost, LightGBM, and Logistic Regression (LR); the system was evaluated based on the precision and recall of each model on the validation set. Precision reflects the proportion of samples predicted to be agonist Y=1 that are actually active: ; Wherein, TP represents a true positive, a sample predicted as 1 that is actually 1; FP represents a false positive, a sample predicted as 1 that is actually 0; TN represents a true negative; and FN represents a false negative. Recall reflects the proportion of all actually active samples that the model successfully retrieves: ; The F1-Score is derived from the harmonic mean of precision and recall. This metric comprehensively measures the model's discriminative performance under uneven distribution of positive and negative samples. The calculation formula is as follows: ; Subsequently, ROC and AUC are the most crucial statistical indicators for evaluating the model's agonistic and inhibitory efficacy; among them, the ROC curve is based on... The false positive rate (FPR) is plotted by shifting the classification threshold (Threshold) to depict the dynamic performance of the model's discrimination. The FPR calculation formula is as follows: ; AUC is the area under the ROC curve, used to quantify the model's ability to distinguish between different categories; its value ranges from 0.5 to 1, and the closer it is to 1, the more accurate the model's discrimination. in, It is the sum of the ranks of all positive sample activation classes in the predicted probability ranking; , These represent the number of positive and negative samples, respectively. The essence of AUC is the probability that the model will rank a positive sample before a negative sample when randomly selecting one positive and one negative sample. Based on AUC, F1-Score, and Recall performance, the best classifier is automatically determined or a weighted ensemble is performed. Subsequently, the Focal Loss function is introduced for model optimization. in, It is the target loss function value during model training; It is the confidence probability that the model predicts the current sample belongs to its actual class; It is the weighting balance coefficient; It's the focus parameter; setting the weight parameter. Focus parameters ; (32) Automatic correction of the discrimination threshold based on the Youden index; The Youden index is used to automatically correct the decision boundary between activating Y=1 and inhibiting Y=0: ; in, This represents the Yoden index when the predicted probability threshold is t. It is the true rate, representing the proportion of correctly identified effective functional interactions; It is the false positive rate, representing the proportion of items incorrectly judged; the system automatically searches for... Maximize the threshold As the final decision point; (33) Model selection and ensemble strategy based on independent validation sets; After training each base model, the system uses an independent validation set (not involved in training and parameter tuning) to perform a final performance evaluation of the model, establishing the best core classification model with the highest generalization ability. The system calculates the F1-Score based on the predicted probability P output from the validation set. Then, it performs a rank-sum test on the predicted probabilities of the validation set samples and calculates the area under the ROC curve to quantify the model's ability to distinguish between activation and inhibition classes. The system iterates through all trained classifiers and ranks them according to their AUC (Adjustment Value) on the independent validation set to determine the final classification model. Subsequently, based on the best model, a decision enrichment analysis is performed, determining that the accuracy threshold for the best model's inference is 30%, with an EF (Expectation Factor) approaching 100% prediction accuracy.

5. The drug-protein function regulation prediction method based on artificial intelligence according to claim 1, characterized in that: It also includes step 5, a multi-target systems pharmacology assessment based on network propagation; details are as follows: (51) Initial signal construction and high-confidence locking; The initial signal construction is based on the prediction direction and confidence level output from the system in step 4. Tensor quantization is performed to lock 30% of the prediction results; an initial perturbation vector is defined. Its elements The original effect value of the drug on target i: ; in, It is a sign function, and its physical meaning is to quantify the predicted direction of regulation: +1 for activation and -1 for inhibition. It is the predictive model's determination of the functional direction of the drug's action on target i; This is the confidence probability score of the prediction result; the higher the value, the more reliable the initial signal. (52) Topological cascade propagation: random walk diffusion model; Signals are updated iteratively within the network: ; in, This represents the initial state vector. If the protein is a direct target of the drug, it is labeled as... , its in The value at the corresponding position is usually set to 1; if the protein is not labeled as a direct target, it is labeled as... , its in The value at the corresponding position is usually set to 0; In the At the next iteration, the effect value state vector of all protein nodes in the network; It is the network state vector from the previous time step; It is the normalized adjacency matrix of the PPI network, representing the probability distribution of signal transmission between adjacent proteins; It is the restart probability, used to balance the direct target effect and the neighborhood diffusion effect; This is the coefficient for signal propagation to neighboring nodes; when the iteration reaches convergence, i.e. The system obtains the global signal distribution under steady state, i.e., the perturbation score of all components. : ; in, It is the identity matrix. It represents the inverse operation of a matrix; its mathematical essence is to find the steady-state convergent solution of signal propagation over an infinite time span. It is a perturbation vector containing all protein nodes in the network. Its magnitude reflects the influence strength of the drug signal after topological diffusion, and the positive or negative sign represents the direction of the final effect caused by the cascade reaction. Through this calculation, the system can effectively identify indirect drug effect signals generated by network compensation or cascade amplification effects. (53) Soft penalty factor and final score; A direction penalty factor (DPF) was introduced; this factor is based on a logistic regression function and is determined according to the number of correct target points regulated by the drug. Number of target points in the wrong direction Perform dynamic weighting: ; Wherein, DPF is the direction penalty factor, and its value ranges from (0, 1); It is a natural constant; This is a penalty sensitivity coefficient; the larger the value, the lower the system's tolerance for orientation errors. It is the number of correct targets whose predicted direction aligns with the needs of disease treatment. It is the number of incorrect targets whose predicted direction is contrary to the needs of disease treatment; (54) Final overall score; Final score It integrates system disturbance strength, path consistency, and direction determinism: ; in, This represents the pathway consistency weight; ultimately, it will result in a ranking of all drugs. The higher the score, the greater the drug's therapeutic potential for the target disease and the clearer its direction.