Method, system and computer-readable storage medium for predicting protein interaction regulators based on knowledge-enhanced language model

By combining the knowledge-enhanced language model of molecular structure encoding and gene ontology map, the problem of insufficient utilization of external knowledge in existing methods is solved, and accurate prediction and efficient performance improvement of protein interaction regulators are achieved.

CN119993283BActive Publication Date: 2025-07-22YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510464934.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-22
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

Existing methods for predicting protein interaction regulators fail to effectively utilize external knowledge, resulting in low prediction performance, especially when the rich description of biomolecular properties and functionalities in scientific literature and biological databases are ignored.

Method used

Using a knowledge-enhanced language model method, by introducing molecular structure encoder and gene ontology maps, combining gene-based language models and Transformer architectures, integrating external knowledge and language models, constructing regulators and protein feature representations, and using multi-task learning and feature fusion blocks for prediction.

Benefits of technology

Accurate prediction of protein interaction regulators is achieved, performance in small sample scenarios is improved, functional mechanisms are captured, and the generalization ability and prediction accuracy of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993283B_ABST
    Figure CN119993283B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting protein interaction regulators based on a knowledge-enhanced language model, including: constructing a regulator model: pre-training a text encoder, introducing a fingerprint-based structure encoder, and splicing the text embedding output by the text encoder and various fingerprint embeddings output by the structure encoder to generate regulator features; constructing a protein model: collecting GO graphs and GO annotations, training a graph embedding model based on the GO graphs, using the trained graph embedding model to generate structural features for each GO term node, combining the structural features of the GO terms to output GO term embedding vectors, and generating protein features based on multiple GO term embedding vectors; constructing a prediction model: using the regulator features and protein features as inputs, outputting interaction features, and outputting a prediction result based on the interaction features. This solution effectively improves the prediction performance of protein interaction regulators through the integration of the subtle expressions of natural language and the structural properties of biomolecules.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and particularly relates to a method, system and computer-readable storage medium for predicting protein interaction regulators based on a knowledge-enhanced language model. Background Art

[0002] The regulation of protein-protein interaction (PPI) plays a key role in disease treatment and drug development. Accurately predicting small molecule compounds (regulators) that can modulate specific PPIs has become one of the core challenges in computational biology and drug design. Existing methods for predicting protein interaction regulators include target-free methods (utilizing only regulator features) and target-based methods (utilizing both regulator features and protein interaction features).

[0003] Target-free methods cannot elucidate the interaction mechanism, which limits the precise design and optimization of regulators. Target-based methods need to be further enhanced in terms of biomolecular representation. Existing methods only predict protein interaction regulators based on molecular or protein sequences, often ignoring rich external knowledge. In the field of biomedicine, scientific literature, biological databases, and knowledge graphs, as important knowledge sources, provide detailed descriptions of the properties, functions, and interactions of various biomolecules, which cannot be directly inferred solely from molecular or protein sequences. For example, the sequences of two proteins may be highly similar, but if they are annotated as "nuclear" and "mitochondrial" localizations respectively, the probability of actual interaction is extremely low, and the prediction performance of existing methods is not high due to the lack of explicit modeling of such knowledge. Summary of the Invention

[0004] The object of the present invention is to propose a method, system and computer-readable storage medium for predicting protein interaction regulators based on a knowledge-enhanced language model for the problems existing in the prior art.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A method for predicting protein interaction regulators based on a knowledge-enhanced language model, the method includes,

[0007] Regulator Model Construction

[0008] Collect the SMILES sequences and text descriptions of molecules, and expand the text descriptions through a generative language model to construct a pre-training corpus;

[0009] Pre-train a text encoder based on the pre-training corpus;

[0010] Introduce a fingerprint-based structure encoder, use multiple fingerprints to represent the molecular structure, and use a separate linear projection head to map each fingerprint to the embedding space of the text encoder;

[0011] Concatenate the text embedding output by the text encoder and the multiple fingerprint embeddings output by the structure encoder to generate the regulator features;

[0012] Protein model construction

[0013] Collect the GO graph and GO annotations, and use the BERT model as the backbone network;

[0014] GO represents Gene Ontology, and the GO annotations include the names and definitions of the corresponding GO terms;

[0015] Each node of the GO graph represents a GO term, and the edges represent the relationships between GO terms;

[0016] Train a graph embedding model based on the GO graph;

[0017] Use the trained graph embedding model to generate structural features for each GO term node;

[0018] The backbone network converts the names and definitions of GO terms into word embeddings, and the structural features of GO terms are injected into the backbone network through the [SK] token to output the GO term embedding vector;

[0019] Aggregate multiple GO term embedding vectors to generate protein features;

[0020] Prediction model construction

[0021] It includes a fusion block and a two-layer perceptron. The fusion block takes the regulator features output by the regulator model and the protein features output by the protein model as inputs, outputs the interaction features, and the two-layer perceptron outputs the prediction result based on the interaction features.

[0022] In the above protein interaction regulator prediction method based on a knowledge-enhanced language model, for the regulator model, first pre-train the text encoder using a pre-trained corpus, and then introduce the structure encoder during the downstream prediction task, and train the structure encoder and the text encoder simultaneously;

[0023] For the protein model, first train a graph embedding model using the GO graph, then use the trained graph embedding model, as well as the GO graph and GO annotations, to train the GO term encoder to output the GO term embedding vector containing the GO term structural features. Finally, during the downstream prediction task, use the GO term embedding vector output by the trained GO term encoder to train the GO term aggregator;

[0024] For the prediction model, a pre-trained regulator model, protein model, and a pre-constructed PPI regulator dataset are used for training.

[0025] It should be noted that not the entire regulator model and protein model are pre-trained here. For the regulator model, only its text encoder is pre-trained, and for the protein model, only its GO term encoder is pre-trained.

[0026] In the above-mentioned method for predicting protein interaction regulators based on a knowledge-enhanced language model, the structural features of GO terms are injected into the backbone network in the following specific way:

[0027] Use a tokenizer to perform word embedding on the text sequence of GO annotations and define a special token [SK];

[0028] Use the structural features generated by the graph embedding model to replace the embedding of [SK] for encoding, and output the GO term embedding vector;

[0029] Train the above-mentioned GO term encoder through multi-task learning;

[0030] And the training tasks include:

[0031] Prediction of the direct neighbors of GO terms, where the direct neighbors are defined as the union of sub-terms and parent terms;

[0032] Prediction of the members of the sub-ontology of GO terms, where the sub-ontology categories include cellular components, molecular functions, and BP biological processes.

[0033] In the above-mentioned method for predicting protein interaction regulators based on a knowledge-enhanced language model, the regulator model has a fusion block. After text embedding and various fingerprint embeddings are concatenated, they are input into the fusion block to generate the regulator features;

[0034] The protein model has a fusion block. The protein embedding output by the GO term aggregator is input into the fusion block to generate the protein features;

[0035] The regulator model, protein model, and prediction model reuse the same fusion block.

[0036] In the above-mentioned method for predicting protein interaction regulators based on a knowledge-enhanced language model, the construction process of the GO term aggregator includes:

[0037] Retrieve the GO annotations of each protein and construct a GO term set , where corresponds to a single GO term related to the protein;

[0038] Output GO term embedding vectors for each item in set S using the trained GO term encoder;

[0039] Concatenate all GO term embedding vectors to construct the protein embedding ;

[0040] The protein embedding E is fed into the fusion block to generate the protein features described above.

[0041] In the above method for predicting protein interaction regulators based on a knowledge-enhanced language model, the fusion block is constructed based on the Transformer encoder architecture, including a multi-head self-attention layer, a SwiGLU feed-forward network layer, and a fully connected layer, and layer normalization and residual connections are used between layers.

[0042] In the above method for predicting protein interaction regulators based on a knowledge-enhanced language model, the regulator model uses the RoBERTa model as the text encoder, and retrains the tokenizer of the RoBERTa model using byte-level byte pair encoding, pre-trains the RoBERTa model through a masked language modeling task, and calibrates the molecular structure representation using a fingerprint-based structure encoder;

[0043] The structure encoder represents the molecular structure using extended connectivity fingerprints, MACCS fingerprints, atom pair fingerprints, and topological fingerprints.

[0044] In the above method for predicting protein interaction regulators based on a knowledge-enhanced language model, the protein model uses struc2vec as the graph embedding model, generates embedding vectors for each GO term node using the trained struc2vec, and projects the embedding vectors into the text space of the BERT model using a linear transformation and then injects them into the BERT model as structural features.

[0045] A system for predicting protein interaction regulators based on a knowledge-enhanced language model, comprising:

[0046] A regulator encoding unit configured to execute the regulator model constructed by the method for predicting protein interaction regulators based on a knowledge-enhanced language model;

[0047] A protein encoding unit configured to execute the protein model constructed by the method for predicting protein interaction regulators based on a knowledge-enhanced language model;

[0048] A fusion prediction unit configured to execute the prediction model constructed by the method for predicting protein interaction regulators based on a knowledge-enhanced language model;

[0049] An interaction unit for receiving a SMILES sequence of a molecule to be predicted and a text description extended by a generative language model for processing by a modulator encoding unit and outputting modulator features, receiving GO annotations of a protein target for processing by a protein encoding unit and outputting protein features, and displaying a prediction result output based on the modulator features and protein features described above.

[0050] A computer-readable storage medium storing a computer program thereon, and the computer program, when executed by a processor, implements the protein interaction modulator prediction method described in item.

[0051] The advantages of the present invention are as follows:

[0052] 1) The knowledge-enhanced protein interaction modulator prediction method proposed in this solution constructs an overall view of modulator-PPI interaction by fusing the subtle expressions (functional descriptions) of natural language with the structural attributes of biomolecules (SMILES sequences, fingerprints). By integrating external knowledge, the model can capture functional mechanisms rather than just sequence or structural similarities. Proven by experiments on the Caspase-9 / XIAP interaction, it can achieve precise elucidation of key roles;

[0053] 2) This solution collects SMILES sequences and text descriptions from the PubChem database and enriches them through a generative language model such as GPT-3.5, supplements chemical properties, enhances the semantic density of the corpus, and establishes a large-scale small molecule pre-training corpus covering a wide chemical space. By enhancing data diversity and semantic density, this corpus effectively alleviates the performance degradation problem of traditional methods in few-shot scenarios;

[0054] 3) The modulator representation model in this solution adopts a combined design of a RoBERTa text encoder and a molecular fingerprint structure encoder to capture the semantic associations of SMILES sequences and text descriptions and the precise chemical properties of molecular topologies respectively. By linearly projecting to align the embedding spaces of the two modalities, while ensuring semantic coherence, it calibrates the errors that may be introduced by the generative text description. It has been verified that the pre-trained text encoder in the modulator representation model performs excellently in cross-task transfer and outperforms existing methods in the potency prediction of 7 / 9 PPI families;

[0055] 4) This solution constructs a protein representation model based on Gene Ontology, which combines factual knowledge from the GO graph and GO annotations. The structural information of each GO term is injected into the term-based language model through a special token and processed together with the text annotations; furthermore, a multi-task learning framework for direct neighbor prediction and sub-ontology classification is designed to capture the logical dependencies between terms; in the experimental case of Caspase-9 / XIAP interaction, the key GO terms highlighted by the model are highly consistent with the experimentally verified functional mechanisms, verifying its effectiveness;

[0056] 5) This solution designs a general feature fusion block based on the Transformer architecture. Through the multi-head self-attention mechanism and the SwiGLU feed-forward network, it captures the complex interactions and dependencies between input features, flexibly adapts to different input features, and significantly improves compared with existing methods. Brief Description of the Drawings

[0057] Figure 1 is the overall framework diagram of the protein interaction regulator prediction method based on the knowledge-enhanced language model of this solution;

[0058] Figure 2 is the structural schematic diagram of the fusion block in the protein interaction regulator prediction method based on the knowledge-enhanced language model of this solution;

[0059] Figure 3 is the training process of the protein interaction regulator prediction method based on the knowledge-enhanced language model of this solution;

[0060] Figure 4 is the model interpretability study of the embodiments provided in the protein interaction regulator prediction method based on the knowledge-enhanced language model of this solution;

[0061] Figure 5 is the structural block diagram of the protein interaction regulator prediction system based on the knowledge-enhanced language model of this solution.

[0062] Reference Numerals: Regulator Encoding Unit 1; Protein Encoding Unit 2; Fusion Prediction Unit 3; Interaction Unit 4. Detailed Description of the Invention

[0063] This solution provides a protein-protein interaction (PPI) regulator prediction method based on a knowledge-enhanced language model, which integrates independent regulator and protein representation models. By combining external knowledge with the language modeling architecture, the expression ability of biomolecular properties is improved. Specifically:

[0064] For the modulator model, first collect the SMILES sequences and text descriptions from the PubChem database, and enrich these text descriptions through generative language models such as the GPT-3.5 model to obtain more detailed chemical property and functional group information, so as to establish a large-scale corpus. In this embodiment, the GPT-3.5 model is used as an example. When put into use, other generative language models can also be used. Then, pre-train a Roberta model on this corpus and use it as the backbone text encoder. Due to the lack of detailed descriptions of many modulator molecules in PubChem and the inaccuracy problems in the generation of large language models, this method additionally introduces a fingerprint-based structure encoder to calibrate the knowledge provided only by the GPT-3.5 model.

[0065] For the protein model, this solution combines the factual knowledge from the Gene Ontology (GO) graph and GO annotations. Both the GO graph and GO annotations contain the corresponding GO terms. The structural information of each GO term is injected into the BERT-based language model through a special token and processed together with the text annotations. In addition, direct neighbor prediction and sub-ontology member prediction tasks are designed to enhance the language model's understanding of the basic semantics and context relationships of the terms.

[0066] Finally, apply these two representation models to the protein interaction modulator prediction task, where the modulator representation and the protein representation are integrated through a carefully designed attention-based fusion block. The specific process is as Figure 1 shown:

[0067] The first step is data set construction

[0068] First, collect the SMILES sequences and corresponding text descriptions of 301,566 molecules from the PubChem database to establish a pre-training corpus for the modulator model.

[0069] Then, use GPT-3.5 to enrich the previously collected molecular text descriptions to obtain detailed information about chemical properties and functional groups. For example, the original text description of "ethanol" is "Ethanol is a common organic solvent with antibacterial properties". Provide the original SMILES sequence "CCO" of ethanol and its description "Ethanol is a common organic solvent with antibacterial properties" to GPT-3.5 to obtain the extended text description: "Ethanol, chemical formula , is a monohydric alcohol containing a hydroxyl (-OH) functional group. This molecule has significant hydrophilicity, with a logP value of -0.18. Ethanol has a molecular weight of 46.07 g / mol and a boiling point of 78.37°C. It is often used as a solvent in pharmaceutical preparations. Its hydroxyl functional group gives it weak acidity (pKa ≈ 15.9), enabling it to participate in proton transfer reactions. The volatility and lipophilicity of ethanol allow it to penetrate cell membranes and disrupt the lipid bilayer structure of microorganisms, thereby exerting a broad-spectrum antibacterial effect.”

[0070] Then, the latest Gene Ontology data released in September 2024 was collected, and GO terms marked as "obsolete" were removed, resulting in 44,261 terms, including 8,888 "biological process (BP, describing biological activities such as "cellular respiration" and "DNA repair")" terms, 11,177 "cellular component (CC, describing the cellular location where gene products are located, such as "mitochondria" and "ribosome")" terms, and 4,196 "molecular function (MF, describing molecular-level functions such as "ATPase activity" and "protein binding")" terms. The set of all GO terms and their inclusion relationships was organized into a directed acyclic graph, i.e., the GO graph, where each node represents a GO term, the edges represent the relationships between terms, and each term is labeled with its direct neighbors and subontology members.

[0071] Finally, the interactions between small molecule regulators and PPI targets were extracted from the public database DLiP to construct a benchmark dataset for PPI regulator prediction. Using a filtering procedure consistent with existing methods, data entries of illegal PPI targets and non-human species were further removed. After this processing, a total of 11,145 regulator-PPI target interaction (PPIMI) tuples were generated, including 9,343 small molecule regulators and 117 PPI targets. These experimentally verified PPIMI tuples were regarded as positive samples. Correspondingly, negative samples were generated by replacing the regulators or PPI targets in the positive samples. The ratio of positive samples to negative samples was controlled at 1:1, and it was ensured that the sampled negative samples did not appear in the positive samples for effective training.

[0072] The second step, construction of the regulator model

[0073] The regulator model consists of two different sub-modules: a backbone text encoder and a structure encoder for knowledge calibration. They work together to generate a comprehensive and stable representation of the regulator.

[0074] The text encoder is implemented based on the RoBERTa architecture and is used to process the SMILES sequences and text descriptions of molecules. The tokenizer of RoBERTa is retrained using Byte Level Byte-Pair Encoding to effectively process SMILES strings and domain-specific vocabulary, making it suitable for the tokenization requirements of chemical symbols and biomedical proprietary terms. The definition of special tokens is consistent with RoBERTa, and [SEP] is added as the exact boundary between the SMILES sequence and the text description to help the model understand the input text. Then, using the masked language modeling task, the text encoder is pre-trained based on the pre-trained corpus enriched with text descriptions by GPT-3.5 constructed in the first step, where each token is randomly masked with a probability of 15% and reconstructed based on the context, finally obtaining the pre-trained text encoder.

[0075] As a general large language model, GPT cannot avoid occasional inaccuracies, which will reduce the reliability of structure-activity analysis. To alleviate this problem, this solution introduces a fingerprint-based structure encoder for knowledge calibration when using the modulator model for the downstream task of PPI modulator prediction. Specifically, multiple types of molecular fingerprints are used to represent various aspects of the molecular structure. These include Extended-Connectivity Fingerprints (ECFPs) for capturing the characteristics of the local atomic environment in the molecule; MACCS fingerprints for encoding predefined substructure patterns; Atom-Pair Fingerprints (APFPs) for representing atom pairs in the molecule and their topological distances; and Topological Fingerprints (TFPs) for encoding features based on the molecular graph. By leveraging the structure encoder to synthesize the information contained in different fingerprints, the model is calibrated with knowledge in the downstream task, thereby enhancing the model's understanding and prediction ability of the molecular structure.

[0076] Then, a separate linear projection head is used to align each type of molecular fingerprint with the embedding space of the RoBERTa text encoder to ensure the effective integration of structural and text information. Finally, the four fingerprint embeddings and one text embedding are concatenated, a total of five molecular embeddings, and they are fed into the fusion block to generate the final molecular representation as the modulator feature.

[0077] Step 3: Protein model construction

[0078] The protein model consists of two stacked modules:

[0079] 1) GO term encoder, which processes and encodes GO terms, is constructed based on the BERT model and is responsible for converting the text information and structural information of GO terms into rich GO term embeddings;

[0080] 2) GO term aggregator, which jointly models multiple GO term embedding vectors to generate the protein representation.

[0081] The specific steps for constructing the GO term encoder are as follows:

[0082] Use the BERT model as the backbone network of the protein model, and inject the structural information of the extracted GO terms into its text representation for subsequent joint modeling. Initialize the model weights using ouBioBERT, which has been pre-trained on large corpora in disciplines such as biology and medicine. The BERT model includes Figure 1 the BERTEncoder and BERT Embeder in, and in this solution, the two parts are used separately to inject structural embeddings.

[0083] Train the graph embedding model struc2vec on the entire GO graph to extract the context information implicit in GO terms and their interconnections. Generate embedding vectors for each term node using the pre-trained struc2vec. These embedding vectors represent the structural features of GO terms and are injected into the backbone network as structural knowledge to enhance the GO representation. To bridge the gap between the GO structure and the annotation text, the struc2vec embeddings are projected into the text space of the BERT model using a linear transformation before injection.

[0084] For the text information of GO terms, concatenate the name and definition of each GO term to form a text sequence that provides a detailed description of protein function. Process the aforementioned text sequence using the WordPiece tokenization method and add special tokens [CLS], [PAD], and [SEP]. In addition, define a special token [SK] as a placeholder. The embedding of the special token [SK] will be replaced by the structural features encoded by struc2vec in order to improve the GO representation in subsequent information propagation and inject the structural knowledge into the GO representation framework. In the above way, the structural information and text information of GO terms are combined together in a specific format to form a comprehensive input sequence, which can be in the form of {[CLS][SK][SEP]Name + Def[SEP]}, where:

[0085] [CLS], is a special token indicating the start of the entire sequence;

[0086] [SK], is a specially defined token as a placeholder for injecting the structural information of GO terms. During the model processing, the embedding of [SK] will be replaced by the structural features encoded by struc2vec;

[0087] [SEP], is a separator used to separate different parts, that is, to separate the text part of the [SK] GO term;

[0088] Name+Def is a text sequence consisting of the name and definition of GO terms.

[0089] To enhance the joint representation of semantics and structure, different prediction heads are added on top of the GO term encoder, and the entire architecture is trained through multi-task learning. The designed training tasks include:

[0090] Direct neighbor prediction of GO terms, where direct neighbors are defined as the union of subterms and parent terms;

[0091] Sub-ontology member prediction of GO terms, where sub-ontology categories include CC (cellular component), MF (molecular function), and BP (biological process).

[0092] Both use cross-entropy as the loss function and are jointly optimized by adding different classification heads.

[0093] The specific steps for constructing the GO term aggregator are as follows:

[0094] In this method, protein representations are created through GO term embeddings, which are obtained through the GO term encoder. First, the GO annotations of each protein are retrieved from the Uniprot database to construct a GO term set , where corresponds to a single GO term related to the protein. Then, the trained GO term encoder is used to infer the embeddings of each item in the set S and assemble them to construct a comprehensive embedding .

[0095] The specificity (importance) of GO terms is affected by several factors, especially the cellular processes related to protein interactions. Therefore, E is further input into the feature fusion block to generate protein features in the context of its functional annotation terms.

[0096] Fourth step, construction of the fusion block

[0097] The construction of the fusion block is based on the standard Transformer encoder architecture and enhanced by the SwiGLU feed-forward network. Figure 2 As shown, it mainly includes a multi-head self-attention layer (Multi-Head Attention), a SwiGLU feed-forward network (SwiGLU FFN) layer, and a fully connected layer (FC). In addition, layer normalization (LN) and residual connections are applied between each layer. This method reuses this fusion block in the regulator model, protein model, and subsequent prediction models to capture the complex interactions and dependencies between input features.

[0098] Fifth step, construction of the prediction model

[0099] This model uses a two-layer perceptron to perform the PPIMI prediction task. First, the regulator feature vector output by the regulator model is concatenated with the chaperone protein feature vector output by the protein model to form a joint feature vector. Then, this joint feature vector is input into the fusion block, and through the processing of the fusion block, a feature vector that can characterize their interaction is generated. Next, this interaction feature vector is input into the two-layer perceptron for further processing and transformation, and finally the prediction result is obtained. During the training process, binary cross-entropy is used as the loss function. Figure 3 The overall training process of this method is summarized. During this training process, the structure encoder of the regulator model is introduced, the text encoder of the regulator model is further trained, and the structure encoder is trained synchronously; the GO term aggregator of the protein model is introduced, the GO term encoder of the protein model no longer participates in the training, and the GO term aggregator is trained synchronously.

[0100] Finally, the performance of this method is evaluated, and this method is compared with four currently state-of-the-art methods on the benchmark dataset. As shown in Table 1, for the transduction setting, all methods perform excellently, and the AUROC and AUPR values exceed 0.9. Nevertheless, this method (KEPPIMI) still achieves further improvement, with AUROC and AUPR reaching as high as 0.99 and 0.989 respectively. For the inductive setting, this method has made substantial improvements. Compared with the second-best MultiPPIMI, the AUROC of S3 and S4 has increased by 4.3% and 6.8% respectively.

[0101] Table 1 Performance evaluation of this method and other methods under different experimental settings

[0102]

[0103] By applying the two pre-trained components in the model, the pre-trained text encoder in the regulator model and the pre-trained GO term encoder in the protein model, to the relevant downstream tasks (potency prediction and PPI prediction tasks) respectively, their generalization abilities are demonstrated. The comparison results are shown in Tables 2 and 3. It can be observed that on 7 out of 9 PPI families, the pre-trained component (KEPPIMI-MTE) in the regulator model shows better performance than pdCSM-PPI. This result verifies its accuracy and generalization ability, and emphasizes its great potential to facilitate and accelerate the discovery of PPI regulators. It can also be observed that the pre-trained component (KEPPIMI-GOE) in the protein model is superior to all benchmarks in the BFS and DFS settings, demonstrating its ability to understand protein properties by integrating biological knowledge, which is crucial for robust applications in the real world.

[0104] Table 2 Performance of the pre-trained components in the regulator model for the potency prediction task

[0105]

[0106] Table 3 Performance of the pre-trained components in the protein model for the PPI prediction task

[0107]

[0108] Furthermore, the interpretability of the proposed method is demonstrated by using the case of the Caspase-9 / XIAP interaction target and its regulator C17H20N4O2S. As Figure 4 shown in A of Figure 4 , some GO terms obtained higher attention scores and are presented in a vertical bar pattern. For the Caspase-9 protein, the two most important terms are GO:0008047 (enzyme activator activity) and GO:0097153 (cysteine-type endopeptidase activity involved in apoptotic process). For the XIAP protein, the two top-ranked terms are GO:0004869 (cysteine-type endopeptidase inhibitor activity) and GO:0043027 (cysteine-type endopeptidase inhibitor activity involved in apoptotic process). Due to space limitations, not all GO term names are shown in the figure, and the scale shows every few GO term names, such as the two GO terms GO:0097153 and GO:0004869 are not shown in the figure. Obviously, the key terms determined by the proposed method are closely related to the functional characteristics of the Caspase9 / XIAP interaction. In addition, the distribution and importance of different types of GO terms among all 117 PPI targets in the S1 scenario were analyzed. As Figure 4 shown in B of 17 H 20 and C of Figure 4 , the BP and CC terms generally contribute more to the interaction prediction than the MF terms. Especially for the CC terms, although their background frequency is low, the CC terms have higher attention scores. This finding is in line with expectations, given that PPI is highly dependent on physical proximity by nature, and the cellular components of proteins are more important when modeling the interaction between two proteins. If two proteins are located in different cellular regions, spatial separation poses a significant obstacle to their interaction. Finally, the attention scores of C 17 H <000002> N4O2S were calculated in the prediction model, and different parts of the input text were colored based on these scores. As Figure 4As shown in D in [reference], on the left is the partial structure of alanine, in the middle is the partial structure of proline, and on the right is the aromatic group. The darker the colored atoms indicate higher attention values. It can be seen that most atoms in alanine, proline, and the aromatic group are assigned relatively high attention values. These residues are closely related to the actual binding coordination. In addition, this method emphasizes terms such as "lipophilicity", "functional groups", "aromatic", and "cyclization" in the text description, and these terms are consistent with the key elements of C 17 H 20 N4O2S (as shown in E in Figure 4 [reference]). This result indicates that this method can effectively understand the chemical properties and structures of molecules and selectively focus on functional groups.

[0109] Furthermore, this embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned protein interaction regulator prediction method.

[0110] In addition, this embodiment also provides a protein interaction regulator prediction system based on a knowledge-enhanced language model, as shown in Figure 5 [reference], including:

[0111] A regulator encoding unit 1, configured to execute the regulator model constructed by the above-mentioned protein interaction regulator prediction method based on a knowledge-enhanced language model;

[0112] A protein encoding unit 2, configured to execute the protein model constructed by the above-mentioned protein interaction regulator prediction method based on a knowledge-enhanced language model;

[0113] A fusion prediction unit 3, configured to execute the prediction model constructed by the above-mentioned protein interaction regulator prediction method based on a knowledge-enhanced language model;

[0114] An interaction unit 4, used to receive the SMILES sequence of the molecule to be predicted and the text description extended by a generative language model for the regulator encoding unit to process and output regulator features, receive the GO annotation of the protein target for the protein encoding unit to process and output protein features, and display the prediction result output by the fusion prediction unit based on the above-mentioned regulator features and protein features.

[0115] This solution predicts protein interaction regulators for the first time by integrating external knowledge from scientific literature, biological databases, and knowledge graphs, rather than just using biological sequences. It also proposes independent regulator models and protein representation models. The protein model performs semantic parsing on the molecular SMILES sequence and text description, and calibrates the structure of chemical properties using molecular fingerprints to ensure the reliability of chemical information. The regulator model injects the structured hierarchical relationship of gene ontology into the language model and jointly encodes it with the text annotation of proteins to enhance the context awareness of functional semantics. Both models can be successfully applied to other downstream tasks and have strong performance. In addition, by retraining the tokenizer of the RoBERTa model to adapt to chemical symbols and biological terms, and introducing a fusion module based on the attention mechanism to dynamically capture the interaction between regulators and proteins, high-precision and highly interpretable prediction are finally achieved. It has been verified that, compared with existing methods, this solution has a significant improvement in performance, verifying the effectiveness and superiority of the model.

[0116] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art to which the present invention pertains can make various modifications or supplements to the described specific embodiments or use similar methods for substitution, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for predicting protein interaction regulators based on a knowledge-enhanced language model, characterized in that, The method includes: Regulator model construction Collect the SMILES sequences and text descriptions of molecules, and expand the text descriptions through a generative language model to construct a pre-training corpus; Pre-train a text encoder based on the pre-training corpus; Introduce a fingerprint-based structure encoder, represent the molecular structure using multiple fingerprints, and use separate linear projection heads to map each fingerprint to the embedding space of the text encoder; Concatenate the text embeddings output by the text encoder and the multiple fingerprint embeddings output by the structure encoder to generate regulator features; Protein model construction Collect GO graphs and GO annotations, and use the BERT model as the backbone network; GO represents Gene Ontology, and the GO annotations include the names and definitions of corresponding GO terms; Each node of the GO graph represents a GO term, and the edges represent the relationships between GO terms; Train a graph embedding model based on the GO graph; Use the trained graph embedding model to generate structural features for each GO term node; The backbone network converts the names and definitions of GO terms into word embeddings, and the structural features of GO terms are injected into the backbone network through the [SK] token to output GO term embedding vectors; Aggregate multiple GO term embedding vectors to generate protein features; Prediction model construction It includes a fusion block and a two-layer perceptron. The fusion block takes the regulator features output by the regulator model and the protein features output by the protein model as inputs, outputs interaction features, and the two-layer perceptron outputs a prediction result based on the interaction features.

2. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 1, wherein For the regulator model, first pre-train the text encoder using the pre-training corpus, and then introduce the structure encoder during the downstream prediction task, and train the structure encoder and the text encoder simultaneously; For the protein model, first train a graph embedding model using the GO graph, then use the trained graph embedding model, as well as the GO graph and GO annotations to train a GO term encoder to output GO term embedding vectors containing GO term structural features. Finally, during the downstream prediction task, use the output of the trained GO term encoder to train a GO term aggregator; For the prediction model, use the pre-trained regulator model, protein model, and a pre-constructed PPI regulator prediction dataset for training.

3. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 2, wherein, Specifically, the structural features of GO terms are injected into the backbone network in the following way: Use a tokenizer to perform word embedding on the text sequence of GO annotations and define a special token [SK]; Use the structural features generated by the graph embedding model to replace the embedding of [SK] for encoding and output GO term embedding vectors; Train the GO term encoder through multi-task learning; And the training tasks include: Direct neighbor prediction of GO terms, where direct neighbors are defined as the union of sub-terms and parent terms; Sub-ontology member prediction of GO terms, where the sub-ontology categories include cellular components, molecular functions, and BP biological processes.

4. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 2, wherein The regulator model has a fusion block, and after concatenating the text embedding and multiple fingerprint embeddings, they are input into the fusion block to generate the regulator features; The described protein model has a fusion block, and the protein embeddings output by the GO term aggregator are input into the fusion block to generate the protein features; The described regulator model, protein model, and prediction model reuse the same fusion block.

5. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 4, wherein The construction process of the GO term aggregator includes: Retrieve the GO annotations for each protein and construct a set of GO terms , where corresponds to a single GO term associated with the protein; Using the trained GO term encoder to output GO term embedding vectors for each item in the set S; Construct protein embeddings by combining all GO term embedding vectors ; The protein embedding E is fed into the fusion block to generate the protein features.

6. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 4, wherein The described fusion block is constructed based on the Transformer encoder architecture, including a multi-head self-attention layer, a SwiGLU feed-forward network layer, and a fully connected layer, and layer normalization and residual connections are used between the layers.

7. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 1, wherein The described regulator model uses the RoBERTa model as the text encoder, retrains the tokenizer of the RoBERTa model using byte-level byte pair encoding, pre-trains the RoBERTa model through a masked language modeling task, and calibrates the molecular structure representation using a fingerprint-based structure encoder; The described structure encoder represents the molecular structure using extended connectivity fingerprints, MACCS fingerprints, atom pair fingerprints, and topological fingerprints.

8. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 1, wherein The described protein model uses struc2vec as the graph embedding model, generates embedding vectors for each GO term node using the trained struc2vec, and projects the embedding vectors into the text space of the BERT model using a linear transformation and then injects them into the BERT model as structural features.

9. A protein interaction regulator prediction system based on a knowledge-enhanced language model, characterized in that, Including: A regulator encoding unit configured to execute the regulator model constructed by the protein interaction regulator prediction method based on the knowledge-enhanced language model according to any one of claims 1-8; A protein encoding unit configured to execute the protein model constructed by the protein interaction regulator prediction method based on the knowledge-enhanced language model according to any one of claims 1-8; A fusion prediction unit configured to execute the prediction model constructed by the protein interaction regulator prediction method based on the knowledge-enhanced language model according to any one of claims 1-8; An interaction unit for receiving the SMILES sequence of the molecule to be predicted and the text description extended by the generative language model for the regulator encoding unit to process and output regulator features, receiving the GO annotations of the protein target for the protein encoding unit to process and output protein features, and displaying the prediction results output by the fusion prediction unit based on the regulator features and protein features.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The described computer program, when executed by a processor, implements the protein interaction regulator prediction method according to any one of claims 1 to 8.