Protein interaction regulator prediction method and system based on knowledge enhanced language model, and computer readable storage medium
Through the method based on knowledge-enhanced language model, integrating molecular structure information, text description and gene ontology knowledge, and constructing a protein interaction regulator prediction model, solving the problem that existing methods cannot effectively utilize external knowledge and achieving more accurate and explainable prediction effects.
Patent Information
- Application Number
- CN202510464934.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing methods of protein interaction regulator prediction cannot effectively utilize external knowledge, resulting in low prediction performance and inability to elucidate the interaction mechanism, limiting the precise design and optimization of regulators.
Using a knowledge-enhanced language model method, a more comprehensive protein interaction characteristics are generated by constructing regulator models and protein models, molecular structure information, text descriptions and gene ontology knowledge are integrated to generate more comprehensive protein interaction characteristics.
More precise predictions of protein interaction regulators are achieved, capable of capturing functional mechanisms rather than just sequence or structural similarities, improving prediction accuracy and interpretability.
Smart Images

Figure CN119993283A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of bioinformatics, and in particular relates to a protein interaction regulator prediction method, system and computer-readable storage medium based on a knowledge-enhanced language model. Background Art
[0002] The regulation of protein-protein interactions (PPIs) plays a key role in disease treatment and drug development. Accurately predicting small molecule compounds (modulators) that can regulate specific PPIs has become one of the core challenges of computational biology and drug design. Existing protein-protein interaction modulator prediction methods include target-free methods (using only modulator features) and target-based methods (using both modulator features and protein-protein interaction features).
[0003] Target-free methods cannot elucidate the interaction mechanism, which limits the precise design and optimization of regulators. Target-based methods need to be further enhanced in terms of biomolecule representation. Existing methods predict protein interaction regulators based only on molecular or protein sequences, often ignoring rich external knowledge. In the biomedical field, scientific literature, biological databases, and knowledge graphs are important sources of knowledge, providing detailed descriptions of the properties, functions, and interactions of various biomolecules, which cannot be directly inferred from molecular or protein sequences alone. For example, the sequences of two proteins may be highly similar, but if they are annotated as "nucleus" and "mitochondria" localization, respectively, the probability of actual interaction is extremely low, and the existing methods lack explicit modeling of such knowledge, resulting in low prediction performance. Summary of the invention
[0004] The purpose of the present invention is to propose a protein interaction regulator prediction method, system and computer-readable storage medium based on a knowledge-enhanced language model to address the problems existing in the prior art.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: A method for predicting protein interaction regulators based on a knowledge-enhanced language model, the method comprising: Modulator model construction Collecting SMILES sequences and text descriptions of molecules, and expanding the text descriptions through a generative language model to build a pre-trained corpus; Pre-training a text encoder based on the pre-training corpus; Introducing a fingerprint-based structural encoder, which uses multiple fingerprints to represent molecular structures and uses a separate linear projection head to map each fingerprint to the embedding space of the text encoder; Concatenate the text embedding output by the text encoder and the multiple fingerprint embedding output by the structure encoder to generate the modulator features; Protein model building Collect GO graphs and GO annotations, and use the BERT model as the backbone network; GO stands for Gene Ontology, and the GO annotation includes the names and definitions of corresponding GO terms; Each node of the GO graph represents a GO term, and the edge represents the relationship between the GO terms; Training a graph embedding model based on the GO graph; Use the trained graph embedding model to generate structural features for each GO term node; The backbone network converts the names and definitions of GO terms into word embeddings. The structural features of GO terms are injected into the backbone network through [SK] tokens, and the GO term embedding vectors are output; Multiple GO term embedding vectors were aggregated to generate protein features; Prediction model building It includes a fusion block and a double-layer perceptron. The fusion block takes the regulator features output by the regulator model and the protein features output by the protein model as inputs and outputs interaction features. The double-layer perceptron outputs prediction results based on the interaction features.
[0006] In the above-mentioned protein interaction regulator prediction method based on knowledge-enhanced language model, for the regulator model, the text encoder is first pre-trained using the pre-training corpus, and then the structure encoder is introduced in the downstream prediction task, and the structure encoder and the text encoder are trained simultaneously; For protein models, the GO graph is first used to train the graph embedding model. Then, the trained graph embedding model, GO graph, and GO annotations are used to train the GO term encoder to output GO term embedding vectors containing GO term structural features. Finally, in the downstream prediction task, the GO term embedding vectors output by the trained GO term encoder are used to train the GO term aggregator. For the prediction model, pre-trained regulator models, protein models, and previously constructed PPI regulator datasets were used for training.
[0007] It should be noted that the regulator model and protein model here are not the entire models that are pre-trained. The text encoder in the regulator model is pre-trained, and the GO term encoder in the protein model is pre-trained.
[0008] In the above-mentioned protein interaction regulator prediction method based on knowledge-enhanced language model, the structural features of GO terms are injected into the backbone network in the following way: Use the tokenizer to embed the text sequence annotated by GO into words and define a special token [SK]; Use the structural features generated by the graph embedding model to replace the embedding of [SK] for encoding and output the GO term embedding vector; Training the GO term encoder via multi-task learning; The training tasks include: Direct neighbor prediction of GO terms, where direct neighbors are defined as the union of the child term and the parent term; Sub-ontology membership prediction of GO terms, where sub-ontology categories include cellular component, molecular function, and BP biological process.
[0009] In the above-mentioned protein interaction regulator prediction method based on knowledge-enhanced language model, the regulator model has a fusion block, and the text embedding and multiple fingerprint embeddings are spliced and input into the fusion block to generate the regulator features; The protein model has a fusion block, and the protein embedding output by the GO term aggregator is input into the fusion block to generate the protein features; The regulator model, protein model and prediction model reuse the same fusion block.
[0010] In the above-mentioned protein interaction regulator prediction method based on knowledge-enhanced language model, the construction process of GO term aggregator includes: Retrieve GO annotations for each protein and construct a GO term set ,in corresponds to a single GO term associated with a protein; Use the trained GO term encoder to output the GO term embedding vector for each item in the set S; Construct protein embedding by concatenating all GO term embedding vectors ; The protein embedding E is fed into the fusion block to generate the protein signature.
[0011] In the above-mentioned protein interaction regulator prediction method based on knowledge-enhanced language model, the fusion block is constructed based on the Transformer encoder architecture, including a multi-head self-attention layer, a SwiGLU feed-forward network layer and a fully connected layer, and layer normalization and residual connections are used between layers.
[0012] In the above-mentioned protein interaction regulator prediction method based on knowledge-enhanced language model, the regulator model uses RoBERTa model as the text encoder, and retrains the tagger of RoBERTa model using byte-level byte pair encoding, pre-trains RoBERTa model through masked language modeling task, and calibrates molecular structure representation using fingerprint-based structure encoder; The structure encoder uses extended connection fingerprint, MACCS fingerprint, atom pair fingerprint and topological fingerprint to represent the molecular structure.
[0013] In the above-mentioned protein interaction regulator prediction method based on knowledge-enhanced language model, the protein model uses struc2vec as a graph embedding model, and uses the trained struc2vec to generate an embedding vector for each GO term node. The embedded vector is projected into the text space of the BERT model using a linear transformation and then injected into the BERT model as a structural feature.
[0014] A protein interaction regulator prediction system based on a knowledge-enhanced language model, comprising: A regulator encoding unit, configured to execute the regulator model constructed by the protein interaction regulator prediction method based on the knowledge enhanced language model; A protein encoding unit, configured to execute the protein model constructed by the protein interaction regulator prediction method based on the knowledge-enhanced language model; A fusion prediction unit, configured to execute the prediction model constructed by the protein interaction regulator prediction method based on the knowledge enhanced language model; The interactive unit is used to receive the SMILES sequence of the molecule to be predicted and the text description expanded by the generative language model for processing by the regulator encoding unit and output the regulator feature, receive the GO annotation of the protein target for processing by the protein encoding unit and output the protein feature, and display the prediction result output by the fusion prediction unit based on the regulator feature and protein feature.
[0015] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the protein interaction regulator prediction method described in item 1.
[0016] The advantages of the present invention are: 1) The knowledge-enhanced protein interaction regulator prediction method proposed in this scheme constructs a holistic view of regulator-PPI interactions by integrating subtle expressions of natural language (functional description) with biomolecular structural properties (SMILES sequence, fingerprint). By integrating external knowledge, the model can capture functional mechanisms rather than just sequence or structural similarities. Experiments on Caspase-9 / XIAP interaction have proven that it can accurately clarify key roles; 2) This solution collects SMILES sequences and text descriptions from the PubChem database, enriches them through generative language models such as GPT-3.5, supplements chemical properties, improves the semantic density of the corpus, and establishes a large-scale small molecule pre-trained corpus covering a wide range of chemical space. This corpus effectively alleviates the performance degradation problem of traditional methods in few-sample scenarios by enhancing data diversity and semantic density; 3) The modulator representation model of this scheme adopts a combination of RoBERTa text encoder and molecular fingerprint structure encoder to capture the semantic association of SMILES sequence and text description and the precise chemical properties of molecular topology. By aligning the embedding space of the two modalities through linear projection, while ensuring semantic coherence, the errors that may be introduced by the generated text description are calibrated. It has been verified that the pre-trained text encoder in the modulator representation model performs well in cross-task transfer and surpasses existing methods in the efficacy prediction of 7 / 9 PPI families; 4) This scheme constructs a protein representation model based on gene ontology, which combines factual knowledge from GO graphs and GO annotations. The structural information of each GO term is injected into the term-based language model through a special tag and processed together with the text annotation; and a multi-task learning framework of direct neighbor prediction and sub-ontology classification is further designed to capture the logical dependencies between terms; in the experimental case of Caspase-9 / XIAP interaction, the key GO terms highlighted by the model are highly consistent with the functional mechanism verified by the experiment, which verifies its effectiveness; 5) This solution designs a universal feature fusion block based on the Transformer architecture. Through the multi-head self-attention mechanism and the SwiGLU feed-forward network, it captures the complex interactions and dependencies between input features, flexibly adapts to different input features, and is significantly improved over existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is the overall framework diagram of the protein interaction regulator prediction method based on the knowledge enhanced language model in this scheme; Figure 2 This is a schematic diagram of the structure of the fusion block in the protein interaction regulator prediction method based on the knowledge-enhanced language model in this scheme; Figure 3 The training process of the protein interaction regulator prediction method based on the knowledge-enhanced language model in this scheme; Figure 4 This is a study on the model interpretability of the embodiments provided in the protein interaction regulator prediction method based on the knowledge-enhanced language model in this scheme; Figure 5 This is the structural block diagram of the protein interaction regulator prediction system based on the knowledge enhanced language model in this scheme.
[0018] Figure numerals: regulator encoding unit 1; protein encoding unit 2; fusion prediction unit 3; interaction unit 4. DETAILED DESCRIPTION
[0019] This paper provides a method for predicting protein-protein interaction (PPI) regulators based on knowledge-enhanced language models, which integrates independent regulator and protein representation models. The expressive power of biomolecular properties is improved by combining external knowledge with the language modeling framework. Specifically: For the regulator model, SMILES sequences and text descriptions are first collected from the PubChem database, and these text descriptions are enriched by generative language models, such as the GPT-3.5 model to obtain more detailed chemical properties and functional group information to establish a large-scale corpus. This embodiment uses the GPT-3.5 model as an example. When it is put into use, other generative language models can also be used. Then, a Roberta model is pre-trained on this corpus and used as the backbone text encoder. Due to the lack of detailed descriptions of many regulator molecules in PubChem and the inaccuracy problems in the generation of large language models, this method additionally introduces a fingerprint-based structure encoder to calibrate the knowledge provided only by the GPT-3.5 model.
[0020] For protein models, this scheme combines factual knowledge from Gene Ontology (GO) graphs and GO annotations. The corresponding GO terms are contained in both the GO graph and the GO annotations. The structural information of each GO term is injected into the BERT-based language model through a special marker and processed together with the text annotations. In addition, direct neighbor prediction and sub-ontology member prediction tasks are designed to enhance the language model's understanding of the basic semantics and contextual relationships of the terms.
[0021] Finally, these two representation models are applied to the protein interaction regulator prediction task, where the regulator representation and protein representation are integrated through a carefully designed attention-based fusion block. The specific process is as follows Figure 1 As shown: The first step is to build a data set First, SMILES sequences and corresponding text descriptions of 301,566 molecules were collected from the PubChem database to build a pre-training corpus for the regulator model.
[0022] Next, GPT-3.5 is used to enrich the previously collected molecular text descriptions to obtain detailed information about chemical properties and functional groups. For example, the original text description of "ethanol" is "ethanol is a common organic solvent with antibacterial properties." The original SMILES sequence "CCO" of ethanol and its description "ethanol is a common organic solvent with antibacterial properties" are provided to GPT-3.5 to obtain an expanded text description: "Ethanol (Ethanol), chemical formula is , is a monohydric alcohol containing a hydroxyl (-OH) functional group. The molecule has significant hydrophilicity, with a logP value of -0.18. The molecular weight of ethanol is 46.07 g / mol, the boiling point is 78.37°C, and it is often used as a solvent in pharmaceutical preparations. Its hydroxyl functional group gives it weak acidity (pKa≈15.9), which allows it to participate in proton transfer reactions. The volatility and lipid solubility of ethanol enable it to penetrate cell membranes and destroy the lipid bilayer structure of microorganisms, thereby exerting a broad-spectrum antibacterial effect."
[0023] Then, the latest gene ontology data released in September 2024 were collected, and the GO terms marked as "obsolete" were deleted, resulting in 44,261 terms, including 8,888 "biological process (BP, describing biological activities (such as "cellular respiration", "DNA repair"))" terms, 11,177 "cellular component (CC, describing the cellular location of gene products (such as "mitochondria", "ribosome"))" terms, and 4,196 "molecular function (MF, describing molecular-level functions (such as "ATPase activity", "protein binding"))" terms. The set of all GO terms and inclusion relationships is organized into a directed acyclic graph, namely a GO graph, in which each node represents a GO term, the edge represents the relationship between the terms, and each term is labeled with its direct neighbors and sub-ontology members.
[0024] Finally, the interactions between small molecule modulators and PPI targets were extracted from the public database DLiP to construct a benchmark dataset for PPI modulator prediction. Using a filtering procedure consistent with existing methods, data entries for illegal PPI targets and non-human species were further removed. After this processing, a total of 11,145 modulator-PPI target interaction (PPIMI) tuples were generated, including 9,343 small molecule modulators and 117 PPI targets. These experimentally verified PPIMI tuples were regarded as positive samples. Accordingly, negative samples were generated by replacing the modulators or PPI targets in the positive samples. The ratio of positive and negative samples was controlled to 1:1, and it was ensured that the sampled negative samples would not appear in the positive samples for effective training.
[0025] Step 2: Construction of the regulator model The modulator model consists of two different submodules: a backbone text encoder and a structure encoder for knowledge calibration. They work together to produce a comprehensive and stable modulator representation.
[0026] The text encoder is implemented based on the RoBERTa architecture to process SMILES sequences and text descriptions of molecules. RoBERTa's tagger is retrained using Byte Level Byte-Pair Encoding to efficiently process SMILES strings and domain-specific vocabulary, adapting it to the needs of word segmentation of chemical symbols and biomedical terminology. The definition of special tokens is consistent with RoBERTa, and [SEP] is added as a precise boundary between SMILES sequences and text descriptions to help the model understand the input text. Then, the masked language modeling task is used to pre-train the text encoder based on the pre-training corpus enriched with text descriptions by GPT-3.5 built in the first step, where each token is randomly masked with a probability of 15% and reconstructed based on the context, resulting in a pre-trained text encoder.
[0027] As a general-purpose large-scale language model, GPT cannot avoid occasional inaccuracies, which will reduce the reliability of structure-activity analysis. To alleviate this problem, this scheme introduces a fingerprint-based structure encoder for knowledge calibration when using the modulator model for the downstream task of PPI modulator prediction. Specifically, multiple types of molecular fingerprints are used to represent various aspects of molecular structure. Including extended connectivity fingerprints (ECFPs), which are used to capture the characteristics of the local atomic environment in molecules; MACCS fingerprints, which encode predefined substructure patterns; atom pair fingerprints (APFPs), which represent atomic pairs and their topological distances in molecules; topological fingerprints (TFPs), which encode features based on molecular graphs. By using the structure encoder to synthesize the information contained in different fingerprints, the model is calibrated in downstream tasks, thereby improving the model's understanding and prediction capabilities of molecular structure.
[0028] Then, a separate linear projection head is used to align each type of molecular fingerprint with the embedding space of the RoBERTa text encoder to ensure the effective integration of structural and textual information. Finally, four fingerprint embeddings and one text embedding, a total of five molecular embeddings, are concatenated and fed into the fusion block to generate the final molecular representation as the modulator feature.
[0029] Step 3: Protein model construction The protein model consists of two stacked modules: 1) GO term encoder, which processes and encodes GO terms. It is built based on the BERT model and is responsible for converting the textual and structural information of GO terms into rich GO term embeddings; 2) GO term aggregator, which jointly models multiple GO term embedding vectors to generate protein representations.
[0030] The specific steps for constructing the GO term encoder are as follows: The BERT model is used as the backbone network of the protein model, and the structural information of the extracted GO terms is injected into its textual representation for subsequent joint modeling. The model weights are initialized using ouBioBERT, which has been pre-trained on large corpora in disciplines such as biology and medicine. The BERT model includes Figure 1 BERTEncoder and BERT Embeder in this paper, in order to inject structural embedding, the two parts are used separately.
[0031] The graph embedding model struc2vec is trained on the entire GO graph to extract the implicit contextual information in GO terms and their interconnections. The pre-trained struc2vec is used to generate embedding vectors for each term node, which represent the structural features of the GO terms and are injected into the backbone network as structural knowledge to enhance GO representation. To bridge the gap between GO structure and annotated text, the struc2vec embedding is projected into the text space of the BERT model using a linear transformation before injection.
[0032] For the text information of GO terms, the name and definition of each GO term are connected to form a text sequence that provides a detailed description of protein function. The aforementioned text sequence is processed using the WordPiece tokenization method, and special tags [CLS], [PAD], and [SEP] are added. In addition, a special token [SK] is defined as a placeholder, and the embedding of the special token [SK] will be replaced by the structural features encoded by struc2vec, so as to improve the GO representation in subsequent information propagation and inject structural knowledge into the GO representation framework. In the above way, the structural information and text information of the GO terms are combined in a specific format to form a comprehensive input sequence, which can be in the form of {[CLS][SK][SEP]Name + Def[SEP]}, where: [CLS] is a special marker that indicates the beginning of the entire sequence; [SK] is a specially defined token that acts as a placeholder to inject structural information of GO terms. During model processing, the embedding of [SK] will be replaced by the structural features encoded by struc2vec. [SEP] is a separator used to separate different parts, i.e., the text part used to separate the [SK]GO terms; Name+Def is a text sequence consisting of the name and definition of a GO term.
[0033] To enhance the joint representation of semantics and structure, different prediction heads are added on top of the GO term encoder, and the whole architecture is trained via multi-task learning. The designed training tasks include: Direct neighbor prediction of GO terms, where direct neighbors are defined as the union of the child term and the parent term; Sub-ontology membership prediction of GO terms, where sub-ontology categories include CC (cellular component), MF (molecular function), and BP (biological process).
[0034] Both use cross entropy as the loss function and are jointly optimized by adding different classification heads.
[0035] The specific steps of constructing the GO term aggregator are as follows: In this method, protein representation is created by GO term embedding, which is obtained by GO term encoder. First, the GO annotation of each protein is retrieved from Uniprot database to construct GO term set. ,in Corresponding to a single GO term related to the protein. Then, the trained GO term encoder is used to infer the embedding of each term in the set S and assemble them to construct the comprehensive embedding .
[0036] The specificity (importance) of GO terms is influenced by several factors, especially the cellular processes associated with protein interactions. Therefore, E was further input into a feature fusion block to generate protein features in the context of their functional annotation terms.
[0037] Step 4: Construction of fusion block The construction of the fusion block is based on the standard Transformer encoder architecture and enhanced by a SwigLU feed-forward network. Figure 2 As shown in Figure 1, it mainly contains a multi-head self-attention layer (Multi-Head Attention), a SwiGLU feed-forward network (SwiGLU FFN) layer, and a fully connected layer (FC). In addition, layer normalization (LN) and residual connections are applied between each layer. This method reuses this fusion block in regulator models, protein models, and subsequent prediction models to capture complex interactions and dependencies between input features.
[0038] Step 5: Construction of prediction model This model uses a two-layer perceptron to perform the PPIMI prediction task. First, the regulator feature vector output by the regulator model is concatenated with the partner protein feature vector output by the protein model to form a joint feature vector. Then, this joint feature vector is input into the fusion block, and the fusion block is used to generate a feature vector that can characterize the interaction between the two. Next, this interaction feature vector is input into the two-layer perceptron for further processing and transformation, and finally the prediction result is obtained. During the training process, binary cross entropy is used as the loss function. Figure 3 The overall training process of this method is summarized. During the training process, the structure encoder of the regulator model is introduced, the text encoder of the regulator model is further trained, and the structure encoder is trained synchronously; the GO term aggregator of the protein model is introduced, the GO term encoder of the protein model no longer participates in the training, and the GO term aggregator is trained synchronously.
[0039] Finally, the performance of our method is evaluated by comparing it with four state-of-the-art methods on the benchmark datasets. As shown in Table 1, for the transduction setting, all methods perform well, with AUROC and AUPR values exceeding 0.9. Nevertheless, our method (KEPPIMI) still achieves further improvements, with AUROC and AUPR as high as 0.99 and 0.989, respectively. For the induction setting, our method achieves substantial improvements, with AUROC improvements of 4.3% and 6.8% for S3 and S4, respectively, compared to the second-best MultiPPIMI.
[0040] Table 1 Performance evaluation of this method and other methods under different experimental settings
[0041] The generalization ability of the two pre-trained components in the model, the pre-trained text encoder in the regulator model and the pre-trained GO term encoder in the protein model, was demonstrated by applying them to relevant downstream tasks (potency prediction and PPI prediction tasks), respectively. The comparison results are shown in Tables 2 and 3. It can be observed that the pre-trained component in the regulator model (KEPPIMI-MTE) shows better performance than pdCSM-PPI on 7 out of 9 PPI families. This result verifies its accuracy and generalization ability and highlights its great potential to promote and accelerate the discovery of PPI regulators. It can also be observed that the pre-trained component in the protein model (KEPPIMI-GOE) outperforms all baselines in BFS and DFS settings, demonstrating its ability to understand protein properties by integrating biological knowledge, which is crucial for robust applications in the real world.
[0042] Table 2 Performance of pre-trained components in the modulator model in the efficacy prediction task
[0043] Table 3 Performance of pre-trained components in protein models in PPI prediction tasks
[0044] Furthermore, the interpretability of this method is demonstrated by using the case of Caspase-9 / XIAP interaction targets and their regulator C17H20N4O2S. Figure 4 As shown in A in , some GO terms received higher attention scores and are presented in a vertical line pattern. For Caspase-9 protein, the two most important terms are GO:0008047 (enzyme activator activity) and GO:0097153 (cysteine endopeptidase activity involved in the apoptotic process). For XIAP protein, the two highest-ranked terms are GO:0004869 (cysteine endopeptidase inhibitor activity) and GO:0043027 (cysteine endopeptidase inhibitor activity involved in the apoptotic process). Due to space limitations, the attached figure does not show all GO term names. The scale is displayed every few GO term names, such as GO:0097153 GO:0004869. The two GO terms are not reflected in the figure. Obviously, the key terms determined by this method are closely related to the functional characteristics of Caspase9 / XIAP interaction. In addition, the distribution and importance of different types of GO terms in all 117 PPI targets in the S1 scenario were analyzed. As shown Figure 4 B and Figure 4 As shown in C in Figure 1, BP and CC terms contribute more to interaction prediction than MF terms in general, especially CC terms, which have a higher attention score despite its lower background frequency. This finding is in line with expectations, given that PPIs are intrinsically highly dependent on physical proximity, and the cellular composition of proteins is more important when modeling the interaction between two proteins. If two proteins are located in different cellular regions, spatial separation poses a significant obstacle to their interaction. Finally, C is calculated in the prediction model. 17 H 20 N4O2S attention scores and colorize different parts of the input text based on these scores. Figure 4As shown in D in the figure, the left side is the partial structure of alanine, the middle is the partial structure of proline, and the right side is the aromatic group. The darker the atom, the higher the attention value. It can be seen that most atoms in alanine, proline, and aromatic groups are assigned high attention values. These residues are closely related to the actual binding coordination. In addition, this method emphasizes the terms "lipophilicity", "functional groups", "aromatic", and "cyclization" in the text description, which are related to C 17 H 20 The key elements of N4O2S are consistent ( Figure 4 This result indicates that the present method can effectively understand the chemical properties and structures of molecules and selectively focus on functional groups.
[0045] Furthermore, this embodiment also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned protein interaction regulator prediction method is implemented.
[0046] In addition, this embodiment also provides a protein interaction regulator prediction system based on a knowledge-enhanced language model, such as Figure 5 As shown, including: A regulator encoding unit 1 is configured to execute the regulator model constructed by the protein interaction regulator prediction method based on the knowledge enhanced language model; A protein encoding unit 2, configured to execute the protein model constructed by the protein interaction regulator prediction method based on the knowledge enhanced language model; A fusion prediction unit 3, configured to execute the prediction model constructed by the protein interaction regulator prediction method based on the knowledge enhanced language model; Interaction unit 4 is used to receive the SMILES sequence of the molecule to be predicted and the text description expanded by the generative language model for processing by the regulator encoding unit and output the regulator feature, receive the GO annotation of the protein target for processing by the protein encoding unit and output the protein feature, and display the prediction result output by the fusion prediction unit based on the regulator feature and protein feature.
[0047] For the first time, this scheme predicts protein interaction regulators by integrating external knowledge from scientific literature, biological databases, and knowledge graphs instead of just using biological sequences, and proposes independent regulator models and protein representation models. The protein model semantically parses the molecular SMILES sequence and text description, and uses molecular fingerprints to calibrate the structure of chemical properties to ensure the reliability of chemical information; the regulator model injects the structured hierarchical relationship of the gene ontology into the language model and jointly encodes it with the text annotation of the protein, thereby enhancing the contextual awareness of functional semantics. Both models can be successfully applied to other downstream tasks and have strong performance. In addition, by retraining the tagger of the RoBERTa model to adapt chemical symbols and biological terms, and introducing a fusion module based on the attention mechanism, the interaction between regulators and proteins is dynamically captured, and finally a high-precision and highly interpretable prediction is achieved. It has been verified that compared with existing methods, this scheme has significantly improved performance, verifying the effectiveness and superiority of the model.
[0048] The specific embodiments described herein are merely examples of the spirit of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in similar ways, but they will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.
Claims
1. A method for predicting protein interaction regulators based on a knowledge-enhanced language model, characterized in that: The method includes, Modulator model construction Collecting SMILES sequences and text descriptions of molecules, and expanding the text descriptions through a generative language model to build a pre-trained corpus; Pre-training a text encoder based on the pre-training corpus; Introducing a fingerprint-based structural encoder, which uses multiple fingerprints to represent molecular structures and uses a separate linear projection head to map each fingerprint to the embedding space of the text encoder; Concatenate the text embedding output by the text encoder and the multiple fingerprint embedding output by the structure encoder to generate the modulator features; Protein model building Collect GO graphs and GO annotations, and use the BERT model as the backbone network; GO stands for Gene Ontology, and the GO annotation includes the names and definitions of corresponding GO terms; Each node of the GO graph represents a GO term, and the edge represents the relationship between the GO terms; Training a graph embedding model based on the GO graph; Use the trained graph embedding model to generate structural features for each GO term node; The backbone network converts the name and definition of GO terms into word embeddings. The structural features of GO terms are injected into the backbone network through [SK] tokens, and the GO term embedding vector is output. Multiple GO term embedding vectors were aggregated to generate protein features; Prediction model building It includes a fusion block and a double-layer perceptron. The fusion block takes the regulator features output by the regulator model and the protein features output by the protein model as inputs and outputs interaction features. The double-layer perceptron outputs prediction results based on the interaction features.
2. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 1, characterized in that: For the regulator model, the text encoder is first pre-trained using the pre-training corpus, and then the structure encoder is introduced in the downstream prediction task, and the structure encoder and text encoder are trained at the same time; For protein models, the GO graph is first used to train a graph embedding model. Then, the trained graph embedding model, GO graph, and GO annotations are used to train a GO term encoder to output a GO term embedding vector containing the structural features of the GO terms. Finally, in the downstream prediction task, the output of the trained GO term encoder is used to train a GO term aggregator. For the prediction model, pre-trained regulator models, protein models and previously constructed PPI regulator prediction datasets were used for training.
3. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 2, characterized in that: Specifically, the structural features of GO terms are injected into the backbone network in the following way: Use the tokenizer to embed the text sequence annotated by GO into words and define a special token [SK]; Use the structural features generated by the graph embedding model to replace the embedding of [SK] for encoding and output the GO term embedding vector; Training the GO term encoder via multi-task learning; The training tasks include: Direct neighbor prediction of GO terms, where direct neighbors are defined as the union of the child term and the parent term; Sub-ontology membership prediction of GO terms, where sub-ontology categories include cellular component, molecular function, and BP biological process.
4. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 2, characterized in that: The modulator model has a fusion block, and the text embedding and multiple fingerprint embeddings are concatenated and input into the fusion block to generate the modulator feature; The protein model has a fusion block, and the protein embedding output by the GO term aggregator is input into the fusion block to generate the protein features; The regulator model, protein model and prediction model reuse the same fusion block.
5. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 4, characterized in that: The construction process of GO term aggregator includes: Retrieve GO annotations for each protein and construct a GO term set ,in corresponds to a single GO term associated with a protein; Use the trained GO term encoder to output the GO term embedding vector for each item in the set S; Construct protein embedding by concatenating all GO term embedding vectors ; The protein embedding E is fed into the fusion block to generate the protein signature.
6. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 4, characterized in that: The fusion block is built based on the Transformer encoder architecture, including a multi-head self-attention layer, a SwiGLU feed-forward network layer and a fully connected layer, and layer normalization and residual connections are used between layers.
7. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 1, characterized in that: The regulator model uses the RoBERTa model as the text encoder, retrains the RoBERTa model's tagger using byte-level byte pair encoding, pretrains the RoBERTa model through a masked language modeling task, and calibrates the molecular structure representation using a fingerprint-based structure encoder; The structure encoder uses extended connection fingerprint, MACCS fingerprint, atom pair fingerprint and topological fingerprint to represent the molecular structure.
8. The method for predicting protein interaction regulators based on a knowledge-enhanced language model according to claim 1, characterized in that: The protein model uses struc2vec as a graph embedding model, generates an embedding vector for each GO term node using the trained struc2vec, and injects the embedding vector into the BERT model as a structural feature after being projected into the text space of the BERT model using a linear transformation.
9. A protein interaction regulator prediction system based on a knowledge-enhanced language model, characterized in that: include: A regulator encoding unit, configured to execute the regulator model constructed by the protein interaction regulator prediction method based on the knowledge enhanced language model according to any one of claims 1 to 8; A protein encoding unit, configured to execute the protein model constructed by the method for predicting protein interaction regulators based on a knowledge-enhanced language model according to any one of claims 1 to 8; A fusion prediction unit, configured to execute the prediction model constructed by the protein interaction regulator prediction method based on the knowledge enhanced language model according to any one of claims 1 to 8; The interactive unit is used to receive the SMILES sequence of the molecule to be predicted and the text description expanded by the generative language model for processing by the regulator encoding unit and output the regulator feature, receive the GO annotation of the protein target for processing by the protein encoding unit and output the protein feature, and display the prediction result output by the fusion prediction unit based on the regulator feature and protein feature.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting a protein interaction regulator according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Novel hepatitis B virus capsid assembly regulator de novo design and virtual screening method based on generative model and computational chemistry
CN116504302A
Protein function prediction method and device based on topology perception attention network
CN118711672A
Perceptual representation learning method for protein conformations based on pre-trained language model
US20240136021A1