Enzyme-reaction relation prediction method based on representation learning

Through the multimodal fusion architecture's enzyme-reaction relationship prediction model, the problems of time-consuming, high cost and low accuracy in the prediction of enzyme-reaction relationship are solved, and a deep understanding of the functional inference of novel protein sequences and the mechanism of enzymatic reactions are achieved.

CN120431997APending Publication Date: 2025-08-05TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510495577.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The prior art has problems such as time-consuming, high cost, low accuracy and lack of interpretability in the prediction of enzyme-reaction relationships, especially for the functional inference of novel or unique protein sequences, which is difficult to effectively perform.

Method used

A multimodal fusion architecture is constructed using a method based on representation learning, integrating protein language models based on sequence and structure and biochemical reaction language models based on functional groups, and combining interpretability mechanisms to construct an enzyme-reaction relationship prediction model.

Benefits of technology

It reduces the time and cost of the prediction process, improves prediction accuracy, enables functional inference of novel or unique protein sequences, and has a deep understanding of the mechanisms and mechanisms of enzymatic reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431997A_ABST
    Figure CN120431997A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological prediction, and discloses an enzyme-reaction relationship prediction method based on representation learning, which comprises the following steps: acquiring and preprocessing related data of an enzyme-reaction relationship, and constructing an enzyme-reaction relationship database; a multi-modal fusion framework is adopted, a protein language model based on a sequence and a structure and a biochemical reaction language model improved based on functional groups are integrated, and an enzyme-reaction relation prediction model is constructed in combination with an interpretable mechanism; and outputting a binary classification prediction result corresponding to the simplified molecular linear input standard expression of the protein enzyme sequence and the biochemical reaction based on the enzyme-reaction relationship prediction model. According to the method, time consumption and cost in the prediction process can be effectively reduced, function inference on certain novel or unique protein sequences can be realized, in addition, the prediction accuracy can be effectively improved, and researchers are helped to understand the enzymatic reaction mechanism after prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological prediction technology, and in particular to an enzyme-reaction relationship prediction method based on representation learning. Background Art

[0002] Proteins play a vital role in organisms. They are not only the material basis for biological growth, heredity, and biological functions, but also play a core role in the entire life process. As the main executors of complex biochemical processes in organisms, proteins are involved in almost all cellular activities, including key processes such as energy metabolism, signal transduction, maintenance of cell structure, and immune response. Therefore, a deep understanding of the function of proteins is of great significance to many scientific fields such as molecular biology, biochemistry, genetics, and pharmacology. In the fields of research such as protein interaction mechanisms, the design of biosynthetic pathways, metabolic engineering, and biopharmaceuticals, the study of protein function is particularly critical because these fields are directly related to the understanding of life processes.

[0003] Enzymes are a special class of proteins whose primary function is to catalyze biochemical reactions, accelerating them by lowering their activation energy. This catalytic action enables organisms to carry out complex chemical changes with greater efficiency and control. Therefore, accurately predicting the relationship between enzymes and the reactions they catalyze is crucial for uncovering the specific mechanisms of the thousands of enzymatic reactions in living systems. The rapid advancement of sequencing technology in recent years has enabled the acquisition of protein sequence data at an unprecedented speed and scale. However, this rapid growth in data also presents new challenges. Traditional biological experimental methods are unable to cope with this massive amount of data, making it difficult to effectively and timely annotate its functions. This has led to a significant gap between protein sequence and structural information and its functional information.

[0004] To address these challenges, scientists are seeking new methods to accelerate the process of annotating protein function. This includes developing more efficient experimental techniques and using computational methods to predict protein function. Computational methods are particularly attractive because they can quickly process large amounts of data and can reveal hidden relationships between protein sequence and function. Through these methods, protein sequence data can be converted into useful biological information more quickly, thereby promoting the development of scientific research and applications. Enzyme reaction relationship prediction is an important part of this effort. By predicting the relationship between an enzyme and the specific reaction it catalyzes, we can gain a deeper understanding of the enzyme's function and discover new catalytic reactions, which are of great significance for developing new drugs, improving industrial processes, and solving environmental problems. Therefore, enzyme reaction relationship prediction is not only a challenging scientific problem, but also has important application value.

[0005] Although the study of enzyme-reaction relationship prediction has a long history, current methods for predicting these relationships still face several challenges:

[0006] (1) Methods based on biological experiments often require a lot of time, manpower, and resources to predict enzyme-reaction relationships, which makes the research process time-consuming and expensive.

[0007] (2) Sequence alignment-based enzyme-reaction relationship prediction methods mainly infer the potential function of unknown protein sequences by comparing the similarity between them and protein sequences with known functions in the database. For those novel or unique protein sequences, due to the lack of comparable known sequences, sequence alignment-based methods may not be able to provide effective functional inference.

[0008] (3) Artificial intelligence-based methods are divided into two categories. One category uses EC numbers as a bridge for predicting enzyme-reaction relationships. However, the accuracy of EC number predictions still needs to be improved, and the many-to-many relationships between enzymes and EC numbers, and between EC numbers and reactions, may lead to a large number of false positives. The other category directly predicts enzyme-reaction relationships, such as ESP. Although the ESP model can accurately predict enzyme-substrate pairs, due to the lack of interpretability, researchers often cannot fully understand the enzymatic reaction mechanisms and mechanisms behind these predictions.

[0009] Therefore, how to provide an enzyme-reaction relationship prediction method based on representation learning is an urgent problem to be solved. Summary of the Invention

[0010] The embodiment of the present invention provides an enzyme-reaction relationship prediction method based on representation learning to solve the above-mentioned technical problems in the prior art.

[0011] According to an embodiment of the present invention, a method for predicting enzyme-reaction relationships based on representation learning is provided, comprising:

[0012] Based on a preset biological database, relevant data of enzyme-reaction relationships are obtained and preprocessed, and an enzyme-reaction relationship database is constructed based on the preprocessed relevant data of enzyme-reaction relationships;

[0013] A multimodal fusion architecture is used to integrate sequence- and structure-based protein language models and functional group-based improved biochemical reaction language models, and an enzyme-reaction relationship prediction model is constructed in conjunction with an interpretability mechanism.

[0014] Based on the enzyme-reaction relationship prediction model, the binary prediction results corresponding to the simplified molecular linear input canonical expression of protein enzyme sequence and biochemical reaction are output.

[0015] In one embodiment, based on a preset biological database, obtaining and preprocessing relevant data on enzyme-reaction relationships, and constructing an enzyme-reaction relationship database based on the preprocessed relevant data on enzyme-reaction relationships includes:

[0016] Based on biological database 1, extract the reaction information of the compounds involved in the reaction and the related enzymes;

[0017] According to the compound identifiers obtained in the biological database 1, the corresponding compound information is extracted from the biological database 2, and the compound information is associated with the reaction information;

[0018] According to the enzyme identifier provided by biological database 1, the sequence information of the relevant enzyme is extracted from biological database 3, and the structural information is obtained from biological database 4 for matching, and the sequence information, functional information and structural information of the enzyme are integrated;

[0019] The data extracted from biological database 1, biological database 2, biological database 3 and biological database 4 were integrated and cleaned and formatted to obtain the enzyme-reaction relationship database.

[0020] In one embodiment, a multimodal fusion architecture is used to integrate a sequence- and structure-based protein language model and a functional group-based improved biochemical reaction language model, and an interpretability mechanism is combined to construct an enzyme-reaction relationship prediction model, including:

[0021] According to the functional groups of biochemical reactions, the simplified molecular linear input specification representation of the reaction is divided. Combined with the word elements defined by functional groups and the corresponding decomposition tools, an improved biochemical reaction language model based on functional groups is constructed.

[0022] Using pre-training and fine-tuning strategies, combined with protein sequences represented in the form of amino acid feature sequences, a sequence-based protein language model is constructed;

[0023] Based on Foldseek 3D structure information encoding and protein language model, a structure-based protein language model was constructed;

[0024] Using a multimodal feature fusion strategy, we integrated the functional group-based improved biochemical reaction language model, the sequence-based protein language model, and the structure-based protein language model to obtain a matrix containing enzyme-reaction relationships, and constructed a classifier for enzyme-reaction relationship prediction.

[0025] In one embodiment, a simplified molecular linear input specification representation of a biochemical reaction is divided according to the functional groups of the reaction, and a functional group-defined word unit and corresponding decomposition tool are combined to construct an improved biochemical reaction language model based on functional groups, including:

[0026] Obtain atom type tables from core databases in the field of bioinformatics and modify them using the SMARTS format to ensure that each functional group can be uniquely identified, thus obtaining a functional group attribute table based on the SMARTS format.

[0027] Get the functional group list represented by the simplified molecular linear input specification according to the functional group attribute table based on the SMARTS format, and split the simplified molecular linear input specification string of the reaction into separate compound lists;

[0028] Split and simplify the data of the functional group list and compound list represented by the molecular linear input specification, and use chemical information analysis tools to divide the rows into functional groups to obtain a vocabulary list based on functional groups;

[0029] The reaction data of the molecular linear input canonical representation is simplified based on the functional group list division to obtain several words, and the embedding representation of several words is input into the hidden layer of the RoBERTa model to obtain the vector representation corresponding to each word, thereby generating a vector representation of the biochemical reaction.

[0030] In one embodiment, the data of the functional group list and compound list represented by the simplified molecular linear input specification is split and the functional groups are divided using a chemical information analysis tool to obtain a vocabulary list based on functional groups including:

[0031] Based on the first function in the chemical information analysis tool, each item in the data of the functional group list and the compound list represented by the simplified molecular linear input specification is converted into a molecular object to obtain a molecular object list of the functional group and the compound;

[0032] Using a second function in the chemical information analysis tool, identify the common substructures between each molecule in the molecular object list of the compound and the functional group in the molecular object list of the functional group, and generate a list of atomic numbers matching each compound with the functional group;

[0033] Traverse the matching atomic number list, find the chemical bonds of each atomic number list item in the corresponding compound and the chemical bonds of adjacent atoms in each substructure, and save the chemical bond information separately;

[0034] By calculating the difference set of chemical bonds, the chemical bonds that need to be broken are determined, and the chemical bonds in the difference set are broken using the third function in the chemical information analysis tool, and the break positions are marked. The results are converted into a simplified molecular linear input standard format;

[0035] The simplified molecular linear input canonical string is split according to the labeled break positions to obtain a functional group-based vocabulary list.

[0036] In one embodiment, the partitioning expression for the reaction data of the simplified molecular linear input canonical representation is:

[0037] X=[X1,X2,...X N ]

[0038] The expression of word embedding is:

[0039] V i =W(X i )+S+P(i),V i ∈R d

[0040] The expression of the vector representation corresponding to the word unit is:

[0041] Z=[Z1,Z2,...,Z N ]

[0042] In the formula, X represents word unit, X i Represents each word after division, N represents the number of words, V i Represents the embedding representation of the word, W(X i ) represents the word embedding of the word unit, S represents the sentence embedding, P(i) represents the position embedding, R represents the real number set, d represents the embedding dimension, Z represents the vector representation corresponding to the word unit, Z i Represents the vector representation of the i-th word after being processed by the RoBERTa model.

[0043] In one embodiment, constructing a structure-based protein language model based on Foldseek 3D structure information encoding and a protein language model includes:

[0044] In the Foldseek 3D structure information encoding stage, the acquired features are used as input to the encoder, and the encoder is used to convert discrete features into continuous features;

[0045] The vectorization process maps the continuous representation of the latent space to the nearest center of mass, obtaining discretized descriptors of several protein conformations;

[0046] The decoder uses the vectorized results to reconstruct the input data and predict the probability distribution of the descriptors of the residues that are structurally aligned with each residue to obtain the structural label of each residue;

[0047] Given a protein sequence, the structural tag information of Foldseek is integrated to represent the sequence with structural tags; each residue is regarded as a word unit, and a mask signal is introduced to form a perception table;

[0048] The sequence is masked, and a predetermined proportion of residues are selected to replace with mask marks to obtain a masked pattern sequence; the masked pattern sequence is converted into an embedding vector sequence using an embedding layer, and the embedding vector sequence is input into the BERT model to obtain a protein embedding matrix containing structural information.

[0049] In one embodiment, the encoder function is expressed as:

[0050]

[0051] The expression for the protein conformation class of a residue is:

[0052] F=(f1,f2,...,f n )

[0053] The expression for a sequence with a structure tag is:

[0054] P=[P1,P2,...P N ]

[0055] The expression of the embedding vector sequence is:

[0056] E=[e(p1'),e(p'2),...e(p' n )]

[0057] The expression of the protein embedding matrix is:

[0058] M P =BERT(E),M P ∈R l×D

[0059] Where Enc represents the encoder function, x n represents the input of the encoder, z n Representation is a continuous representation of the latent space, D z represents the dimension of the latent space, F represents the protein conformation category of the residue, and f i represents the structural label of the i-th residue, P represents the protein sequence, and p i represents the residue at position i, E represents the embedding vector sequence, e(.) represents the embedding function, and p n ' represents the sequence of the mask pattern of the nth residue, M P represents the protein embedding matrix, BERT represents the BERT model, l represents the length of the protein sequence with structural information, and D represents the final protein vector representation dimension.

[0060] In one embodiment, a multimodal feature fusion strategy is used to integrate a functional group-based improved biochemical reaction language model, a sequence-based protein language model, and a structure-based protein language model to obtain a matrix containing enzyme-reaction relationships, and a classifier for enzyme-reaction relationship prediction is constructed, including:

[0061] Obtain protein sequence vectors and biochemical reaction vectors, and capture the complex relationship between biochemical reactions and protein sequences based on linear transformation, vector interaction, and enzyme-reaction identifier construction technology to obtain a matrix containing enzyme-reaction relationships;

[0062] Based on the matrix containing the enzyme-reaction relationship, a classifier for predicting the enzyme-reaction relationship is constructed, wherein the classifier consists of a long short-term memory network layer, a multi-layer perceptron layer and a softmax.

[0063] In one embodiment, the linear transformation includes linearly transforming the biochemical reaction vector to match the dimension of the protein sequence vector;

[0064] Vector interaction involves calculating the interaction matrix based on the protein sequence vector and the converted biochemical reaction vector to capture the interaction information between the chemical reaction and the protein sequence;

[0065] The construction of the enzyme-reaction identifier involves taking the interaction matrix as the input of the Transformer encoder and using the Transformer encoder to capture the fusion information to obtain a matrix containing the enzyme-reaction relationship.

[0066] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:

[0067] The present invention constructs a new deep learning algorithm for enzyme-reaction relationships based on representation learning. This method adopts a multimodal fusion architecture, integrates a protein language model based on sequence and structure, and a biochemical reaction language model based on functional group improvement, and adds an interpretability mechanism to explore the mechanism and mechanism of enzymatic reactions. It can deeply explore the complex relationship between enzymes and reactions. Compared with traditional technologies, the present invention can not only effectively reduce the time and cost in the prediction process, but also realize functional inference of certain novel or unique protein sequences. In addition, it can also effectively improve the prediction accuracy, which helps researchers understand the enzymatic reaction mechanism and mechanism behind the prediction.

[0068] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0070] Figure 1 is a flow chart of an enzyme-reaction relationship prediction method based on representation learning according to an exemplary embodiment;

[0071] Figure 2 is a schematic diagram of constructing a sample set according to an exemplary embodiment;

[0072] Figure 3 is a schematic diagram of a structure of an enzyme-reaction relationship prediction model according to an exemplary embodiment;

[0073] Figure 4 is a schematic diagram of protein representation based on the ESM-2 model according to an exemplary embodiment;

[0074] Figure 5 is a schematic diagram of a reaction comparison model architecture according to an exemplary embodiment;

[0075] Figure 6 is a schematic diagram illustrating DeepRP model performance according to an exemplary embodiment;

[0076] Figure 7 is a schematic diagram of a binary confusion matrix according to an exemplary embodiment;

[0077] Figure 8 is a schematic diagram of the true and predicted distributions of a binary classification test set according to an exemplary embodiment;

[0078] Figure 9 is a schematic diagram showing comparison of model performance according to an exemplary embodiment. DETAILED DESCRIPTION

[0079] The following description and the accompanying drawings sufficiently illustrate the specific embodiments of this invention to enable those skilled in the art to practice them. Portions and features of some embodiments may be included in or replace portions and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims, and all available equivalents of the claims.

[0080] Figure 1 An embodiment of an enzyme-reaction relationship prediction method based on representation learning of the present invention is shown.

[0081] In this optional embodiment, the enzyme-reaction relationship prediction method based on representation learning includes:

[0082] Step S1: Based on a preset biological database, relevant data of enzyme-reaction relationships are obtained and preprocessed, and an enzyme-reaction relationship database is constructed based on the preprocessed relevant data of enzyme-reaction relationships;

[0083] The step of obtaining and preprocessing enzyme-reaction relationship data based on a preset biological database, and constructing an enzyme-reaction relationship database based on the preprocessed enzyme-reaction relationship data includes:

[0084] Based on biological database 1 (i.e., Rhea database), reaction information of compounds involved in the reaction and related enzymes was extracted;

[0085] The Rhea database is a comprehensive, non-redundant resource of expertly curated biochemical reactions, detailing the interactions between enzymes and reactions, which is crucial for understanding biochemical pathways. Rhea's FTP site functionality supports efficient downloading of large datasets, enabling this example to quickly obtain extensive reaction details, such as reaction IDs, equations, ChEBI IDs of related compounds, and UniProt identifiers for proteins.

[0086] Rhea data preprocessing was performed as follows: First, two key data files were downloaded from the Rhea database: rhea_basic_infor.tsv and rhea2uniprot_sprot.tsv. The rhea_basic_infor.tsv file contains basic reaction information, such as the reaction ID (Reaction_id), chemical equation (Equation), ChEBI ID (Chebi_id), and ChEBI name (Chebi_name). The rhea2uniprot_sprot.tsv file provides information such as the reaction ID (Reaction_id), reaction direction (Direction), corresponding UniProt ID (Uniprot_id), and corresponding EC number (EC_number). Next, the data from these two files were merged using the reaction ID (Reaction_id) to form a more complete Rhea database dataset. During the merging process, special attention was paid to reversible reactions (labeled as UN Direction) and records with empty equations (Equation) were removed. Through this step, the Rhea dataset was finally obtained.

[0087] According to the compound identifiers obtained from the biological database 1, the corresponding compound information is extracted from the biological database 2 (i.e., the ChEBI database), and the compound information is associated with the reaction information;

[0088] The ChEBI (Chemical Entities of Biological Interest) database provides a detailed understanding of biochemical entities, ranging from simple small molecules to complex drugs. As an open resource maintained by the European Bioinformatics Institute (EBI), ChEBI provides information such as structure, biological function, and reactions in which compounds participate, which is crucial for understanding the role of proteins in different biological processes.

[0089] ChEBI data preprocessing is as follows: In order to integrate and optimize data resources, the ChEBI (Chemical Entities of Biological Interest) database was carefully preprocessed with the aim of integrating it with the data of the Rhea database. First, two files, compounds.tsv and structures.tsv, were downloaded from the ChEBI database, which contain basic information and structural data of the compounds, such as compound ID, ChEBI access number, SMILES representation and InChI code, respectively. Next, a data set containing Compounds_id, Chebi_accession, SMILES and InChI was screened and sorted out using a method similar to that used to process the Rhea database. Since the Chebi_accession field in the ChEBI database matches the Chebi_id field in the Rhea database, these two attributes are used as connection keys to integrate the information of the two databases. This process ultimately produces a comprehensive data set that contains both enzyme-reaction information in the Rhea database and detailed structures and properties of compounds in the ChEBI database, providing a rich and multi-dimensional source of information for this embodiment.

[0090] According to the enzyme identifier provided by biological database 1, the sequence information of the relevant enzyme was extracted from biological database 3 (i.e., UniProt database), and the structural information was obtained from biological database 4 (i.e., PDB database) for matching, and the sequence information, functional information and structural information of the enzyme were integrated;

[0091] The UniProt database provides detailed information on protein sequences and structures, which is crucial for understanding protein functions and their roles in biochemical reactions. Using the UniProt API, this example was able to obtain a large number of protein amino acid sequences and functional descriptions, particularly those involved in enzymatic reactions documented in the Rhea database, allowing for in-depth exploration of how these proteins influence or regulate specific biochemical pathways.

[0092] Uniprot data preprocessing is as follows: key data are extracted and integrated from the Rhea and ChEBI databases, and the next focus is to collect additional protease information related to these data. This step is implemented through Python's requests library, using UniProtID to send a request to the UniProt database to obtain supplementary information on the relevant proteases. The collected information covers key data such as protein sequence (seq). Afterwards, using UniProtID as a key, these newly obtained protease data are effectively integrated with the ChEBI data. Further, in order to enhance the depth of subsequent model interpretability analysis, this embodiment also downloaded the binding site (Activate_site) and active site (Binding_site) information of each protease from the UniProt database. The resulting Uniprot data not only lists the basic characteristics of each protease in detail, but also includes data on the various biochemical reactions in which they participate, laying a solid foundation for subsequent research work.

[0093] The Protein Data Bank (PDB) is a valuable resource for providing three-dimensional structural information on biological macromolecules, which is key to understanding protein function and designing bioengineering applications. The structure of each protein in the PDB has been precisely determined, providing important information about its interactions with other molecules. For proteins whose structures have not yet been determined in the PDB, this example used the structural information predicted by AlphaFold.

[0094] PDB data preprocessing is as follows: in the structural data processing stage, the present embodiment downloaded the pdb_seqres.txt file from the PDB (Protein Data Bank) database, which is an important basis for protein structure analysis. It contains key information such as PDBID (Pdb_id), molecule type (Mol), sequence (Seq) and sequence length (Length). These data are then used to match the protease sequences sorted in the PDB data, with the aim of obtaining the PDB structural information of a specific protease. Further, through Python's requests library, using Pdb_id as an index, the corresponding PDB file information is directly obtained from the PDB database and stored locally. This series of steps is summarized to form the PDB data, which includes the basic information of the protease and its extended structural data, providing a solid data foundation for deeper biochemical research.

[0095] The data extracted from biological database 1, biological database 2, biological database 3 and biological database 4 were integrated and cleaned and formatted to obtain the enzyme-reaction relationship database.

[0096] Although the structural information of a batch of proteases has been successfully obtained from the PDB database, many proteases lack experimentally verified crystal structures. To solve this problem, the structural information predicted by AlphaFold is used in this embodiment to make up for those proteases that lack structural information in the PDB database. The specific implementation strategy is to use Python's requests library, combined with the Uniport_id of each protease, to retrieve and download the corresponding protein structure prediction information from the AlphaFold database. Through this method, this embodiment can still conduct an in-depth analysis of the structural characteristics of proteases in the absence of experimental structural data.

[0097] By integrating data from various databases, the database created in this example serves as a comprehensive source of information for studying enzyme catalytic mechanisms, protein structural properties, and their roles in biochemical reactions. The establishment of this database provides a powerful impetus for in-depth understanding and advancement in protein science and enzyme engineering, laying a solid foundation for the subsequent research and practical application of enzyme-reaction relationship prediction.

[0098] Data Storage: Information is stored in Feather format files. Feather is renowned for its exceptional read / write speed and cross-platform compatibility, making it particularly well-suited for rapidly processing and transferring large datasets. To this end, a comprehensive table containing extensive reaction details, protein sequence information, and structural data was compiled and stored using the DataFrame.to_feather method in the Pandas library. Using the to_feather method in Pandas, a DataFrame can be efficiently written to a Feather format file.

[0099] In addition, in this embodiment, the reactions that the enzyme can catalyze are defined as positive samples, and these data are directly from the existing database. However, this leads to an important imbalance problem in the dataset: the lack of negative samples representing reactions that the enzyme cannot catalyze. In order to overcome this problem, a random sampling strategy is adopted in this embodiment to generate these missing negative samples. Specifically, a comprehensive reaction set and enzyme set are first obtained through the previous data processing. Figure 2As shown, if a certain reaction Ra can be catalyzed by enzymes Ea and Ec (this is information obtained from the database), then the combination of Ra with Ea and Ec is considered to be a positive sample for the reaction. For the generation of negative samples, other enzymes are randomly selected from the enzyme set that does not include Ea and Ec. For example, Eb and Ed randomly selected from the enzyme set will form a negative sample pair with Ra. In this embodiment, a positive-negative sample ratio of 1:1 is adopted, that is, for each reaction, if it has two positive sample enzymes, two enzymes that do not catalyze the reaction will also be randomly selected as negative samples. This method is intended to balance the training data and improve the generalization ability and prediction accuracy of the model when dealing with practical problems. Through the above processing steps, this embodiment successfully constructed a comprehensive dataset containing positive and negative samples. Each record in the dataset includes several key fields: Reaction_id, Uniprot_id, SMILES, EC_number and other information.

[0100] Step S2: Using a multimodal fusion architecture, integrating a sequence- and structure-based protein language model and a functional group-based improved biochemical reaction language model, and combining it with an interpretability mechanism to construct an enzyme-reaction relationship prediction model;

[0101] The multimodal fusion architecture integrates a sequence- and structure-based protein language model and a functional group-based improved biochemical reaction language model, and combines an interpretability mechanism to construct an enzyme-reaction relationship prediction model (DeepRP model), including:

[0102] 1) Based on the functional groups of biochemical reactions, a simplified molecular linear input canonical representation of the reaction is divided. Combined with the functional group-defined words and corresponding decomposition tools, an improved biochemical reaction language model based on functional groups is constructed;

[0103] Specifically, this embodiment divides the reactions into SMILES (Simplified Molecular Input Line Entry System) representations according to their functional groups, customizes tokens (words) and tokenizers (tokenizers are tools that decompose text or SMILES strings into tokens) based on the functional groups, and uses xml-roberta-large based on RoBERTa to represent biochemical reactions.

[0104] The simplified molecular linear input specification representation of the reaction divided according to the functional groups of the biochemical reaction, combined with the word elements defined by the functional groups and the corresponding decomposition tools, constructs a biochemical reaction language model based on the functional group improvement, including:

[0105] With the help of the open source chemical informatics software library, RDKIT, a tool for processing chemical data and performing chemical information analysis, provides functions for creating, manipulating, and analyzing molecular structures, and can also process data in SMILES format. In addition, the atom type table (i.e., the atomtype table) was obtained from the core database in the field of bioinformatics (i.e., the KEGG database) and modified using the SMARTS format to ensure that each functional group can be uniquely identified. The functional group attribute table based on the SMARTS format was obtained, containing information such as Group_atom (atom type), Group_name (atom name), Description (functional group description), and Group_pattern (functional group SMARTS format);

[0106] Obtain the functional group list represented by SMILES from the functional group attribute table based on the SMARTS format, i.e., the Group_pattern column, denoted as smarts_list. Split the reaction SMILES string into a separate compound list, called compounds_list.

[0107] Split and simplify the functional group list and compound list data represented by the linear input specification of the molecule, and use chemical information analysis tools to divide the rows into functional groups to obtain a vocabulary list based on functional groups. The main steps are as follows:

[0108] Based on the first function in the chemical information analysis tool (i.e., Chem.MolFromSmiles function), each item in smarts_list and compounds_list is converted into a molecule object (mol object), and the molecule object lists of functional groups and compounds (i.e., mol_smarts_list and mol_compounds_list) are obtained;

[0109] Use the second function in the chemical information analysis tool (i.e., GetStructMatches function) to identify the common substructures between each molecule in the molecular object list of the compound (mol_compounds_list) and the functional groups in the molecular object list of the functional groups (mol_smarts_list). The result of this step is a list of atomic numbers that match each compound and the functional group, which is recorded as matches.

[0110] After obtaining matches, traverse matches and find the chemical bonds of each match item in the corresponding compound. These bonds will be saved in the list bonds1. At the same time, traverse matches to obtain the chemical bonds of adjacent atoms in each substructure and save them in another list bonds2.

[0111] Define the difference between bonds1 and bonds2, denoted as df_bonds. This step is to determine the chemical bonds that need to be broken. Then, use the third function (i.e., FragmentOnBonds function) in the chemical information analysis tool to break the chemical bonds in df_bonds. The break position is indicated by an asterisk (*). Then, convert the result into SMILES format.

[0112] By splitting these SMILES strings with asterisks as delimiters, a functional group-based vocabulary list 1 was obtained (partial data is shown).

[0113] Table 1 Vocabulary based on functional groups

[0114]

[0115]

[0116] The reaction data of the molecular linear input canonical representation is simplified based on the functional group list division to obtain several words, and the embedding representation of several words is input into the hidden layer of the RoBERTa model to obtain the vector representation corresponding to each word, thereby generating a vector representation of the biochemical reaction.

[0117] Specifically, for reaction representation learning, the xml-robertqa-large model was selected. It is based on the Roberta pre-trained language model architecture. Roberta optimizes the BERT pre-training process to improve model performance. Strictly speaking, Roberta is the complete BERT. Roberta adopts a dynamic masking mechanism and improves the random mask application mechanism in the preprocessing stage. It abandons BERT's next sentence prediction task and introduces a larger byte-level BPE vocabulary for model training. In addition, RoBERTa uses a larger batch size during training, and has experimented with batch sizes ranging from 256 to 8000. Reaction representation is divided into the following steps:

[0118] In the RoBERTa-based chemical reaction characterization method, SMILES reaction data is partitioned based on the functional group list. Assuming the input SMILES is converted into a series of tokens after tokenizer processing, the partition expression of the reaction data of the simplified molecular linear input canonical representation is:

[0119] X=[X1,X2,...X N ]

[0120] In the formula, X represents word unit, X iRepresents each word after division, N represents the number of words; the embedding representation of each word is composed of word embedding, sentence embedding and position embedding. The embedding representation of the word is expressed as:

[0121] V i =W(X i )+S+P(i),V i ∈R d

[0122] Where V i Represents the embedding representation of the word, W(X i ) represents the word embedding of the word unit, S represents the sentence embedding, P(i) represents the position embedding, R represents the set of real numbers, and d represents the embedding dimension;

[0123] After that, V is input into the hidden layer of RoBERTa to get the output E. The expression of the vector representation corresponding to the word unit is:

[0124] Z=[Z1,Z2,...,Z N ]

[0125] In the formula, Z represents the vector representation corresponding to the word unit, Z i The vector representation of the i-th word after being processed by the RoBERTa model. Through the above processing, the vector representation of the biochemical reaction can be obtained.

[0126] 2) Using pre-training and fine-tuning strategies, combined with protein sequences represented as amino acid feature sequences, to construct a sequence-based protein language model;

[0127] In addition to the structure-based protein feature representation method, this embodiment also uses a sequence-based protein feature representation method, namely the ESM-2 model. Like other models, the ESM-2 model also adopts a strategy of combining pre-training models with fine-tuning. During the training process of ESM-2, the model receives a protein sequence represented in the form of an amino acid feature sequence. During training, the model randomly masks approximately 15% of the amino acids in the sequence. The core task of this training is to predict these randomly masked amino acid parts, so that the model can deeply understand the structure and function of the protein sequence. In this way, ESM-2 can learn patterns and regularities in protein sequences, thereby improving the ability to recognize and understand new sequences. As Figure 4 As shown in Figure 2, the ESM-2 model generates a vector matrix of size l*1280, where l represents the length of the protein sequence and 1280 represents the dimension of the feature vector generated by the ESM-2 model. Each feature vector provides a comprehensive representation of the amino acid at a specific position in the sequence, reflecting its surrounding environment and possible biological function.

[0128] 3) Constructing a structure-based protein language model SaProt based on Foldseek 3D structure information encoding and protein language model;

[0129] SaProt consists of two parts: Foldseek 3D structural information encoding and protein language model. When performing protein feature representation, this embodiment adopts a strategy of pre-training model combined with fine-tuning (PretrainModel & Fine-tuning). For SaProt and ESM-2 models, the strategy includes freezing the parameters of certain layers and fine-tuning other parameters. The main reason is that the primary layers of the pre-trained model usually capture more general and universal features, which are useful for a variety of tasks. Freezing these layers can retain the general features that have been learned while reducing the risk of overfitting the model on specific tasks.

[0130] Specifically, the construction of a structure-based protein language model based on Foldseek3D structure information encoding and a protein language model includes:

[0131] In the Foldseek 3D structure information encoding stage, the 10 features obtained are used as the input of VQ-VAE. As the input of the encoder, D x is the dimension of the input feature, which is 10 here. Then the discrete features are converted into continuous features. The expression of the encoder function is:

[0132]

[0133] Where Enc represents the encoder function, which takes the input x n Mapped to the latent space, x n represents the input of the encoder, z n Representation is a continuous representation of the latent space, D z represents the dimensionality of the latent space;

[0134] The vectorization process Q is to map the continuous representation of the latent space to the nearest centroid q n , i.e., discretized descriptors of 20 protein conformations (20 structural representations);

[0135] The decoder uses the vectorized result q n To reconstruct the input data x n , which is used to predict the residue y that is structurally aligned with each residue n The probability distribution of the descriptor of , the protein conformation category of the residue can be obtained through the output of the encoder. The conformation category of each residue can be represented by F; the expression of the protein conformation category of the residue is:

[0136] F=(f1,f2,...,f n )

[0137] Where F represents the protein conformation category of the residue, f i represents the structural label of the i-th residue;

[0138] Given a protein sequence P, its primary sequence can be represented as (s1, s2, …, sn), where si represents the residue at position i, and the structure tag information of Foldseek is integrated to represent the sequence with structure tags; the expression of the sequence with structure tags is:

[0139] P=[P1,P2,...,P N ]

[0140] Where P represents the protein sequence, p i represents the residue at position i, and each p i It is composed of (sifi), where si represents the i-th residue, and f i represents the structural label of the i-th residue;

[0141] Through the above processing, each p i As a token, “#” is introduced as a mask signal, where “si#” and “#fi” respectively indicate that only residue or structure tags are available. Finally, a 21×21 perception table (Sa) is formed, and p is used as the input of the language model;

[0142] Based on the above, we can obtain a sequence containing structural markers. Mask the sequence and select a certain proportion of residues to replace them with mask markers "si#" and "#fi" to obtain the mask pattern sequence p'. The embedding layer is used to convert the mask pattern sequence into an embedded vector sequence. The expression of the embedded vector sequence is:

[0143] E=[e(p1'),e(p'2),...e(p' n )]

[0144] Where E represents the embedding vector sequence, e(.) represents the embedding function, which converts the amino acid residues or masks carrying structural labels into n ′ represents the sequence of the mask pattern of the nth residue;

[0145] Input the embedding vector sequence into the BERT or ESM model. Through the above processing, the protein embedding matrix containing structural information can be obtained. The expression of the protein embedding matrix is:

[0146] M P =BERT(E),M P ∈R l×D

[0147] Where M P represents the protein embedding matrix, BERT represents the BERT model, l represents the length of the protein sequence with structural information, and D represents the final protein vector representation dimension, which is 446 here.

[0148] 4) Utilizing a multimodal feature fusion strategy, we integrated the functional group-based improved biochemical reaction language model, the sequence-based protein language model, and the structure-based protein language model to obtain a matrix containing enzyme-reaction relationships, and constructed a classifier for enzyme-reaction relationship prediction.

[0149] Specifically, by using the feature representations of proteases and biochemical reactions, vector representations of proteins and reactions were obtained, respectively. A multimodal feature fusion strategy was adopted, which aims to capture the complex relationship between chemical reactions and protein sequences and enhance the communication and interaction between them. Based on the previous protease and biochemical reaction feature extraction, the vector representation of chemical reactions was obtained to be 1024-dimensional, while the vector representation of protein sequences was 446-dimensional and 1028-dimensional. To align these two vectors and promote information exchange, this example performed the following processing (using the vector representation of ESM-2 for illustration):

[0150] Linear transformation: biochemical reaction vector S∈R lx1024 Perform a linear transformation, where l is the number of chemical reaction tokens, to match the dimension of the protein sequence vector. The transformed dimension is 1028, and the calculation formula is:

[0151] S'=S·W

[0152] Where W∈R 1024x1028 is the transformation matrix, S'∈R lx1028 ;

[0153] Vector interaction: According to the protein sequence vector P∈R mx446 and the transformed biochemical reaction vector S′ to calculate the interaction matrix M, where M∈R mxl Represents the interaction between chemical reactions and protein sequences to capture the interaction information between chemical reactions and protein sequences. The calculation formula is:

[0154] M=P·S' T

[0155] Where T represents transpose;

[0156] Enzyme-reaction identifier construction: The interaction matrix M is used as the input of the Transformer encoder, and the Transformer encoder is used to capture the fusion information to obtain the matrix H∈R containing the enzyme-reaction relationship mxl .

[0157] Based on the matrix containing the enzyme-reaction relationship, a classifier for predicting the enzyme-reaction relationship is constructed, wherein the classifier consists of a long short-term memory network layer, a multi-layer perceptron layer and a softmax.

[0158] After the above processing, a matrix containing the enzyme-reaction relationship is obtained, namely: H∈R mxl Next, we build a classifier to predict the enzyme-reaction relationship. This classifier consists of an LSTM layer, a multi-layer perceptron layer (MLP), and a softmax layer.

[0159] Long Short-Term Memory Network (LSTM) layer: After the above processing, the matrix H is obtained and used as the input of the LSTM layer. LSTM can further extract temporal patterns and key information by its ability to capture long-term dependencies. The calculation formula is as follows: lstm =LSTM(H);

[0160] Multi-layer Perceptron (MLP) layer: The LSTM output is then passed to the MLP for further processing. The MLP consists of several linear layers and nonlinear activation functions to extract more abstract and useful features from the LSTM output.

[0161] Softmax layer: Finally, the output of the model is classified through a softmax layer. The softmax layer converts the output of the MLP into a probability distribution representing the probability of different enzyme-reaction pairs.

[0162] The loss function used in this example is as follows: This example primarily predicts enzyme-reaction relationships, a binary classification task, so a loss function commonly used for multi-classification problems is used. This function calculates the cross entropy between the true distribution and the predicted distribution, using the following formula:

[0163] H(A,B)=∑A(x)logB(x)

[0164] Where A is the true distribution, B is the predicted distribution, Σ is the sum over all categories, x is a specific category, A(x) is the probability of category x in the true distribution, and B(x) is the probability of category x in the predicted distribution.

[0165] In this embodiment, in the classification model evaluation, accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score (F1 score) are used. Among them, accuracy is the most intuitive classification performance indicator, which is defined as the ratio of the number of correctly classified samples to the total number of samples. Accuracy refers to the proportion of samples that are actually positive among all samples predicted by the model to be positive. High accuracy means that fewer negative samples are misclassified as positive samples. Recall refers to the proportion of samples that are correctly predicted to be positive by the model among all samples that are actually positive. A high recall means that the model can capture most of the actual positive samples. The F1 score is the harmonic mean of precision and recall, which is used to comprehensively consider the two. The F1 score provides a balance between precision and recall, and is particularly suitable for cases of class imbalance.

[0166] Step S3: Based on the enzyme-reaction relationship prediction model, output a binary classification prediction result corresponding to the simplified molecular linear input canonical expression of the protein enzyme sequence and the biochemical reaction.

[0167] The enzyme-reaction relationship prediction model takes the protein enzyme sequence and the SMILES expression of the biochemical reaction as input and outputs a binary classification result. The output of the model is 0 or 1, where 0 means that the protein enzyme does not catalyze the reaction, and 1 means that the protein enzyme can catalyze the reaction.

[0168] In addition, this embodiment also includes the construction of a comparative model based on reaction representation: the comparative experiment of this embodiment is to compare DeepRP's RoBERTa-based reaction representation with three other advanced reaction representation learning methods - Drfp, MolR, and Rxnfp. In terms of experimental design, Figure 5 As shown, this example replaces the original reaction representation in DeepRP with three different representation learning methods: Drfp, MolR, and Rxnfp. The protein representation uses the ESM-2 model. This method allows for intuitive comparison of the performance differences of each representation method under the same architecture.

[0169] In addition, the present embodiment also includes the following experimental part:

[0170] 1. Experimental environment:

[0171] The experimental environment of this example is built on the Red Hat Enterprise Linux (RHEL) operating system, providing a stable and reliable platform for the experiment. Computing resources are supported by hardware equipped with an NVIDIA A800-SXM4-80GB GPU. All development work is performed in the Python language environment, and the development and training of the deep learning model are based on the PyTorch 2.0.1 framework. Data visualization relies on Matplotlib and seaborn tools.

[0172] 2. Experimental data:

[0173] Data on enzyme-reaction associations were extracted from the Rhea database, totaling 359,633 records. After screening, reversible reactions were identified, yielding 16,270 pieces of relevant information. These reversible reactions encompass 7,501 different reactions and 69,914 proteins. The collected data included reaction IDs, ChEBI database IDs, EC numbers, and PubMed citations. Furthermore, information on 224,763 compounds, including their structural information, was obtained from the ChEBI database. After preprocessing these data, an initial enzyme-reaction database was formed, containing a total of 297,362 records (including negative samples).

[0174] Next, we retrieved the corresponding structural information from the PDB database based on the enzyme sequence information. For enzymes for which structural information was not found in the PDB, we selected the AlphaFold database. Furthermore, we calculated the amount of crystal structure data available in the PDB for enzymes with different EC numbers, as shown in Table 2.

[0175] Table 2 Protein PDB data corresponding to reaction EC

[0176] EC number Total protein data Proteins with PDB structures percentage(%) EC: 1 18834 620 3.291 EC: 2 44820 977 2.180 EC: 3 20759 617 2.972 EC: 4 13717 335 2.440 EC: 5 11804 253 2.143 EC: 6 14913 281 1.880 EC: 7 2555 60 2.348

[0177] In this example, the data of enzymes with EC:4 were selected as the training set. According to previous statistics, there are a total of 13,717 records of enzymes with EC:4. Among these records, 335 have crystal structure information obtained from the PDB database, accounting for approximately 2.44%. The data of cleavage reactions were selected as the research object mainly based on the following considerations: (1) EC:4 enzymes specialize in cleavage reactions, which play a key role in biochemical processes. Focusing on this type of reaction can help us gain a deeper understanding of the specific biochemical processes of this category. (2) The amount of data for cleavage reactions is moderate, neither too large to be difficult to process nor too small to lack statistical significance. This balance helps to conduct effective data analysis and model training. (3) Since the enzyme data of cleavage reactions contain relatively more PDB structural information, it provides valuable input for building and training prediction models. This structural information can be used to improve the accuracy and reliability of prediction models. (4) Cleavage reactions are very important in many biological processes, such as metabolic pathways and signal transduction. A deep understanding of these reactions can promote the discovery of new drugs and the study of disease treatment methods.

[0178] 3. Parameter settings:

[0179] The model combines biochemical reaction and protein sequence data to construct a multimodal architecture that integrates reaction and protein representation learning. Feature fusion and subsequent classification are then performed on this basis. To improve the model's generalization and reduce the risk of overfitting, some layers were frozen, and the remaining layers were fine-tuned. During the experiment, the model parameters were carefully selected and tuned: a small learning rate of 4e-5 was set to allow for more nuanced weight tuning; 200 training epochs were chosen to ensure sufficient model training; and a batch size of 36 was chosen to strike a balance between computational efficiency and memory capacity. The embedding vector sizes for protein sequences and chemical reaction features were set to 446 (SaProt) and 1280 (ESM-2), respectively, to ensure sufficient information representation. Furthermore, the model was configured with a 4-head multi-head attention mechanism, enabling the model to learn complex features in parallel across multiple subspaces. A dropout rate of 0.5 was introduced to enhance robustness when handling new data. With this parameter configuration, the model achieved higher accuracy in predicting the relationship between enzymes and reactions.

[0180] 4. DeepRP model results:

[0181] Training strategy: During the training process, this embodiment adopts the strategy of pretraining model and fine-tuning to further improve the performance and generalization ability of the model.

[0182] The training data set is: 13,594 enzyme-reaction relationship data from cleavage reactions, that is, reactions with an EC number starting with 4, were selected. 70% of them were used as training sets, 15% as validation sets, and 15% as test sets. The test set obtained was 2,040 samples: 1,114 of them were positive samples and 926 were negative samples.

[0183] exist Figure 6 In this paper, a comparison of two protein feature representation methods is presented: the structure-based SaProt model and the sequence-based ESM-2 model, corresponding to DeepRP_esm2 and DeepRP_saprot, respectively. The two models demonstrated similar performance metrics, both demonstrating excellent predictive capabilities. This demonstrates that both structural and sequence-based protein feature representations effectively support the high performance of the DeepRP model in enzyme-reaction relationship prediction. DeepRP_esm2 slightly outperformed DeepRP_saprot in accuracy, precision, recall, and F1 score, although the differences were not significant. Specifically, DeepRP_esm2 achieved slightly higher accuracy, indicating a slightly better performance in classifying correct and incorrect predictions. Overall, DeepRP_esm2's superior overall performance may be due to its data processing approach or certain optimizations during training. This result confirms that the proposed model can achieve atomic-level 3D structure prediction when its parameters are scaled from 8 million to 15 billion.

[0184] In order to further explore the prediction of positive and negative samples, this embodiment visualizes the confusion matrix diagram, such as Figure 7 As shown in the figure, the main diagonal of the matrix (from top left to bottom right) shows the number of correct predictions, namely true positives and true negatives, corresponding to labeled positive and negative samples, respectively. The off-diagonal elements indicate prediction errors, including false positives and false negatives. The blue part in the figure represents the prediction results of DeepRP_esm2, and the red part represents the prediction results of DeepRP_saprot. Judging from the numbers on the main diagonal, both models demonstrated high prediction accuracy. DeepRP_esm2 correctly predicted 1075 positive samples and 877 negative samples, while DeepRP_saprot correctly predicted 1068 positive samples and 844 negative samples. From this result, it can be concluded that the model performance of DeepRP_esm2 is slightly better than that of DeepRP_saprot.

[0185] In addition, this embodiment uses the principal component analysis (PCA) method to reduce the dimensionality of the features on the test set, so as to view the specific differences in the prediction of the enzyme-reaction relationship on the test set between DeepRP_esm2 and DeepRP_saprot. PCA is a statistical method that can simplify the complexity of the data set by reducing the dimensionality of the data while retaining the variability in the original data as much as possible. In this case, PCA can extract the two most important components from the original data set containing many variables. These two main components capture most of the data variability and represent each data point on two new axes, which are usually linear combinations of the original feature space. Figure 8 Three sets of scatter plots are shown, representing the actual labels, the labels predicted by DeepRP_esm2, and the labels predicted by DeepRP_saprot. The first plot shows the distribution of the true labels, which can be considered the baseline or ground truth. The second plot shows the labels predicted by the DeepRP_esm2 model. Comparing this plot with the first plot, we can see that the DeepRP_esm2 model is able to capture the distribution characteristics of the two data types to a certain extent, but there is confusion in certain areas, resulting in classification errors. These errors may be due to the model not clearly demarcating the boundaries of the data points in those areas. The third plot (right) shows the predictions of the DeepRP_saprot model. By observing the comparison with the true labels, we can see that the DeepRP_saprot model mixes the two data types in some areas and does not distinguish them well, especially in areas where there is significant overlap between the two data types.

[0186] In the enzyme-reaction prediction model, this example selected three reaction representation models, Drfp, MolR, and Rxnfp, for comparison, and the protein representation model used ESM-2. Figure 9The results presented show that DeepRP_esm2, using XLM-RoBERTa as its reaction representation model, outperforms three other functional group-based reaction representation models in terms of accuracy, precision, recall, and F1 score. DeepRP_esm2 achieved an F1 score of 0.9564, compared to 0.8007 for Drfp_esm2, 0.6198 for Molr_esm2, and 0.7549 for Rxnfp_esm2. On this comprehensive performance metric, DeepRP_esm2 outperformed the three models by 15.57%, 33.66%, 20.15%, and 20.22%, respectively. These performance metrics highlight the superior performance of DeepRP_esm2 using XLM-RoBERTa in enzyme-reaction prediction tasks and demonstrate its effectiveness in capturing the relationship between enzymes and reactions. The XLM-RoBERTa model analyzes functional groups similar to "words" in language. Compared with traditional text analysis methods, this method adds more biochemical information and more accurately predicts enzyme-reaction interactions.

[0187] The present invention is not limited to the structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A method for predicting enzyme-reaction relationships based on representation learning, characterized in that: include: Based on a preset biological database, relevant data of enzyme-reaction relationships are obtained and preprocessed, and an enzyme-reaction relationship database is constructed based on the preprocessed relevant data of enzyme-reaction relationships; A multimodal fusion architecture is used to integrate sequence- and structure-based protein language models and functional group-based improved biochemical reaction language models, and an enzyme-reaction relationship prediction model is constructed in conjunction with an interpretability mechanism. Based on the enzyme-reaction relationship prediction model, the binary prediction results corresponding to the simplified molecular linear input canonical expression of protein enzyme sequence and biochemical reaction are output.

2. The enzyme-reaction relationship prediction method based on representation learning according to claim 1, characterized in that The method of obtaining relevant data on enzyme-reaction relationships based on a preset biological database and preprocessing the data, and constructing an enzyme-reaction relationship database based on the preprocessed relevant data on enzyme-reaction relationships, comprises: Based on biological database 1, extract the reaction information of the compounds involved in the reaction and the related enzymes; According to the compound identifiers obtained in the biological database 1, the corresponding compound information is extracted from the biological database 2, and the compound information is associated with the reaction information; According to the enzyme identifier provided by biological database 1, the sequence information of the relevant enzyme is extracted from biological database 3, and the structural information is obtained from biological database 4 for matching, and the sequence information, functional information and structural information of the enzyme are integrated; The data extracted from biological database 1, biological database 2, biological database 3 and biological database 4 were integrated and cleaned and formatted to obtain the enzyme-reaction relationship database.

3. The enzyme-reaction relationship prediction method based on representation learning according to claim 1, characterized in that The multimodal fusion architecture is used to integrate the protein language model based on sequence and structure and the biochemical reaction language model based on functional group improvement, and the enzyme-reaction relationship prediction model is constructed in combination with the interpretability mechanism, including: According to the functional groups of biochemical reactions, the simplified molecular linear input specification representation of the reaction is divided. Combined with the word elements defined by functional groups and the corresponding decomposition tools, a biochemical reaction language model based on functional group improvement is constructed. Using pre-training and fine-tuning strategies, combined with protein sequences represented in the form of amino acid feature sequences, a sequence-based protein language model is constructed; Based on Foldseek 3D structure information encoding and protein language model, a structure-based protein language model was constructed; By utilizing a multimodal feature fusion strategy, we integrated the functional group-based improved biochemical reaction language model, the sequence-based protein language model, and the structure-based protein language model to obtain a matrix containing enzyme-reaction relationships, and constructed a classifier for enzyme-reaction relationship prediction.

4. The enzyme-reaction relationship prediction method based on representation learning according to claim 3, characterized in that The simplified molecular linear input specification representation of the reaction divided according to the functional groups of the biochemical reaction, combined with the word elements defined by the functional groups and the corresponding decomposition tools, constructs a biochemical reaction language model based on the functional group improvement, including: Obtain atom type tables from core databases in the field of bioinformatics and modify them using the SMARTS format to ensure that each functional group can be uniquely identified, thus obtaining a functional group attribute table based on the SMARTS format. Get the functional group list represented by the simplified molecular linear input specification according to the functional group attribute table based on the SMARTS format, and split the simplified molecular linear input specification string of the reaction into separate compound lists; Split and simplify the data of the functional group list and compound list represented by the molecular linear input specification, and use chemical information analysis tools to divide the rows into functional groups to obtain a vocabulary list based on functional groups; The reaction data of the molecular linear input canonical representation is simplified based on the functional group list division to obtain several words, and the embedding representation of several words is input into the hidden layer of the RoBERTa model to obtain the vector representation corresponding to each word, thereby generating a vector representation of the biochemical reaction.

5. The enzyme-reaction relationship prediction method based on representation learning according to claim 4, characterized in that The splitting simplifies the data of the functional group list and compound list represented by the molecular linear input specification, and uses chemical information analysis tools to divide the functional groups into rows, and obtains a vocabulary list based on functional groups including: Based on the first function in the chemical information analysis tool, each item in the data of the functional group list and the compound list represented by the simplified molecular linear input specification is converted into a molecular object to obtain a molecular object list of the functional group and the compound; Using a second function in the chemical information analysis tool, identify the common substructures between each molecule in the molecular object list of the compound and the functional group in the molecular object list of the functional group, and generate a list of atomic numbers matching each compound with the functional group; Traverse the matching atomic number list, find the chemical bonds of each atomic number list item in the corresponding compound and the chemical bonds of adjacent atoms in each substructure, and save the chemical bond information separately; By calculating the difference set of chemical bonds, the chemical bonds that need to be broken are determined, and the chemical bonds in the difference set are broken using the third function in the chemical information analysis tool, and the break positions are marked. The results are converted into a simplified molecular linear input standard format; The simplified molecular linear input canonical string is split according to the labeled break positions to obtain a functional group-based vocabulary list.

6. The enzyme-reaction relationship prediction method based on representation learning according to claim 4, characterized in that The partitioning expression of the reaction data of the simplified molecular linear input specification is: X=[X1,X2,...X N ] The expression of word embedding is: V i =W(X i )+S+P(i),V i ∈R d The expression of the vector representation corresponding to the word unit is: Z=[Z1,Z2,...,Z N ] In the formula, X represents word unit, X i Represents each word after division, N represents the number of words, V i Represents the embedding representation of the word, W(X i ) represents the word embedding of the word unit, S represents the sentence embedding, P(i) represents the position embedding, R represents the real number set, d represents the embedding dimension, Z represents the vector representation corresponding to the word unit, Z i Represents the vector representation of the i-th word after being processed by the RoBERTa model.

7. The enzyme-reaction relationship prediction method based on representation learning according to claim 3, characterized in that: The method of constructing a structure-based protein language model based on Foldseek 3D structure information encoding and protein language model includes: In the Foldseek 3D structure information encoding stage, the acquired features are used as input to the encoder, and the encoder is used to convert discrete features into continuous features; The vectorization process maps the continuous representation of the latent space to the nearest center of mass, obtaining discretized descriptors of several protein conformations; The decoder uses the vectorized results to reconstruct the input data and predict the probability distribution of the descriptors of the residues that are structurally aligned with each residue to obtain the structural label of each residue; Given a protein sequence, the structural tag information of Foldseek is integrated to represent the sequence with structural tags; each residue is treated as a word unit and a mask signal is introduced to form a perception table; The sequence is masked, and a predetermined proportion of residues are selected to replace with mask marks to obtain a masked pattern sequence; the masked pattern sequence is converted into an embedding vector sequence using an embedding layer, and the embedding vector sequence is input into the BERT model to obtain a protein embedding matrix containing structural information.

8. The enzyme-reaction relationship prediction method based on representation learning according to claim 7, characterized in that: The expression of the encoder function is: The expression for the protein conformation class of a residue is: F=(f1,f2,...,f n ) The expression for a sequence with a structure tag is: P=[P1,P2,...P N ] The expression of the embedding vector sequence is: E=[e(p1'),e(p'2),...e(p' n )] The expression of the protein embedding matrix is: M P =BERT(E),M P ∈R l×D Where Enc represents the encoder function, x n represents the input of the encoder, z n Representation is a continuous representation of the latent space, D z represents the dimension of the latent space, F represents the protein conformation category of the residue, and f i represents the structural label of the i-th residue, P represents the protein sequence, and p i represents the residue at position i, E represents the embedding vector sequence, e(.) represents the embedding function, and p n ' represents the sequence of the mask pattern of the nth residue, M P represents the protein embedding matrix, BERT represents the BERT model, l represents the length of the protein sequence with structural information, and D represents the final protein vector representation dimension.

9. The enzyme-reaction relationship prediction method based on representation learning according to claim 3, characterized in that: The multimodal feature fusion strategy is used to integrate the functional group-based improved biochemical reaction language model, the sequence-based protein language model, and the structure-based protein language model to obtain a matrix containing enzyme-reaction relationships, and to construct a classifier for enzyme-reaction relationship prediction, including: Obtain protein sequence vectors and biochemical reaction vectors, and capture the complex relationship between biochemical reactions and protein sequences based on linear transformation, vector interaction, and enzyme-reaction identifier construction technology to obtain a matrix containing enzyme-reaction relationships; Based on the matrix containing the enzyme-reaction relationship, a classifier for predicting the enzyme-reaction relationship is constructed, wherein the classifier consists of a long short-term memory network layer, a multi-layer perceptron layer and a softmax.

10. The enzyme-reaction relationship prediction method based on representation learning according to claim 9, characterized in that: The linear transformation includes linearly transforming the biochemical reaction vector so that its dimension matches the dimension of the protein sequence vector; The vector interaction includes calculating an interaction matrix based on the protein sequence vector and the converted biochemical reaction vector to capture the interaction information between the chemical reaction and the protein sequence; The enzyme-reaction identifier construction includes taking the interaction matrix as the input of the Transformer encoder and using the Transformer encoder to capture the fusion information to obtain a matrix containing the enzyme-reaction relationship.

Citation Information

Cited By

  • Enzyme mining method and system and storage medium

    CN121415892A

  • Transform network-based biological metabolic pathway inverse synthesis design method and system

    CN121617479A