Enzyme kinetic parameter prediction model construction method and system
Through multi-source data integration and multi-task learning methods, the data sparsity and insufficient feature fusion in the prediction of enzyme kinetic parameters are solved, the prediction accuracy of enzyme kinetic parameters is improved, the design and transformation of enzymes is supported, and the development of biosynthesis strategy is promoted.
Patent Information
- Application Number
- CN202510503379.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-05
AI Technical Summary
Computer-aided prediction of enzyme kinetic parameters in the prior art faces the problems of data sparsity, insufficient feature fusion, and neglected correlation between Michal constant and enzyme turnover number, resulting in high cost of screening of enzyme databases and long periods, which makes it difficult to meet large-scale demands.
Through multi-source data integration, the enzyme kinetic parameter set is constructed, data cleaning and enhancement are carried out, and the enzyme kinetic parameter prediction model for multi-task learning is constructed. The low-rank bilinear attention network and the learnable bias enhancement network are used, and the enzyme kinetic parameter prediction model is optimized.
It improves the accuracy of the prediction of enzyme kinetic parameters, assists in virtual screening and rational design of auxiliary strains, accelerates the development of biosynthesis strategy, and enhances the understanding of enzyme-substrate binding and catalytic correlation.
Smart Images

Figure CN120432037A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of enzyme kinetic parameter prediction, and in particular to a method and system for constructing an enzyme kinetic parameter prediction model. Background Art
[0002] As efficient and specific biocatalysts, enzymes play an increasingly important role in synthetic biology, biomanufacturing, and green chemistry. m and enzyme turnover number k cat ) quantifies the binding strength and catalytic efficiency of enzymes and substrates, and is a key quantitative indicator for analyzing enzyme catalytic mechanisms and optimizing industrial enzyme performance. However, traditional experimental determination methods are limited by factors such as high screening costs and long screening cycles, making them difficult to meet the needs of large-scale enzyme library screening.
[0003] With the deep integration of biosynthesis technology and artificial intelligence, existing research has begun to explore computer-assisted prediction of enzyme kinetic parameters. However, these studies still face challenges such as data sparsity and insufficient feature fusion, which are reflected in the following aspects:
[0004] 1. Some enzyme kinetic data in the BRENDA enzyme database lack corresponding protein annotations, and the compound database names are not standardized, resulting in a low retention rate of raw data. Moreover, as a data-driven technology, traditional machine learning's performance ceiling is limited by the quality and quantity of the dataset. During the data cleaning process, protein sequences of unknown origin are introduced during data collection, increasing data noise. In addition, the lack of data features leads to the loss of a large amount of data. This ultimately leads to the sparsity and high noise level of the training dataset.
[0005] 2. Existing computational models (such as DLKcat) typically use a simple splicing method for feature fusion, which cannot capture the deep mapping relationship between proteins and substrates. Although traditional machine learning algorithms are relatively easy to use, they can only accept vector data as input and will lose a lot of information when averaging and pooling high-dimensional features. In addition, the pooled proteins and molecular features can only be simply spliced. Although simple and intuitive and does not introduce additional parameters, this method ignores the weight coefficient and cannot fully reflect the importance of different modalities, resulting in poor feature fusion effect.
[0006] 3. The Michaelis constant and enzyme turnover number jointly describe the steady-state reaction process of the enzyme, and they are interrelated; however, existing studies mostly model a single data set and ignore the correlation between these parameters.
[0007] Therefore, how to provide a method and system for constructing an enzyme kinetic parameter prediction model is an urgent problem to be solved. Summary of the Invention
[0008] The embodiments of the present invention provide a method and system for constructing an enzyme kinetic parameter prediction model to solve the above-mentioned technical problems existing in the prior art.
[0009] To provide a basic understanding of some aspects of the disclosed embodiments, the following is a brief summary. This summary is not intended to be a comprehensive review, identify key or essential elements, or delineate the scope of these embodiments. Its sole purpose is to present some concepts in a simplified form as a prelude to the detailed description that follows.
[0010] According to a first aspect of an embodiment of the present invention, a method for constructing an enzyme kinetic parameter prediction model is provided.
[0011] In one embodiment, the method for constructing an enzyme kinetic parameter prediction model comprises:
[0012] By integrating multi-source data, we construct an enzyme kinetic parameter set, then perform data cleaning and data enhancement in sequence to supplement protein information and match compound data. Based on the data partitioning strategy, we divide the enzyme kinetic parameter set into different types of data sets.
[0013] Constructing a multi-task learning enzyme kinetic parameter prediction model that outputs enzyme kinetic parameter prediction results by modeling the synergistic relationship between the Michaelis constant and enzyme turnover number data; the enzyme kinetic parameter prediction model includes a feature extraction network, a low-rank bilinear attention network, a learnable bias enhancement network, and a decoder;
[0014] Based on a multi-stage training strategy, the enzyme kinetic parameter prediction model is trained in segments. During the training process, label normalization, random inactivation tuning and learning rate warm-up are introduced for optimization to generate an enzyme kinetic parameter prediction model with optimal parameters.
[0015] In one embodiment, the enzyme kinetic parameter set is constructed by integrating multi-source data, and data cleaning and data enhancement are performed in sequence to achieve protein information supplementation and compound data matching; and based on the data partitioning strategy, the enzyme kinetic parameter set is divided into different types of data sets including:
[0016] Download Michaelis constants and enzyme turnover numbers from the Brunswick Enzyme Database, along with corresponding Universal Protein Database IDs, mutation information, reaction substrate names, and literature numbers. Use ID conversion tools to obtain amino acid sequences, and then, based on compound names, obtain ligand identifiers from the compound database. After integrating multi-source data, a set of enzyme kinetic parameters is constructed.
[0017] Delete the data for which the sequence or ligand identifier cannot be obtained, merge the data with the same amino acid sequence and ligand identifier, and perform arithmetic averaging on the data labels that need to be merged;
[0018] Protein information missing from universal protein databases is supplemented through data backfilling, and identical compound names are standardized through multi-database searches.
[0019] The enzyme kinetic parameter set was divided into training set, validation set and test set in a ratio of 8:1:1; and based on whether there were identical or similar proteins between the training set, validation set and test set, the enzyme kinetic parameter set was divided into hot start data set and cold start data set. The hot start test set or cold start test set was then functionally divided to obtain multiple test subsets.
[0020] In one embodiment, in the process of obtaining the amino acid sequence using the ID conversion tool, the amino acid sequence of the mutant protein is modified on the wild-type amino acid sequence according to the mutation information.
[0021] In one embodiment, the method of supplementing protein information missing from the universal protein database ID by data backfilling and normalizing the names of identical compounds by multi-database searching includes:
[0022] The unique identification code of the literature was used to match the universal protein database and search, and the universal protein database knowledge base entries in the literature were captured. The species and enzyme commission numbers that matched the Brunswick enzyme database records were screened by scripts, and the proteins that met the conditions were retained as supplementary data.
[0023] By using the enzyme map tool, simplified molecular linear input specifications can be obtained from the compound library in one stop through the reaction substrate name to improve the accuracy of simplified molecular linear input specification matching.
[0024] In one embodiment, the test subsets include a wild-type and mutant protein test set, a Michaelis constant common data and orphan data test set, an enzyme turnover number common data and orphan data test set, a low Michaelis constant enzyme test set, and a high enzyme turnover number enzyme test set.
[0025] In one embodiment, the multi-task learning enzyme kinetic parameter prediction model is constructed, and the output of the enzyme kinetic parameter prediction results includes:
[0026] A feature extraction network is used to capture protein and molecular features, mapping each input into a two-dimensional feature matrix. Protein and molecular models are then used to map discrete character data into a high-dimensional continuous space.
[0027] A low-rank bilinear attention network is used to calculate the bilinear mapping between protein features and molecular features, generate an attention weight matrix, and then pool the matrix to obtain a joint representation, which is then input into a learnable bias enhancement network.
[0028] The learnable bias-augmented network receives the joint representation of the Michaelis constant and enzyme turnover number data, selects the expert output through the task-specific layer, generates a task-specific representation, and then average-pools it into a linear bias for the task;
[0029] The fused representation is added to the task-specific representation bias, the high-dimensional fused representation is reduced in dimension, and mapped to the final model output through a linear layer to obtain the prediction results of enzyme kinetic parameters.
[0030] In one embodiment, the method of calculating the bilinear mapping between protein features and molecular features using a low-rank bilinear attention network, generating an attention weight matrix, obtaining a joint representation through pooling, and inputting the result into a learnable bias enhancement network includes:
[0031] Mapping the amino acid sequence and molecular structure of a protein into a high-dimensional space to generate a feature representation that captures the complex patterns between the sequence and structure. It also calculates the bilinear mapping between protein features and molecular features, generating an attention weight matrix that reflects the correlation between features.
[0032] A low-rank strategy is introduced to decompose the bilinear mapping into the product of two low-rank matrices, and the attention weight matrix and the feature matrix are integrated through the pooling layer to generate a simplified joint representation.
[0033] In one embodiment, the learnable bias augmentation network receives a joint representation of the Michaelis constant and enzyme turnover number data, selects expert outputs through a task-specific layer, generates a task-specific representation, and then average-pools it into a linear bias for the task, including:
[0034] The shared features of the Michaelis constant and enzyme turnover number data are extracted through the shared layer. The task-specific layer selects the most relevant expert outputs through a gating mechanism and linearly combines them to generate task feature representations.
[0035] The task-specific layer adjusts the learnable bias term by output, provides bias for the prediction results, and connects the shared features to the task-specific network in the form of a learnable bias residual.
[0036] In one embodiment, the stage training strategy includes a multi-task pre-training stage and a single-task fine-tuning stage. The multi-task pre-training stage freezes the protein and molecule pre-training large models and uses the full data of the two sets of data for training; the single-task fine-tuning stage freezes other networks including shared layers and trains the decoder network weights of the task only on a single data of Michaelis constant or enzyme turnover number data.
[0037] According to a second aspect of an embodiment of the present invention, a system for constructing an enzyme kinetic parameter prediction model is provided.
[0038] In one embodiment, the enzyme kinetic parameter prediction model construction system comprises:
[0039] The data acquisition module is used to construct enzyme kinetic parameter sets by integrating multi-source data, and then perform data cleaning and data enhancement in sequence to achieve protein information supplementation and compound data matching; and based on the data partitioning strategy, the enzyme kinetic parameter set is divided into different types of data sets;
[0040] The model building module is used to fuse protein features and compound molecular features using a low-rank bilinear attention network, combined with a learnable bias enhancement network, to build a multi-task learning enzyme kinetic parameter prediction model; a decoder is used to reduce the dimensionality of the high-dimensional fusion and output the prediction results of the enzyme kinetic parameters;
[0041] The model training module is used to perform segmented training on the enzyme kinetic parameter prediction model based on a multi-stage training strategy. During the training process, label normalization, random inactivation tuning, and learning rate warm-up are introduced for optimization to generate an enzyme kinetic parameter prediction model with optimal parameters.
[0042] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0043] 1. Improve the accuracy of enzyme kinetic parameter predictions by building a richer dataset of enzyme kinetic parameters and more rational deep learning models. Accurately predicting enzyme kinetic parameters will assist in virtual screening and rational design of strains, thereby accelerating the development of biosynthetic strategies.
[0044] 2. By applying multi-task learning technology, K m With k cat The two sets of data were used to train the enzyme kinetic parameter prediction model. m With k cat The complementary relationship between the data helps the model extract additional features to enhance predictive capabilities; therefore, the effective verification of the multi-task model will support the correlation between enzyme-substrate binding and catalysis from a computational perspective, providing theoretical support for the design and modification of enzymes.
[0045] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0047] Figure 1 is a flow chart of a method for constructing an enzyme kinetic parameter prediction model according to an exemplary embodiment;
[0048] Figure 2 This is a system principle block diagram of a system for constructing an enzyme kinetic parameter prediction model according to an exemplary embodiment;
[0049] Figure 3 is a flowchart of constructing MTLKP-DB according to an exemplary embodiment;
[0050] Figure 4 is a schematic diagram of a STLKP model framework according to an exemplary embodiment;
[0051] Figure 5 FIG. 4 is a schematic diagram of an MTLKP model framework according to an exemplary embodiment. DETAILED DESCRIPTION
[0052] Figure 1 An embodiment of a method for constructing an enzyme kinetic parameter prediction model of the present invention is shown.
[0053] In this optional embodiment, the method for constructing an enzyme kinetic parameter prediction model includes:
[0054] Step S101: constructing an enzyme kinetic parameter set by integrating multi-source data, performing data cleaning and data enhancement in sequence to achieve protein information supplementation and compound data matching; and dividing the enzyme kinetic parameter set into different types of data sets based on a data partitioning strategy;
[0055] Step S102: constructing a multi-task learning enzyme kinetic parameter prediction model, outputting the prediction results of the enzyme kinetic parameters by modeling the synergistic relationship between the Michaelis constant and the enzyme turnover number data; the enzyme kinetic parameter prediction model includes a feature extraction network, a low-rank bilinear attention network, a learnable bias enhancement network and a decoder;
[0056] Step S103: Based on a multi-stage training strategy, the enzyme kinetic parameter prediction model is segmented trained. During the training process, label normalization, random inactivation tuning and learning rate warm-up are introduced for optimization to generate an enzyme kinetic parameter prediction model with optimal parameters.
[0057] In this optional embodiment, when constructing the enzyme kinetic parameter set by integrating multi-source data, data cleaning and data enhancement are performed in sequence to achieve protein information supplementation and compound data matching; and when the enzyme kinetic parameter set is divided into different types of data sets based on the data partitioning strategy, the Michaelis constant and enzyme turnover number data, as well as the corresponding universal protein database ID, mutation information, reaction substrate name and literature number can be downloaded from the Brunswick enzyme database, the amino acid sequence can be obtained using the ID conversion tool, and the ligand identifier can be obtained from the compound database according to the compound name; after the multi-source data is integrated, the enzyme kinetic parameter set is constructed; and the data for which the sequence or ligand cannot be obtained are deleted. The data of identifiers were merged with the data with the same amino acid sequence and ligand identifier, and the data labels that needed to be merged were processed by arithmetic averaging; the protein information of the missing universal protein database ID was supplemented by data backfill, and the names of the same compounds were normalized through multi-database search; the enzyme kinetic parameter set was divided into training set, validation set and test set in a ratio of 8:1:1; and the enzyme kinetic parameter set was divided into hot start data set and cold start data set according to whether there were identical or similar proteins among the training set, validation set and test set, and then the hot start test set or the cold start test set was functionally divided to obtain multiple test subsets.
[0058] In this optional embodiment, in the process of obtaining the amino acid sequence using the ID conversion tool, the amino acid sequence of the mutant protein is modified on the wild-type amino acid sequence according to the mutation information.
[0059] In this optional embodiment, when supplementing the protein information of the missing universal protein database ID through data backfilling and realizing the normalization of the same compound name through multi-database search, the universal protein database can be matched and searched through the unique identification code of the document to capture the universal protein database knowledge base entries in the document, and the species and Enzyme Commission numbers that match the Brunswick enzyme database records can be screened through a script, and the proteins that meet the conditions are retained as backfill data; using the enzyme map tool, the simplified molecular linear input specification can be obtained from the compound library in a one-stop manner through the reaction substrate name to improve the accuracy of the simplified molecular linear input specification matching.
[0060] In this optional embodiment, the test subsets include a wild-type and mutant protein test set, a Michaelis constant common data and orphan data test set, an enzyme turnover number common data and orphan data test set, a low Michaelis constant enzyme test set, and a high enzyme turnover number enzyme test set.
[0061] In this optional embodiment, in the construction of the enzyme kinetic parameter prediction model for multi-task learning, by modeling the synergistic relationship between the Michaelis constant and the enzyme turnover number data, when outputting the prediction results of the enzyme kinetic parameters, a feature extraction network can be used to capture protein features and molecular features, and each input is mapped to a two-dimensional feature matrix, and the protein and molecule model is used to map the discrete character data to a high-dimensional continuous space; a low-rank bilinear attention network is used to calculate the bilinear mapping of protein features and molecular features, generate an attention weight matrix, and obtain a joint representation through pooling, which is input into a learnable bias enhancement network; the learnable bias enhancement network receives the joint representation of the Michaelis constant and the enzyme turnover number data, selects the expert output through the task-specific layer, generates a task-specific representation, and then averages and pools it into a linear bias for the task; the fused representation is added to the task-specific representation bias, the high-dimensional fused representation is reduced in dimensionality, and mapped to the final model output through a linear layer to obtain the prediction results of the enzyme kinetic parameters.
[0062] In this optional embodiment, the low-rank bilinear attention network is used to calculate the bilinear mapping between protein features and molecular features, generate an attention weight matrix, and obtain a joint representation through pooling. When input into the learnable bias enhancement network, the amino acid sequence and molecular structure of the protein can be mapped to a high-dimensional space to generate a feature representation, capture the complex pattern between the sequence and the structure, and calculate the bilinear mapping between the protein features and the molecular features to generate an attention weight matrix to reflect the correlation between the features; a low-rank strategy is introduced to decompose the bilinear mapping into the product of two low-rank matrices, and the attention weight matrix and the feature matrix are integrated through a pooling layer to generate a simplified joint representation.
[0063] In this optional embodiment, the learnable bias-enhanced network receives a joint representation of the Michaelis constant and the enzyme turnover number data, selects the expert output through the task-specific layer, generates a task-specific representation, and then average-pools it into a linear bias for the task. The common features of the Michaelis constant and the enzyme turnover number data can be extracted through the shared layer, and the task-specific layer selects the most relevant expert output through a gating mechanism and linearly combines them to generate a task feature representation; the task-specific layer adjusts the learnable bias term through the output to provide a bias for the prediction result, and residually connects the shared features to the task-specific network in the form of a learnable bias.
[0064] In this optional embodiment, the stage training strategy includes a multi-task pre-training stage and a single-task fine-tuning stage. The multi-task pre-training stage is to freeze the protein and molecule pre-training large models and use the full data of the two sets of data for training; the single-task fine-tuning stage is to freeze other networks including shared layers and only train the decoder network weights of the task on a single data of Michaelis constant or enzyme turnover number data.
[0065] Figure 2An embodiment of an enzyme kinetic parameter prediction model construction system of the present invention is shown.
[0066] In this optional embodiment, the enzyme kinetic parameter prediction model construction system includes:
[0067] The data acquisition module 201 is used to construct an enzyme kinetic parameter set by integrating multi-source data, and sequentially perform data cleaning and data enhancement to achieve protein information supplementation and compound data matching; and based on the data partitioning strategy, the enzyme kinetic parameter set is divided into different types of data sets and;
[0068] The model construction module 202 is used to fuse protein features and compound molecular features using a low-rank bilinear attention network, combined with a learnable bias enhancement network, to build a multi-task learning enzyme kinetic parameter prediction model; use a decoder to reduce the dimension of the high-dimensional fusion and output the prediction results of the enzyme kinetic parameters;
[0069] The model training module 203 is used to perform segmented training on the enzyme kinetic parameter prediction model based on a multi-stage training strategy, and introduce label normalization, random inactivation tuning and learning rate warm-up during the training process for optimization to generate an enzyme kinetic parameter prediction model with optimal parameters.
[0070] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0071] In view of the sparsity of enzyme kinetics data, the present invention proposes a data enhancement strategy and constructs the largest enzyme kinetics parameter dataset in the current field. In view of the problems of insufficient feature interaction and parameter association learning in the prediction of enzyme catalytic properties in traditional machine learning, the present invention introduces a low-rank bilinear attention mechanism and a learnable bias enhancement network to establish a new method for the interaction between protein-ligand complex features and kinetic data. The model strictly follows the technical route of "feature extraction-interaction modeling-multi-task fusion-model verification" and provides a systematic solution for the accurate prediction of enzyme catalytic properties. The following will systematically introduce the construction process of the kinetics dataset and K m With k cat Predictive model implementation details and its training and testing methods.
[0072] 1. Dataset Construction
[0073] The model proposed in this paper involves three main pieces of information: amino acid sequence, ligand identifier (SMILES) and Michaelis constant (K m ) / enzyme turnover number (k catAfter feature extraction, amino acid sequences and ligand identifiers support model training and inference; the Michaelis constant / enzyme turnover number serve as labels, providing conditions for error backpropagation in supervised learning. The specific composition and source of the dataset are shown in Table 1. This section will comprehensively describe the collection and cleaning process of the enzyme kinetic parameter dataset, the dataset partitioning strategy, and the test set construction method.
[0074] Table 1: Information contained in the enzyme kinetic parameter dataset
[0075]
[0076] Note: “-” represents an attribute that does not exist in the dataset.
[0077] 1.1 Data Collection and Cleaning
[0078] BRENDA is a data platform that provides enzyme-related information. Data is collected from scientific literature, database integration, or user submissions, and is incorporated into the database after expert review. Although the official SOAP API is provided to obtain information in JSON format, in practice, it has been found that the data is not timely enough and does not match the website page content. Therefore, the present invention uses the Requests library to directly download the data stored in the BRENDA page to ensure data accuracy and timeliness.
[0079] like Figure 3 As shown in the figure (√ indicates that the conditions are met; × indicates that the conditions are not met), the present invention can be divided into four steps in terms of data collection and processing. In step 1, the present invention uses the Requests library to download the Michaelis constant (K m ) and enzyme turnover number (k cat ) data, and the corresponding UniProt ID, mutation information, reaction substrate name and literature number were collected at the same time. Then, in step 2, the ID mapping tool (identifier mapping tool / ID conversion tool) provided by UniProt was used to obtain the amino acid sequence. For the sequence of the mutant protein, the wild-type amino acid sequence was modified according to the mutation information. In step 3, the ligand identifier was obtained from the compound database by the compound name. Finally, in step 4, the present invention deleted the data for which the sequence or ligand identifier could not be obtained, merged the data with the same protein sequence and ligand identifier, and performed arithmetic averaging on the data labels that needed to be merged.
[0080] Through data cleaning, each enzyme kinetic data point in the dataset is assigned a unique protein sequence and compound ligand identifier. This dataset will be further partitioned to support model training and testing.
[0081] To address the problems of missing UniProt ID information for some data in the database and inconsistent naming conventions in compound databases, the present invention proposes a data enhancement method that supplements protein information based on literature associations and integrates multiple databases to obtain compound SMILES.
[0082] 1.1.1. Protein information supplementation and multi-database compound matching process
[0083] The low data retention rate of enzyme kinetic parameters is mainly caused by two reasons: first, the missing UniProt ID makes it impossible to obtain protein sequences; second, different compound databases use different normalization methods for the same compound names, making it difficult to obtain identifiers in batches through compound names.
[0084] In response to the above problems, the present invention adopts data backfill and multi-database search measures to improve data retention rate.
[0085] For data missing UniProt ID (Universal Protein Database ID), the present invention performs a matching search from UniProt (Universal Protein Database) using the PMID number (unique identification code) of the document, and uses the Selenium tool (automated testing tool) to capture the relevant UniProtKB entries (Universal Protein Database Knowledge Base entries) provided in the document.
[0086] After capturing a large number of proteins, the present invention uses a script to screen out species and EC numbers (Enzyme Commission Numbers) that match the records in BRENDA (Brunswick Enzyme Database), and finally retains the proteins that meet the conditions as supplementary data. The present invention uses the data with UniProt IDs saved by BRENDA as a test set to evaluate the success rate of obtaining the correct protein through this method. The results show that K m The protein complement accuracy of the dataset was 89.51%, and k cat The data set was 90.06%, which verified the reliability of the protein replenishment scheme.
[0087] To improve the matching rate and processing efficiency of SMILES (Simplified Molecular Linear Input Specification), this paper utilizes the EnzymeMap tool (an enzyme mapping tool) to obtain SMILES from compound libraries such as RDKit, PubChem, ChEBI, and OPSIN in a one-stop manner by substrate name. Compared with previous studies, this method significantly improves the accuracy of SMILES matching.
[0088] The above two measures effectively solved the problem of low data retention rate, thereby significantly improving the retention rate of enzyme kinetics data and laying the foundation for building a larger-scale data set.
[0089] 1.2 Data Partitioning Strategy
[0090] In order to test the prediction ability of the model in different scenarios, the present invention uses two data set division methods. m and k cat The enzyme kinetic parameter data was randomly divided into a training set, a validation set, and a test set in an 8:1:1 ratio to form a warm-start dataset. To test the model's predictive performance for untrained proteins, the present invention adopted a more difficult data partitioning method. The data was also divided into a training set, a validation set, and a test set in an 8:1:1 ratio, but without identical or similar proteins between them, to form a cold-start dataset. The specific partitioning method is as follows:
[0091] (1) Hot start dataset: Use the partitioning tool of the scikit-learn package to randomly partition the dataset.
[0092] (2) Cold start dataset: Use the MMSeq2 tool to perform cluster analysis, cluster sequences with a similarity of more than 90% into one category, and assign the protein data in the cluster to the training set (or validation set / test set).
[0093] 1.3 Test set construction method
[0094] On the basis of the data partitioning strategy, the present invention further divides the test set of the hot start data set (hot start test set) and the test set of the cold start data set (cold start test set).
[0095] Mainly divided into wild type and mutant protein test sets, K m With k cat Mutual data and orphan data test set, low K m Enzyme test set. The specific construction method is as follows:
[0096] Wild-type and mutant protein test sets: Divided into wild-type and mutant data based on the protein mutation information in the dataset.
[0097] K m Common data and orphan data test set: First get k cat The enzyme-substrate pairs in the training set are K m The enzyme-substrate pair present in the test concentration k cat The data of the training set is defined as common data; the enzyme-substrate pair does not exist in k cat The data of the training set is defined as orphan data.
[0098] k catCommon data and orphan data test set: First get K m The enzyme-substrate pairs in the training set are k cat The enzyme-substrate pair in the test concentration is present in K m The data of the training set is defined as common data; the enzyme-substrate pair does not exist in K m The data of the training set is defined as orphan data.
[0099] Low K m Enzyme test set: Statistic K m The label density distribution of the dataset. According to the label density distribution, K m Less than 1e -2 The data is defined as low K m Enzyme data.
[0100] High K cat Enzyme test set: statistic k cat The label density distribution of the dataset. According to the label density distribution, k cat Greater than 1e 2 The data is defined as high k cat Enzyme data.
[0101] 2. Model Construction
[0102] The present invention constructs two enzyme kinetic parameter prediction models: a single-task learning enzyme kinetic parameter prediction model (STLKP) and a multi-task learning enzyme kinetic parameter prediction model (MTLKP).
[0103] The MTLKP model adopts the design strategy of bias-enhanced network based on STLKP. While ensuring the single-task prediction accuracy, it improves the performance of the model through multi-task learning.
[0104] 2.1 Single-task model
[0105] The STLKP model adopts a three-layer architecture to achieve accurate prediction of enzyme kinetic parameters.
[0106] Single-task learning enzyme kinetic parameter prediction model framework Figure 4 As shown in the figure (Figure A on the left shows the overall framework of the STLKP model; Figure B on the right shows the details of the low-rank bilinear attention network), the first level is the feature extraction module, which uses a pre-trained model to map protein sequences and ligand molecules to a high-dimensional semantic space; the second level establishes cross-modal feature interaction through a low-rank bilinear attention network, and uses low-rank matrix parameters to reduce the complexity of attention calculation; the third level realizes feature decoding through a multi-layer perceptron, and introduces a residual connection mechanism in the decoding layer, which effectively alleviates the problem of abnormal prediction distribution.
[0107] 2.1.1 Feature Extraction
[0108] The single-task learning enzyme kinetic parameter prediction model freezes the weights of the protein pre-trained model and the molecule pre-trained model to extract protein and molecule features. Each input is first mapped to a two-dimensional feature matrix. The feature extraction module uses pre-trained large-scale protein and molecule models (such as ProtT5 and Uni-Mol) to map the discrete character data into a high-dimensional continuous space.
[0109]
[0110] Where, X protein and X molecule represent the feature matrices of proteins and molecules respectively, The dimension of the feature matrix is L × D, where L is the token length (i.e., the number of amino acids or ligand atoms) and D is the feature dimension of each token. To maintain uniform length, tokens are padded with a special token (PAD). During training, the feature matrix is directly fed into the low-rank bilinear attention network as input, preserving more information.
[0111] Tables 2 and 3 show the five protein pre-trained large models and three molecular pre-trained large models used in this invention. This invention cross-combines the feature extraction methods of the protein and molecular pre-trained large models and tests them on the XGBoost baseline model to select the optimal features.
[0112] Table 2: Protein pre-training model used in the present invention
[0113]
[0114] Table 3: Molecular pre-training large model used in the present invention
[0115]
[0116] 2.1.2 Low-rank bilinear attention network
[0117] The Low-Rank Bilinear Attention Network (STLKP) is the core of the model. It uses an attention mechanism to weight and score protein and molecular features, effectively capturing the correlation and importance between features. The network calculates the interaction between protein sequence and molecular structure and generates an attention weight matrix that reflects the importance of each position in the sequence.
[0118] The low-rank bilinear attention network maps protein sequences and molecular structures into a high-dimensional space, generates feature representations, captures complex patterns of sequences and structures, and provides rich information for subsequent attention calculations. The network then calculates the bilinear mapping between protein features and molecular features, generates an attention weight matrix that reflects the correlation between features, and thus provides fine feature control for the model.
[0119] To reduce the complexity of model parameters, the low-rank bilinear attention network introduces a low-rank technique to decompose the bilinear mapping into the product of two low-rank matrices, thereby reducing the number of parameters and improving computational efficiency. The attention calculation is as follows:
[0120]
[0121] In this way, the network can reduce the risk of overfitting while maintaining high expressive power. ij is the attention matrix calculated by the low-rank bilinear attention mechanism, v i and q j are the representations of protein sequence and substrate features, respectively, W1 and W2 are learnable weight matrices, and b is the bias term.
[0122] Finally, the low-rank bilinear attention network simplifies the feature representation through the pooling layer. The pooling layer integrates the attention weight matrix with the feature matrix to generate a simplified joint representation. The calculation relationship of the joint representation is as follows:
[0123]
[0124] It not only reduces the number of parameters through compression, but also retains the key information of proteins and molecules, thereby improving the generalization ability of the model.
[0125] 2.1.3 Decoder
[0126] The decoder reduces the dimensionality of the high-dimensional fused representation and outputs a numerical label.
[0127] The decoder receives the fused representation of the low-rank bilinear attention network output.
[0128]
[0129] f fusion It is the fusion representation after residual connection, which is mapped to the final prediction result through 3 layers of fully connected layers. The calculation relationship is as follows:
[0130]
[0131] 2.2 Multi-task Model
[0132] Based on the STLKP model, a multi-task learning model MTLKP is constructed based on the learnable bias enhancement network. The knowledge transfer between parameters is achieved through the learnable bias enhancement network. The results of the MTLKP model are shown in Figure 5 , Figure 5 Figure A on the left shows the overall framework of the MTLKP model; Figure B on the right shows the details of the learnable bias-enhanced network, where Loss represents the loss function; L1Loss represents the L1 norm loss; MLP expert represents an expert with a multi-layer perceptron structure; Linear represents the linear layer; and Softmax represents the Softmax function.
[0133] 2.2.1. Learnable Bias Enhanced Network
[0134] The Learnable-Bias Enhanced Network is a key component of the MTLKP model. It uses the hybrid expert model as a framework to share the task-level K m and k cat The network introduces a learnable bias term to provide linear adjustment for prediction tasks, enhancing the model's ability to predict specific enzyme kinetic parameters.
[0135] Learnable bias-enhanced network extracts K through shared layers m and k cat The common features of the data, this layer consists of multiple expert models, each of which captures different patterns in the data. For multi-task models, the task subscript t is introduced to distinguish tasks (e.g. K m or k cat ), the expert’s calculation formula is as follows:
[0136]
[0137] is the fusion representation of the t-th task, F i is the i-th expert network mapping function, is the output of the i-th expert for the t-th task. The task-specific layer selects the most relevant expert output through a gating mechanism and linearly combines them to generate a task-specific representation. The calculation process is as follows:
[0138]
[0139] G i is the gated score of the i-th expert, Softmax is the Softmax function, Linear is the linear layer transformation, and flatten is Flattened into a one-dimensional vector, N expert is the number of experts, S (t)is the shared representation of the tth task. Finally, the task-specific layer outputs a learnable bias term to provide a bias for the final prediction. The bias is calculated as follows:
[0140] b (t) =mean(S (t) ,dim=-1)
[0141] b (t) is the bias of the t-th task, and mean (dim = -1) is the average of the hidden layer (the penultimate dimension) of the task-specific shared representation, which is then added to the input of the linear layer.
[0142] 2.2.2 Data Flow
[0143] The data flow of the multi-task model is similar to that of the single-task model. The difference is that the low-rank bilinear attention network and the decoder of the fully connected layer of the multi-task model are task-specific, which is K m or k cat Create a separate network to learn the task information. m and k cat After the data are mixed and represented by their respective low-rank bilinear networks, they are input into the learnable bias-enhanced network to jointly adjust the weight parameters of the network.
[0144] This approach enables the multi-task model to not only capture the specific features of each task, but also learn the common patterns of the two sets of data through a shared network, thereby improving generalization ability and prediction accuracy.
[0145] The data flow of the multi-task model is as follows:
[0146] (1) Feature extraction: Similar to the single-task model, the features of proteins and molecules are extracted through a pre-trained model.
[0147] (2) Low-rank bilinear attention network: Calculate the bilinear mapping between protein and molecular features, generate the attention weight matrix, and perform pooling to obtain a joint representation.
[0148] (3) Learnable bias-enhanced network: Shared layer gets K m and k cat The joint representation of the network, the task-specific layer selects the most relevant expert output through the gating mechanism and generates a task-specific representation, which is then average-pooled into the linear bias of the task, which is a specific number.
[0149] (4) Decoder: First, the fused representation and task-specific bias are added together, and then mapped to the final prediction result through a linear layer to obtain K m or k cat Output.
[0150] 3. Model Training
[0151] To improve model performance, this paper employs a variety of training techniques during the training phase to help the model converge more fully and quickly. This section describes in detail the strategies and techniques used in model training.
[0152] 3.1 Training Strategy
[0153] The present invention adopts different training strategies for STLKP and MTLKP, and the training strategies of STLKP and MTLKP will be described below respectively.
[0154] STLKP training strategy: Single-stage training is adopted. First, the protein and molecule pre-training large models are frozen. Only in K m or k cat Train on a single dataset. Save the model when the validation loss reaches a new low.
[0155] MTLKP training strategy: divided into two stages of training: multi-task pre-training and specific task fine-tuning.
[0156] The first stage is multi-task pre-training: freeze the large protein and molecule pre-trained models and train them using the full dataset from both sets. Save the model when the sum of the losses on the two test sets reaches a new low.
[0157] The second stage is the single-task fine-tuning stage: freeze other networks including shared layers and only m or k cat Train the decoder network weights for this task on a single dataset. Save the model when the test set loss for this task reaches a new low.
[0158] 3.2 Training Techniques
[0159] Label normalization, Dropout tuning, and learning rate warmup were used during training.
[0160] Label Normalization: K m With k cat The label range can span more than 10 orders of magnitude, and the fluctuation of data at this order of magnitude is huge. For example, the calculation formula of the fully connected layer is y=wx+b. When the feature x is normalized, the weight w needs to be 1e10 to cover the label to 10 orders of magnitude. This means that any perturbation to x will cause the predicted value change to be scaled by 1e10 times, thereby increasing the difficulty of training loss convergence. Therefore, the present invention has a better understanding of K m and k cat All labels are normalized using log10, which reduces the label span and effectively reduces the difficulty of model training.
[0161] Dropout tuning: Dropout is a regularization technique that randomly shuts down some neurons during training. It aims to improve the generalization ability of a model and is widely used in tasks such as image classification, speech recognition, and natural language processing. However, the proposed model does not use dropout tuning because it can cause performance degradation in both multi-task and single-task models. This is because dropout changes the variance of the model output, causing it to be inconsistent with the variance of the training data, making it unhelpful for regression tasks.
[0162] Learning rate warmup: A large initial learning rate can cause optimization instability and make it difficult to reach the optimal solution. Warm-up techniques use a predefined learning rate decay to allow the model to fully explore the parameter space at a smaller learning rate, making it easier to reach the optimal solution. We tested various warmup strategies and ultimately adopted the cosine annealing strategy.
[0163] 4. Construction and training of other models
[0164] In the feature selection and model performance comparison stages, this paper trains multiple models including the XGBoost baseline. The following sections describe the construction and training methods of each model.
[0165] XGBoost baseline model: This paper uses the XGBRegressor class provided by the xgboost package to build a baseline model for feature cross-evaluation. The default parameters of XGBRegressor are used during training.
[0166] UniKP (papermodel): Open source model and weights of the original UniKP paper.
[0167] UniKP (retrain model): Retrain the UniKP open-source model using the newly constructed dataset. Since the original paper did not perform hyperparameter tuning during training, this paper first uses the GridSearchCV class provided by the scikit-learn package to perform a cross-hyperparameter search, and then retrains at the optimal parameters.
[0168] Km_prediction(papermodel): Open source model and weights of the original Km_prediction paper.
[0169] DLKcat (papermodel): Open source model and weights of the original DLKcat paper.
[0170] DLKcat (retrainmodel): Retrain the DLKcat open-source model using a newly constructed dataset. Since the original paper performed hyperparameter tuning during training, this paper directly uses the optimal parameters from the DLKcat paper for retraining.
[0171] 5. Model Evaluation Metrics
[0172] In K m Prediction and k cat In the regression task of prediction, the present invention adopts the commonly used regression indicators: mean absolute error (MAE), root mean square error (RMSE), coefficient of determination (R 2 ) and Pearson correlation coefficient (PCC).
[0173] MAE measures the mean absolute deviation between predicted and true values, directly reflecting the model's prediction error in real-world applications. Smaller MAE values indicate more accurate model predictions. RMSE, a weighted squared error metric, is significantly affected by larger errors. Compared to MAE, RMSE is more sensitive to fluctuations in model performance when exposed to outliers.
[0174] R 2 The value reflects the correlation between the predicted results and the true values. The closer the value is to 1, the better the model fits the data. PCC is used to measure the strength of the linear relationship between the predicted results and the true values. Its value ranges from -1 to 1. The closer the PCC value is to 1, the stronger the linear relationship between the model prediction results and the true data.
[0175] Typically, R 2 It is suitable for evaluating the fitting effect of the regression model, especially when there is a certain deviation between the predicted value and the true value. 2 It can effectively reflect the accuracy of the model. Therefore, when evaluating regression tasks, other indicators should be combined to comprehensively evaluate the predictive ability of the model.
[0176] The calculation formulas for the above four indicators are as follows:
[0177]
[0178] In actual biological scenarios, using K m and k catThe main purpose is to find enzymes with high binding strength and fast catalytic rate. Therefore, the present invention proposes frequency-weighted indices, namely frequency-weighted mean absolute error (Frequency-weighted MAE) and frequency-weighted root mean square error (Frequency-weighted RMSE), to measure the model's performance on low K m Value or high k cat Value enzyme prediction performance.
[0179] Because K m and k cat It is a continuous value. First, the data is binned according to the values in the data set. The bins in the log10 scale of the present invention are [-10, -9, -8…, 8, 9, 10]. Then the amount of data in each bin N is counted. i , the weight of each box data is calculated as follows:
[0180]
[0181] The frequency-weighted index is calculated as follows:
[0182]
[0183] Where y i represents the true value of the data, Represents the model’s predicted value for the data. represents the mean of the true values, Represents the mean of the predicted value, n represents the number of samples, N represents the number of samples in the box, and bin represents the number of boxes.
[0184] The present invention first fuses protein and molecular features through a low-rank bilinear attention network. Compared with the feature fusion method based on linear splicing, this model-level fusion method has a higher K m The linear splicing is K m With k catFeature fusion is a commonly used method in the prediction field, but no research has yet explored this aspect in depth. The bilinear attention network is a second-order feature fusion method that exhibits more robust performance by selecting features from the high-dimensional space of protein and molecular features. Unlike linear concatenation, which assumes that all amino acids have the same weight, the bilinear attention network introduces an attention mechanism to learn weight coefficients for each amino acid-atom pair. In fact, during the enzyme-substrate binding process, some amino acids play a key role at the binding site. Therefore, the attention mechanism can assign higher weights to important amino acids, thereby better matching the enzyme-substrate binding mechanism.
[0185] Compared with the deep learning model STLKP-concat based on linear splicing features, UniKP performs significantly better than STLKP-concat. Their main differences lie in the feature extraction and fitting algorithms. The cross-experimental results of feature combination show that the feature extraction method (Prot-T5+UniMol) adopted in the present invention has only a slight advantage over the feature extraction method (Prot-T5+SMILES Transformer) of UniKP, so the impact of feature extraction differences on performance can be ruled out. Compared with the fully connected layer, the extreme tree algorithm adopted by UniKP involves complex mathematical processing. For example, decision trees can efficiently decompose high-dimensional data into smaller subsets, and through the combination of multiple decision trees, the model variance is reduced, which helps to improve the generalization ability of the model. By improving the feature fusion method, the performance of STLKP is slightly better than that of UniKP, indicating that the improvement of the feature fusion network plays an important role in the effect of the deep learning model.
[0186] The present invention designs a prediction model based on multi-task learning to make full use of the correlation of enzyme kinetic parameters, thereby further improving the prediction performance. m The R2 on the warm-start test set (0.6403) is 3.3% higher than the previous state-of-the-art UniKP model (0.6201); cat The R2 on the hot start dataset (0.6576) is 5.1% higher than that of UniKP (0.6259). Since the binding of enzymes and substrates and the catalytic process are closely related, whether the model performance is improved depends on the K m and k cat Whether it can explain the enzyme-substrate binding and catalytic process. The results show that the multi-task model effectively improves the prediction performance and verifies the K m and k cat It plays a central role in quantifying enzyme-substrate binding and catalytic processes.
[0187] The present invention is not limited to the structures described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A method for constructing an enzyme kinetic parameter prediction model, characterized in that: include: By integrating multi-source data, we construct an enzyme kinetic parameter set, then perform data cleaning and data enhancement in sequence to supplement protein information and match compound data. Based on the data partitioning strategy, we divide the enzyme kinetic parameter set into different types of data sets. Constructing a multi-task learning enzyme kinetic parameter prediction model that outputs enzyme kinetic parameter prediction results by modeling the synergistic relationship between the Michaelis constant and enzyme turnover number data; the enzyme kinetic parameter prediction model includes a feature extraction network, a low-rank bilinear attention network, a learnable bias enhancement network, and a decoder; Based on a multi-stage training strategy, the enzyme kinetic parameter prediction model is trained in segments. During the training process, label normalization, random inactivation tuning and learning rate warm-up are introduced for optimization to generate an enzyme kinetic parameter prediction model with optimal parameters.
2. The method for constructing an enzyme kinetic parameter prediction model according to claim 1, wherein The enzyme kinetic parameter set is constructed by integrating multi-source data, and data cleaning and data enhancement are performed in sequence to achieve protein information supplementation and compound data matching; and based on the data partitioning strategy, the enzyme kinetic parameter set is divided into different types of data sets including: Download Michaelis constants and enzyme turnover numbers from the Brunswick Enzyme Database, along with corresponding Universal Protein Database IDs, mutation information, reaction substrate names, and literature numbers. Use ID conversion tools to obtain amino acid sequences, and then, based on compound names, obtain ligand identifiers from the compound database. After integrating multi-source data, a set of enzyme kinetic parameters is constructed. Delete the data for which the sequence or ligand identifier cannot be obtained, merge the data with the same amino acid sequence and ligand identifier, and perform arithmetic averaging on the data labels that need to be merged; Protein information missing from universal protein databases is supplemented through data backfilling, and identical compound names are standardized through multi-database searches. The enzyme kinetic parameter set was divided into training set, validation set and test set in a ratio of 8:1:1; and based on whether there were identical or similar proteins between the training set, validation set and test set, the enzyme kinetic parameter set was divided into hot start data set and cold start data set. The hot start test set or cold start test set was then functionally divided to obtain multiple test subsets.
3. The method for constructing an enzyme kinetic parameter prediction model according to claim 2, wherein In the process of obtaining the amino acid sequence using the ID conversion tool, the amino acid sequence of the mutant protein is modified on the wild-type amino acid sequence according to the mutation information.
4. The method for constructing an enzyme kinetic parameter prediction model according to claim 2, wherein The method of supplementing protein information of missing universal protein database IDs through data backfilling and normalizing identical compound names through multi-database searching includes: The unique identification code of the literature was used to match the universal protein database and search, and the universal protein database knowledge base entries in the literature were captured. The species and enzyme commission numbers that matched the Brunswick enzyme database records were screened by scripts, and the proteins that met the conditions were retained as supplementary data. By using the enzyme map tool, simplified molecular linear input specifications can be obtained from the compound library in one stop through the reaction substrate name to improve the accuracy of simplified molecular linear input specification matching.
5. The method for constructing an enzyme kinetic parameter prediction model according to claim 2, wherein The test subsets include a wild-type and mutant protein test set, a Michaelis constant common data and orphan data test set, an enzyme turnover number common data and orphan data test set, a low Michaelis constant enzyme test set, and a high enzyme turnover number enzyme test set.
6. The method for constructing an enzyme kinetic parameter prediction model according to claim 1, wherein The multi-task learning enzyme kinetic parameter prediction model is constructed, and by modeling the synergistic relationship between the Michaelis constant and the enzyme turnover number data, the prediction results of the enzyme kinetic parameters are output including: A feature extraction network is used to capture protein and molecular features, mapping each input into a two-dimensional feature matrix. Protein and molecular models are then used to map discrete character data into a high-dimensional continuous space. A low-rank bilinear attention network is used to calculate the bilinear mapping between protein features and molecular features, generate an attention weight matrix, and then pool the matrix to obtain a joint representation, which is then input into a learnable bias enhancement network. The learnable bias-augmented network receives the joint representation of the Michaelis constant and enzyme turnover number data, selects the expert output through the task-specific layer, generates a task-specific representation, and then average-pools it into a linear bias for the task; The fused representation is added to the task-specific representation bias, the high-dimensional fused representation is reduced in dimension, and mapped to the final model output through a linear layer to obtain the prediction results of enzyme kinetic parameters.
7. The method for constructing an enzyme kinetic parameter prediction model according to claim 6, wherein: The method of using a low-rank bilinear attention network to calculate the bilinear mapping between protein features and molecular features, generating an attention weight matrix, obtaining a joint representation through pooling, and inputting the result into a learnable bias enhancement network includes: Mapping the amino acid sequence and molecular structure of a protein into a high-dimensional space to generate a feature representation that captures the complex patterns between the sequence and structure. It also calculates the bilinear mapping between protein features and molecular features, generating an attention weight matrix that reflects the correlation between features. A low-rank strategy is introduced to decompose the bilinear mapping into the product of two low-rank matrices, and the attention weight matrix and the feature matrix are integrated through the pooling layer to generate a simplified joint representation.
8. The method for constructing an enzyme kinetic parameter prediction model according to claim 6, wherein: The learnable bias augmentation network receives the joint representation of the Michaelis constant and the enzyme turnover number data, selects the expert output through the task-specific layer, generates a task-specific representation, and then average-pools it into a linear bias for the task, including: The shared features of the Michaelis constant and enzyme turnover number data are extracted through the shared layer. The task-specific layer selects the most relevant expert outputs through a gating mechanism and linearly combines them to generate task feature representations. The task-specific layer adjusts the learnable bias term by output, provides bias for the prediction results, and connects the shared features to the task-specific network in the form of a learnable bias residual.
9. The method for constructing an enzyme kinetic parameter prediction model according to claim 1, wherein The stage training strategy includes a multi-task pre-training stage and a single-task fine-tuning stage. The multi-task pre-training stage freezes the protein and molecule pre-training large models and uses the full data of the two sets of data for training; the single-task fine-tuning stage freezes other networks including shared layers and only trains the decoder network weights of the task on a single data set of Michaelis constant or enzyme turnover number data.
10. A system for constructing an enzyme kinetic parameter prediction model, characterized in that: include: The data acquisition module is used to construct enzyme kinetic parameter sets by integrating multi-source data, and then perform data cleaning and data enhancement in sequence to achieve protein information supplementation and compound data matching; and based on the data partitioning strategy, the enzyme kinetic parameter set is divided into different types of data sets; The model building module is used to fuse protein features and compound molecular features using a low-rank bilinear attention network, combined with a learnable bias enhancement network, to build a multi-task learning enzyme kinetic parameter prediction model; a decoder is used to reduce the dimensionality of the high-dimensional fusion and output the prediction results of the enzyme kinetic parameters; The model training module is used to perform segmented training on the enzyme kinetic parameter prediction model based on a multi-stage training strategy. During the training process, label normalization, random inactivation tuning, and learning rate warm-up are introduced for optimization to generate an enzyme kinetic parameter prediction model with optimal parameters.