A method and system for predicting the function of biological reagents

By collecting information from public databases of biological reagents and training neural network models, the problem of low accuracy in traditional biological reagent function prediction methods has been solved, achieving efficient and accurate biological reagent function prediction and promoting the research and application of new biological reagents.

CN119832988BActive Publication Date: 2025-10-28林爱珊
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411904362.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-28
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Traditional methods for predicting the function of biological reagents suffer from low accuracy, as they fail to fully consider the characteristics of molecular structural domains and molecular functional domains, leading to inaccurate predictions of biological reagent functions.

Method used

By collecting information from publicly available databases of biological reagents, a set of biomolecular information was obtained, including protein sequence data, molecular structural domain data, and molecular functional domain data. A functional prediction network framework was constructed using convolutional neural networks (CNN), recurrent neural networks (RNN), and Transformer neural networks. The model was trained by combining the activity, structural stability, and specific functional targets of the biological reagents to make functional predictions. The model was then validated and optimized through experiments.

Benefits of technology

It improves the accuracy and efficiency of predicting the function of biological reagents, enabling the prediction of their functional performance under different biological or pharmacological conditions, simplifying experimental design, reducing uncertainty and reproducibility, and promoting the research and development and application of novel biological reagents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832988B_ABST
    Figure CN119832988B_ABST
Patent Text Reader

Abstract

This invention relates to the field of biological function prediction technology, and more particularly to a method and system for predicting the function of biological reagents. The method includes the following steps: collecting biomolecular information from a publicly available database of biological reagents to obtain a biomolecular information set; performing biofeedback analysis on the biomolecular information set to obtain protein sequence alignment features, molecular structural domain features, and specific molecular functional domain features of the biological reagents; constructing a biological reagent function prediction network framework and training a function prediction model to generate a biological reagent function prediction model; acquiring the molecular features of the biological reagent to be predicted and performing target function prediction to obtain the target biological reagent function prediction result; and performing function verification and optimization based on the target biological reagent function prediction result to generate an optimized target biological reagent function result. This invention enables effective prediction and screening of the function of biological reagents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biological function prediction technology, and in particular to a method and system for predicting the function of biological reagents. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence and machine learning technologies, data-driven methods for predicting the function of biological reagents have gradually become a research hotspot. These methods, by analyzing large-scale biological data and utilizing algorithms such as deep learning, support vector machines, and random forests, can extract complex patterns and perform functional predictions over a wider range. Machine learning methods can better integrate biological information from different sources, such as genomic data, proteomic data, and literature data, improving the accuracy and efficiency of predictions. However, traditional functional prediction methods have some limitations. For example, many biological reagents have diverse functions, exhibiting different biological activities under different environments or conditions, making it difficult to comprehensively consider both molecular structural domain characteristics and molecular functional domain characteristics, thus leading to lower accuracy in predicting the function of biological reagents. Summary of the Invention

[0003] Therefore, it is necessary for the present invention to provide a method and system for predicting the function of biological reagents in order to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objective, a method for predicting the function of a biological reagent includes the following steps:

[0005] Step S1: Collect biomolecular information from the public database of biological reagents to obtain a biomolecular information set of biological reagents, which includes protein sequence data, molecular structural domain data, and functional domain data of biological reagents.

[0006] Step S2: Perform sequence alignment analysis on the biological reagent protein sequence data to obtain the biological reagent protein sequence alignment features; perform protein structure feature prediction analysis on the biological reagent molecular domain data to generate biological reagent molecular domain features; extract specific molecular functional features based on the biological reagent molecular functional domain data to obtain specific molecular functional domain features of the biological reagent.

[0007] Step S3: Construct a biological reagent function prediction network framework, and input the biological reagent protein sequence alignment features, biological reagent molecular structural domain features, and biological reagent specific molecular functional domain features into the biological reagent function prediction network framework to train and construct a function prediction model to generate a biological reagent function prediction model; obtain the molecular features of the biological reagent to be predicted, and input the molecular features of the biological reagent to be predicted into the biological reagent function prediction model to predict the target function to obtain the target biological reagent function prediction result;

[0008] Step S4: Based on the function prediction results of the target biological reagent, perform functional verification and optimization on the corresponding biological reagent to be predicted, so as to generate the function optimization results of the target biological reagent.

[0009] Furthermore, step S1 includes the following steps:

[0010] Step S11: Perform preliminary screening of raw biological reagent data in the public database of biological reagents to obtain a set of raw biological information of biological reagents;

[0011] Step S12: Extract protein sequences from the original biological information set of biological reagents to obtain biological reagent protein sequence data, including the amino acid molecular sequence and gene molecular sequence corresponding to each biological reagent.

[0012] Step S13: Perform molecular domain identification, labeling, and extraction on the corresponding amino acid molecular sequences and gene molecular sequences within the biological reagent protein sequence data to obtain biological reagent molecular domain data, including secondary domain sequence fragments and tertiary domain sequence fragments corresponding to each biological reagent.

[0013] Step S14: Perform molecular functional domain annotation and analysis on the molecular structural domain data of biological reagents to obtain molecular functional domain data of biological reagents, including the enzyme catalytic activity, binding affinity and substance transport and degradation functional domain fragments corresponding to each biological reagent.

[0014] Step S15: Integrate the biological reagent protein sequence data, biological reagent molecular structural domain data, and biological reagent molecular functional domain data to obtain a biological reagent biomolecular information set.

[0015] Furthermore, the original biological information set of biological reagents in step S11 includes original biological information data of enzymes, antibodies, and microbial types.

[0016] Furthermore, step S13 includes the following steps:

[0017] Step S131: Perform molecular domain identification and prediction on the corresponding amino acid molecular sequences and gene molecular sequences within the biological reagent protein sequence data to obtain the amino acid molecular domain region sequences and gene molecular domain region sequences of the biological reagent.

[0018] Step S132: Obtain the corresponding amino acid molecular structure sites and amino acid molecular interaction forces through the amino acid molecular structure domain region sequence of the biological reagent, and perform sequence structure hydrophobicity evaluation analysis on the corresponding amino acid molecular structure domain region sequence of the biological reagent based on the amino acid molecular structure sites and amino acid molecular interaction forces to obtain the hydrophobicity score of the amino acid molecular structure of the biological reagent.

[0019] Step S133: Perform molecular structure functional charge analysis on the domain sequence of the amino acid molecule of the biological reagent to obtain the molecular structure functional charge characteristics of the amino acid molecule of the biological reagent, including the molecular structure charge distribution, the molecular structure functional charge coupling degree, and the molecular structure charge density of the amino acid molecule.

[0020] Step S134: Based on the hydrophobicity score and functional charge characteristics of the amino acid molecular structure of the biological reagent, the corresponding domain sequence of the amino acid molecular structure of the biological reagent is measured using the protein secondary structure measurement formula to obtain the secondary structure degree of the amino acid molecular region sequence; based on the secondary structure degree of the amino acid molecular region sequence, the corresponding domain sequence of the amino acid molecular structure of the biological reagent is subjected to secondary structure labeling to obtain the secondary structure domain sequence fragment corresponding to the biological reagent.

[0021] Step S135: Perform molecular tertiary structure deduction analysis on the molecular structural domain region sequence of the biological reagent gene to obtain the corresponding tertiary structural domain sequence fragment of the biological reagent.

[0022] Furthermore, the specific formula for calculating the protein secondary structure measurement in step S134 is as follows:

[0023]

[0024] In the formula, S s Let L be the secondary structure degree of the amino acid molecule's domain sequence, r be the spatial position parameter of the domain, H(r) be the hydrophobicity score of the amino acid molecule's domain sequence at spatial position r, and α be the degree of secondary structure degree of the amino acid molecule. h Here, C(r) represents the hydrophobicity weighting factor associated with amino acids, C(r) is the functional charge distribution of the amino acid molecular region sequence at spatial position r, O(r) is the functional charge coupling degree of the amino acid molecular region sequence at spatial position r, M(r) is the structural charge density of the amino acid molecular region sequence at spatial position r, and β... c γ is the charge characteristic weighting factor associated with amino acids, F(r) is the functional group characteristic parameter of the amino acid molecular region sequence at spatial position r, and γ is the charge characteristic weighting factor associated with amino acids. f η is the weighting factor for the functional group characteristics associated with amino acids, and η is the correction coefficient for the secondary structure degree of the amino acid molecular region sequence.

[0025] Furthermore, step S135 includes the following steps:

[0026] Gene molecular dynamics simulation analysis was performed on the molecular structural domain sequences of biological reagent genes to generate the biological reagent gene molecular dynamics simulation process.

[0027] Gene molecular folding path deduction was performed on the molecular dynamics simulation process of biological reagent genes to generate a gene molecular folding deduction path diagram of biological reagents;

[0028] Based on the molecular folding deduction path map of biological reagent genes, the corresponding molecular structural domain sequence of biological reagent genes is analyzed for domain folding energy distribution to obtain the molecular structural domain folding energy distribution map of biological reagent genes.

[0029] The folding energy distribution map of the gene molecular structure domain of the biological reagent is used to obtain the corresponding gene molecular folding energy stability score and gene molecular folding energy distribution gradient. Based on the gene molecular folding energy stability score and gene molecular folding energy distribution gradient, the protein tertiary structure matching calculation formula is used to match the corresponding biological reagent gene molecular structure domain region sequence to obtain the tertiary structure matching degree of the gene molecular region sequence.

[0030] The specific formula for calculating protein tertiary structure matching is as follows:

[0031]

[0032] In the formula, M represents the tertiary structure matching degree of the gene molecular region sequence, Ω represents the gene molecular folding space range, and x represents the gene molecular folding space coordinates. Let δ be the Hamiltonian of the folding energy at coordinate x in the gene folding space, and δ be the adjustment coefficient for the gene folding energy. U is the gradient symbol. fold (x) represents the gene folding energy stability score at coordinate x in the gene folding space, and ε is the parameter affecting gene folding energy stability. ξ is the energy distribution gradient of gene molecule folding, θ is the influence coefficient of the energy distribution gradient of folding, and ξ is the correction coefficient of the tertiary structure matching degree of gene molecule region sequence.

[0033] Based on the tertiary structure matching degree of gene molecular region sequences, the tertiary structure of the corresponding biological reagent gene molecular structural domain sequences is labeled to obtain the corresponding tertiary structural domain sequence fragments of the biological reagent.

[0034] Furthermore, step S15 includes the following steps:

[0035] Step S151: Perform functional domain segmentation on the corresponding secondary structure domain sequence fragments within the biological reagent molecule structural domain data to obtain the functional domain type segmentation of the secondary structure of the biological reagent molecule, including the enzyme functional domain fragments and the binding functional domain fragments of the biological reagent molecule.

[0036] Step S152: Obtain the known catalytic enzyme reaction mechanism and enzyme substrate characteristics, and perform catalytic potential prediction analysis on the enzyme functional domain fragments of biological reagent molecules based on the known catalytic enzyme reaction mechanism and enzyme substrate characteristics to generate a biological reagent molecule catalytic potential score; perform enzyme catalytic activity annotation processing on the corresponding biological reagent molecule enzyme functional domain fragments based on the biological reagent molecule catalytic potential score to obtain the enzyme catalytic activity functional domain fragments corresponding to each biological reagent.

[0037] Step S153: By simulating different environmental constraints, and performing dynamic annotation of binding affinity of the binding functional domain fragments of biological reagent molecules based on different environmental constraints, the binding affinity functional domain fragments corresponding to each biological reagent are obtained.

[0038] Step S154: Perform transmembrane transport capability prediction analysis on the corresponding tertiary domain sequence fragments within the molecular structural domain data of biological reagents to obtain the predicted value of transmembrane transport capability of biological reagent molecules.

[0039] Step S155: Based on the predicted transmembrane transport capacity of the biological reagent molecular structure, the corresponding tertiary domain sequence fragments are annotated with material transport and degradation data to obtain the material transport and degradation functional domain fragments corresponding to each biological reagent.

[0040] Furthermore, step S2 includes the following steps:

[0041] Step S21: Perform inter-sequence alignment analysis on each protein sequence in the biological reagent protein sequence data to obtain the inter-sequence alignment matrix of the biological reagent protein sequence, which includes the alignment score, matching length and gap region between each protein sequence.

[0042] Step S22: Based on the alignment matrix between biological reagent protein sequences, perform sequence difference statistical analysis on the corresponding protein sequences within the biological reagent protein sequence data to obtain the alignment features of biological reagent protein sequences, including the differences in amino acid residue substitution, insertion, and deletion between protein sequences.

[0043] Step S23: Perform protein structure feature prediction analysis on the molecular domain data of biological reagents to generate molecular domain features of biological reagents, including molecular domain stability, molecular domain hydrophobic interaction force and molecular domain binding energy features.

[0044] Step S24: Extract specific molecular functional features based on the functional domain data of biological reagents to obtain specific molecular functional domain features of biological reagents, including enzyme activity levels, binding affinity, and substance transport and degradation capabilities of biological reagents.

[0045] Furthermore, step S3 includes the following steps:

[0046] Step S31: Construct a biological reagent function prediction network framework by connecting convolutional neural networks (CNN), recurrent neural networks (RNN), and transformer neural networks;

[0047] Step S32: Input the protein sequence alignment features of the biological reagent, the molecular structural domain features of the biological reagent, and the specific molecular functional domain features of the biological reagent into the biological reagent function prediction network framework to train and construct the function prediction model. At the same time, the activity, structural stability, and specific function prediction targets of the biological reagent are considered to generate the biological reagent function prediction model.

[0048] Step S33: Obtain the molecular features of the biological reagent to be predicted, including the protein sequence alignment features, molecular structural domain features, and molecular functional domain features corresponding to the biological reagent to be predicted.

[0049] Step S34: Input the molecular features of the biological reagent to be predicted into the biological reagent function prediction model to predict the target function and obtain the target biological reagent function prediction result.

[0050] Furthermore, the present invention also provides a biological reagent function prediction system for performing the biological reagent function prediction method described above, the biological reagent function prediction system comprising:

[0051] The biological reagent molecular information acquisition module is used to acquire biomolecular information from public databases of biological reagents to obtain a biological reagent biomolecular information set, which includes biological reagent protein sequence data, biological reagent molecular structural domain data, and biological reagent molecular functional domain data.

[0052] The biological reagent molecular feature analysis module is used to perform sequence alignment analysis on biological reagent protein sequence data to obtain biological reagent protein sequence alignment features; to perform protein structure feature prediction analysis on biological reagent molecular domain data to generate biological reagent molecular domain features; and to extract specific molecular functional features based on biological reagent molecular functional domain data to obtain specific molecular functional domain features of biological reagents.

[0053] The biological reagent target function prediction module is used to construct a biological reagent function prediction network framework. It inputs the biological reagent protein sequence alignment features, biological reagent molecular structural domain features, and biological reagent specific molecular functional domain features into the biological reagent function prediction network framework to train and construct the function prediction model, thereby generating the biological reagent function prediction model. It also acquires the molecular features of the biological reagent to be predicted and inputs these features into the biological reagent function prediction model to predict the target function, thereby obtaining the target biological reagent function prediction result.

[0054] The biological reagent function verification and optimization module is used to perform function verification and optimization on the corresponding biological reagent to be predicted based on the function prediction results of the target biological reagent, so as to generate the function optimization results of the target biological reagent.

[0055] The beneficial effects of this invention are:

[0056] 1. Compared with the prior art, the beneficial effect of the biological reagent function prediction method proposed in this invention lies in the fact that by collecting information from public databases of biological reagents, a set of biomolecular information containing protein sequence data, molecular structural domain data, and molecular functional domain data is obtained. The key to this step is that public databases gather a large amount of verified and researched data, making the collected biomolecular information high-quality and reliable. Protein sequence data is the foundation for studying the structure and function of biomolecules. It can provide basic genetic information of biological reagents and help researchers understand their functions in different biological systems. Molecular structural domain data helps to identify functional regions of proteins, such as enzyme active sites and receptor binding sites, thereby revealing the relationship between molecular function and structure. The collection of functional domain data can help researchers identify the role of molecules in specific biological processes, thereby providing direction for further functional research and application development. This information set not only contains the basic molecular sequence data of each biological reagent, but also systematically summarizes the structural features and functional properties of the molecule, providing rich data for further biological research, avoiding the limitations of a single data source, and thus improving the efficiency of data utilization. Secondly, sequence alignment analysis of biological reagent protein sequence data allows for the identification of sequence alignment characteristics, enabling in-depth exploration of protein sequence variation. This helps reveal the functional differences of biological reagents at the molecular level, identify the regularity of these variations and their functional significance in specific organisms, and thus conduct more in-depth and comprehensive research on the function and variability of biological reagent proteins. Simultaneously, predictive analysis of protein structural features based on biological reagent molecular domain data provides a deeper understanding of protein structural regions and characteristics. Protein domains are fundamental to their function; different domains possess different functional characteristics. Predicting the stability, hydrophobic interactions, and binding energies of these domains provides profound insights into the functional mechanisms of proteins, helping scientists more accurately understand the domain characteristics of the corresponding protein molecules and thus advancing the process of predicting the function of biological reagents. Furthermore, by extracting specific molecular functional characteristics based on the functional domain data of biological reagents, their functional roles in biological processes can be comprehensively evaluated. This process not only involves the extraction of features such as protein enzymatic activity, binding affinity, and material transport and degradation capabilities, but also provides important data support for the functional prediction of biological reagents. This step, through in-depth analysis of these functional domain characteristics, can optimize the application performance of biological reagents, providing strong data support and theoretical basis for the application development of biological reagents, thus fully considering the functional domain characteristics of the corresponding protein molecules of biological reagents.Then, a biological reagent function prediction network framework was constructed by combining convolutional neural networks (CNN), recurrent neural networks (RNN), and Transformer neural networks. Protein sequence alignment features, molecular domain features, and specific molecular functional domain features of biological reagents were input into this framework for training the function prediction model, establishing a multi-dimensional biological reagent function prediction model. Protein sequence alignment features help identify amino acid sequence patterns related to known protein functions, facilitating functional inference of similar protein families. Molecular domain features help capture the tertiary structure and other important functional domains of proteins, providing structural information closely related to molecular function. Specific molecular functional domain features reveal regions related to certain biological functions (such as enzyme activity, receptor binding, signal transduction, etc.). The input of these features significantly improves the model's performance on complex biological data. By combining the activity, structural stability, and specific functional prediction targets of biological reagents for training, the model can go beyond learning surface features and comprehensively analyze molecular functions from different levels, thereby improving the diversity and accuracy of biological reagent prediction results. Furthermore, by inputting the molecular characteristics of the biological reagent to be predicted into a biological reagent function prediction model, the target function prediction result of the target biological reagent is obtained. The key to this step is to combine real biomolecular characteristics with a trained prediction model to provide biologically meaningful prediction results. Through the efficient computation and pattern recognition capabilities of the neural network model, key information can be extracted from a large number of molecular characteristics in a very short time, and the functional performance of the biological reagent in a specific biological or pharmacological environment can be predicted, such as its potential activity in enzyme catalysis, receptor binding, antigen recognition, etc. This helps to quickly screen potential biological reagents, thereby greatly improving the efficiency and success rate of function prediction. Finally, the predicted biological reagents are functionally validated and optimized based on the predicted results of the target biological reagents. The key to this stage is that the functional prediction results based on the prediction model can systematically guide experimental design, reduce uncertainty and repeatability in experiments, and improve validation efficiency. If the experimental results are consistent with the prediction, the reliability of the functional prediction model can be confirmed; if the experimental results are inconsistent with the prediction, the model parameters can be further adjusted or the dataset can be expanded to continuously optimize the model. This process ultimately promotes the qualitative and quantitative analysis of the functions of biological reagents, and also promotes the research and development and practical application of new biological reagents, thereby improving the accuracy of functional prediction.

[0057] 2. The biological reagent function prediction system proposed in this invention is composed of a biological reagent molecular information acquisition module, a biological reagent molecular feature analysis module, a biological reagent target function prediction module, and a biological reagent function verification and optimization module. It can realize any biological reagent function prediction method described in this invention. It is used to combine the operations between the computer programs running on each module to realize the biological reagent function prediction method. The internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower input, and can quickly and effectively provide a more accurate and efficient biological reagent function prediction process, thereby simplifying the operation process of the biological reagent function prediction system. Attached Figure Description

[0058] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0059] Figure 1 This is a schematic diagram of the steps in the biological reagent function prediction method of the present invention;

[0060] Figure 2 for Figure 1 A detailed flowchart of step S1;

[0061] Figure 3 for Figure 2 A detailed flowchart of step S15. Detailed Implementation

[0062] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0063] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a method for predicting the function of biological reagents. In the embodiments of this invention, please refer to... Figure 1 The diagram shown is a flowchart illustrating the steps of the biological reagent function prediction method of the present invention. In this example, the biological reagent function prediction method includes the following steps:

[0064] Step S1: Collect biomolecular information from the public database of biological reagents to obtain a biomolecular information set of biological reagents, which includes protein sequence data, molecular structural domain data, and functional domain data of biological reagents.

[0065] In this embodiment of the invention, raw data related to biological reagents is obtained from public databases of biological reagents (such as UniProt, PDB, GeneBank, etc.). These databases contain a wide range of protein and gene sequence information. The screening process includes filtering relevant records by keywords (such as biological reagent name, category, function, etc.) or according to specific screening conditions (such as species, experimental data type, protein expression system, etc.). The data is captured and screened using a programming language (such as Python) combined with a bioinformatics toolkit (such as Biopython). After screening, a raw data set containing all qualified biological reagent-related proteins and their corresponding gene sequences is obtained, including raw information data of biological reagents corresponding to enzymes, antibodies, and microbial types. Protein sequences and gene sequences are extracted from these data. First, the sequence files in FASTA or GenBank format stored in the database are parsed using tools such as Biopython to extract the protein amino acid sequence and the corresponding gene sequence for each biological reagent, thereby obtaining the biological reagent protein sequence data. By utilizing specialized domain prediction software, molecular domain identification is performed on the extracted amino acid and gene sequences. Specifically, tools such as InterProScan, Pfam, and SMART can be used to label and identify domains in protein amino acid sequences. These tools can automatically identify domains in protein sequences based on known domain databases and protein sequence alignment algorithms. Through domain labeling, the amino acid sequence of each biological reagent can be divided into several functional domain fragments, such as secondary and tertiary domain sequence fragments, thus obtaining molecular domain data for the biological reagent. Simultaneously, based on the previously extracted molecular domain data, the molecular functional domains of each biological reagent are further annotated and analyzed. Specifically, databases such as KEGG, Reactome, and Gene Ontology (GO) can be used to perform functional annotation for each domain. Through domain alignment and annotation, the functional information of each biological reagent is identified, including enzyme catalytic activity and molecular binding affinity (obtained from secondary domains), substance transport and degradation (obtained from tertiary domains), and other functional characteristics, thus obtaining molecular functional domain data for the biological reagent. Then, by completing the comprehensive data integration of biological reagents, the specific operations include integrating the protein sequence data, molecular structural domain data and molecular functional domain data obtained from previous analysis. The integrated biological reagent biomolecular information set contains detailed information about each biological reagent, including amino acid and gene sequences, structural domain division, functional domain annotation, etc., and finally obtains the biological reagent biomolecular information set.

[0066] Step S2: Perform sequence alignment analysis on the biological reagent protein sequence data to obtain the biological reagent protein sequence alignment features; perform protein structure feature prediction analysis on the biological reagent molecular domain data to generate biological reagent molecular domain features; extract specific molecular functional features based on the biological reagent molecular functional domain data to obtain specific molecular functional domain features of the biological reagent.

[0067] In this embodiment of the invention, multiple sequence alignment (MSA) is performed on each protein sequence within the biological reagent protein sequence data. Common alignment tools such as BLAST (Basic Local Alignment Search Tool) or Clustal Omega can be used to align the protein sequences. The input protein sequence data is first preprocessed to remove low-quality sequences and standardize the format of all sequences (e.g., FASTA format). During the alignment process, an appropriate alignment score matrix (e.g., BLOSUM or PAM) is set to calculate the score for amino acid substitutions, and appropriate gap penalty and gap extension are set. The alignment tool uses a penalty to handle insertion and deletion regions. Based on the set parameters, it outputs an alignment matrix, which includes the alignment score, match length, and existing insertion and deletion regions (gap) between each pair of protein sequences. It also performs sequence differential statistical analysis on the corresponding protein sequences. The goal of this step is to identify and statistically analyze the differences between different protein sequences, mainly including amino acid residue substitutions, insertions, and deletions. Using the alignment results, the residue differences between each aligned sequence are first extracted. By calculating the changes in amino acids at each alignment position, the frequency of amino acid substitutions is statistically determined. Simultaneously, the insertion and deletion positions in the sequence are analyzed to identify which regions have undergone insertion or deletion, thus obtaining the biological reagent protein sequence alignment characteristics. Simultaneously, by performing predictive analysis on the protein structure characteristics of the previously extracted biological reagent molecular domain data, the analysis of the protein sequence domains is conducted. First, protein structure prediction software (such as Phyre2, I-TASSER) is used to predict the three-dimensional structure of each protein sequence. This process predicts the spatial structure of the protein sequence based on the sequence information, obtaining the conformation and interaction information of the domains. On this basis, the stability of the domains can be further calculated, for example, by assessing the protein's free energy or by simulating the protein folding process to infer stability. Hydrophobic interaction forces refer to the nonpolar interactions between hydrophobic amino acid residues within a protein molecule. Molecular dynamics simulations (such as GROMACS software) can be used to calculate the energy changes of these interactions, thereby predicting the characteristics of the generated biological reagent molecular domains, including molecular domain stability, molecular domain hydrophobic interaction forces, and molecular domain binding energy characteristics.Then, specific functional characteristics of proteins are extracted using molecular functional domain data based on biological reagents. First, the protein sequence is mapped to a known functional domain database using bioinformatics tools (such as InterProScan and Pfam) to determine its functional domains and their corresponding biological functions. For each determined functional domain, the enzyme activity level is predicted using existing experimental data or computational models (such as Enzyme Commission numbers, EC numbers), and the binding affinity of the protein is calculated. In addition, combined with the transport and degradation capacity characteristics data of biological reagents, the activity of proteins in the process of material transport or degradation can be predicted using bioprocess networks (such as the KEGG database) or molecular simulation tools. These specific molecular functional characteristics include enzyme activity, binding affinity, and transport and degradation capacity, ultimately yielding the specific molecular functional domain characteristics of the biological reagents.

[0068] Step S3: Construct a biological reagent function prediction network framework, and input the biological reagent protein sequence alignment features, biological reagent molecular structural domain features, and biological reagent specific molecular functional domain features into the biological reagent function prediction network framework to train and construct a function prediction model to generate a biological reagent function prediction model; obtain the molecular features of the biological reagent to be predicted, and input the molecular features of the biological reagent to be predicted into the biological reagent function prediction model to predict the target function to obtain the target biological reagent function prediction result;

[0069] In this embodiment of the invention, a neural network framework for predicting the function of biological reagents is constructed by combining convolutional neural networks (CNNs), recurrent neural networks (RNNs), and transformer neural networks. Specifically, CNNs are used for local feature extraction, RNNs handle the temporal dependencies of feature sequence data, and transformer neural networks capture global dependencies through their self-attention mechanism. These features are then fused and output through fully connected layers, thus constructing the biological reagent function prediction network framework. By inputting the protein sequence alignment features, molecular structural domain features, and molecular functional domain features of the biological reagents into the constructed function prediction network framework, these features, after being processed by the aforementioned CNN, RNN, and transformer networks, are fed into an ensemble learning model (such as a multilayer perceptron, MLP) for training the function prediction model. During training, the network considers not only the basic molecular features of the biological reagents but also target information such as the activity, structural stability, and specific functions of the biological reagents. These targets serve as supervisory signals, helping the network to perform multi-task learning, thereby improving the model's prediction accuracy and training a biological reagent function prediction model. By acquiring corresponding molecular features for the biological reagent to be predicted, including protein sequence alignment features, molecular structural domain features, and molecular functional domain features, the molecular features of the biological reagent to be predicted are obtained. Then, by inputting the molecular features of the biological reagent to be predicted into a pre-trained biological reagent function prediction model, specifically by inputting the protein sequence alignment features, molecular structural domain features, and molecular functional domain features of the biological reagent to be predicted into the model according to a predetermined format, the model uses the aforementioned deep neural network (a combination of CNN, RNN, and Transformer) combined with a multi-task learning framework to comprehensively predict the function of the biological reagent. The predicted targets typically include molecular activity, structural stability, and specific functions. After being processed by the model, this information is output in the form of probabilities, scoring each function. The model output provides researchers with the functional prediction results of the target biological reagent, ultimately yielding the functional prediction result of the target biological reagent.

[0070] Step S4: Based on the function prediction results of the target biological reagent, perform functional verification and optimization on the corresponding biological reagent to be predicted, so as to generate the function optimization results of the target biological reagent.

[0071] In this embodiment of the invention, the function of a target biological reagent is verified and optimized by utilizing the results of a biological reagent function prediction model. First, the molecular characteristics of the biological reagent to be predicted are input, and the function prediction result output by the model is obtained. This prediction result can be a classification label (such as "enzyme activity" or "receptor binding") or a function score (such as the probability score of a specific function). The prediction result provides a theoretical basis for subsequent experimental verification. Next, based on the prediction result, the biological reagent to be verified is experimentally verified. For example, for a biological reagent predicted as having "enzyme activity," enzyme activity experiments (such as substrate conversion assays, enzyme inhibition assays, etc.) can be used to verify whether it has the expected enzyme activity. For the "receptor binding" prediction result, further verification can be performed through receptor binding experiments (such as ELISA, flow cytometry, etc.). After experimental verification, if a discrepancy is found between the function prediction and the experimental results, the function prediction model can be optimized and adjusted according to the actual situation. For example, feature extraction can be redone, model architecture can be adjusted, or training data can be expanded to improve the model's prediction accuracy. In addition, the experimental results can also be fed back into the database to optimize the relevant datasets of the biological reagent, further improving the model's accuracy and reliability, and finally generating the optimized function result of the target biological reagent.

[0072] Furthermore, as an embodiment of the present invention, reference is made to... Figure 2 As shown, Figure 1 A detailed flowchart of step S1 is shown below. In this embodiment, step S1 includes the following steps:

[0073] Step S11: Perform preliminary screening of raw biological reagent data in the public database of biological reagents to obtain a set of raw biological information of biological reagents;

[0074] In this embodiment of the invention, raw data related to biological reagents are obtained from public databases of biological reagents (such as UniProt, PDB, GeneBank, etc.). These databases contain a wide range of protein and gene sequence information. The screening process includes filtering relevant records by keywords (such as biological reagent name, category, function, etc.) or according to specific screening conditions (such as species, experimental data type, protein expression system, etc.). The data is captured and screened using a programming language (such as Python) combined with a bioinformatics toolkit (such as Biopython). After screening, a raw data set containing all qualified biological reagent-related proteins and their corresponding gene sequences is obtained, including raw information data of biological reagents corresponding to enzymes, antibodies, and microbial types. Finally, a raw biological information set of biological reagents is obtained.

[0075] Step S12: Extract protein sequences from the original biological information set of biological reagents to obtain biological reagent protein sequence data, including the amino acid molecular sequence and gene molecular sequence corresponding to each biological reagent.

[0076] In this embodiment of the invention, protein and gene sequences are extracted from the previously obtained raw biological information set. First, tools such as Biopython are used to parse the FASTA or GenBank format sequence files stored in the database to extract the protein amino acid sequence and the corresponding gene sequence for each biological reagent. During this process, it is necessary to ensure that each protein sequence and its corresponding gene sequence are correctly matched to avoid data confusion. At the same time, protein sequence databases (such as UniProt) can be used for confirmation and correction to ensure that the extracted sequences are complete protein coding sequences. The obtained protein sequence data includes the amino acid molecular sequence and gene molecular sequence corresponding to each biological reagent, and finally, the biological reagent protein sequence data is obtained.

[0077] Step S13: Perform molecular domain identification, labeling, and extraction on the corresponding amino acid molecular sequences and gene molecular sequences within the biological reagent protein sequence data to obtain biological reagent molecular domain data, including secondary domain sequence fragments and tertiary domain sequence fragments corresponding to each biological reagent.

[0078] In this embodiment of the invention, molecular domain identification is performed on the extracted amino acid sequences and gene sequences using specialized domain prediction software. Specifically, tools such as InterProScan, Pfam, and SMART can be used to label and identify domains in protein amino acid sequences. These tools can automatically identify domains in protein sequences based on known domain databases and protein sequence alignment algorithms. By labeling the domains, the amino acid sequence of each biological reagent can be divided into several domain fragments with specific functions. At the same time, the three-dimensional structural information related to the domain, such as secondary and tertiary domain sequence fragments, can also be identified. The molecular domain data generated in this process will provide a basis for subsequent molecular functional annotation, ultimately yielding the molecular domain data of the biological reagent.

[0079] Step S14: Perform molecular functional domain annotation and analysis on the molecular structural domain data of biological reagents to obtain molecular functional domain data of biological reagents, including the enzyme catalytic activity, binding affinity and substance transport and degradation functional domain fragments corresponding to each biological reagent.

[0080] In this embodiment of the invention, based on the previously extracted molecular structural domain data of biological reagents, the molecular functional domains of each biological reagent are further annotated and analyzed. Specifically, databases such as KEGG, Reactome, and Gene Ontology (GO) can be used to perform functional annotation of each structural domain. By comparing and annotating the structural domains, the functional information of each biological reagent is identified, including functional characteristics such as enzyme catalytic activity and molecular binding affinity (obtained from secondary structural domains), and substance transport and degradation (obtained from tertiary structural domains). For example, through GO annotation, the enzyme catalytic functional domain of a certain protein can be identified, and its corresponding biochemical reaction type can be inferred. By combining the KEGG pathway database, the role of the biological reagent in cell signal transduction can be further inferred. The output of this step is the functional domain data of each biological reagent, specifically including detailed information on functional domains such as enzyme catalytic activity, binding affinity, and substance transport and degradation, ultimately obtaining the molecular functional domain data of the biological reagent.

[0081] Step S15: Integrate the biological reagent protein sequence data, biological reagent molecular structural domain data, and biological reagent molecular functional domain data to obtain a biological reagent biomolecular information set.

[0082] In this embodiment of the invention, comprehensive data integration of biological reagents is achieved. Specifically, this involves integrating previously analyzed protein sequence data, molecular structural domain data, and molecular functional domain data. This process can be accomplished by writing data processing scripts and using the Pandas library in Python to merge data by ID. The data for each biological reagent is integrated into a unified database or data framework, which includes the amino acid sequence of the protein, the distribution of structural domains, and the corresponding functional domain information. The integrated biological reagent biomolecular information set contains detailed information for each biological reagent, including amino acid and gene sequences, structural domain divisions, functional domain annotations, etc., ultimately resulting in the biological reagent biomolecular information set.

[0083] Furthermore, as an embodiment of the present invention, reference is made to... Figure 3 As shown, Figure 2 A detailed flowchart of step S13 is shown in this embodiment. Step S13 includes the following steps:

[0084] Step S131: Perform molecular domain identification and prediction on the corresponding amino acid molecular sequences and gene molecular sequences within the biological reagent protein sequence data to obtain the amino acid molecular domain region sequences and gene molecular domain region sequences of the biological reagent.

[0085] In this embodiment of the invention, molecular domains are identified and predicted for amino acid sequences and their corresponding gene sequences in biological reagent protein sequence data. First, the protein sequence data is input into a domain identification tool (such as Pfam, InterPro, SMART, etc.). Using known domain models in the database, and combining them with the characteristics of the amino acid sequences, the corresponding domain regions are identified. For gene sequences, the same method is used to predict gene domains and identify gene domains associated with the protein amino acid sequences. This process uses multiple alignment methods (such as BLAST, HMMER, etc.) to score and filter the alignment results of known domains, obtaining the amino acid molecular domain region sequences and gene molecular domain region sequences of the biological reagent. These identified domain regions provide necessary spatial location information for subsequent analysis, ultimately yielding the amino acid molecular domain region sequences and the gene molecular domain region sequences of the biological reagent.

[0086] Step S132: Obtain the corresponding amino acid molecular structure sites and amino acid molecular interaction forces through the amino acid molecular structure domain region sequence of the biological reagent, and perform sequence structure hydrophobicity evaluation analysis on the corresponding amino acid molecular structure domain region sequence of the biological reagent based on the amino acid molecular structure sites and amino acid molecular interaction forces to obtain the hydrophobicity score of the amino acid molecular structure of the biological reagent.

[0087] In this embodiment of the invention, based on the structural domain sequence of amino acid molecules in biological reagents, bioinformatics methods are used to further analyze the structural sites of amino acid molecules and the interaction forces between amino acid molecules. First, molecular dynamics simulations of the amino acid sequences are performed using structure prediction software (such as AutoDock, Dock, Rosetta, etc.) to analyze the relative positions of each amino acid in the spatial structure and the corresponding interaction sites. Based on this, the hydrophobicity of the amino acid molecules is quantitatively evaluated by calculating a hydrophobicity score. This process can be achieved by calculating the hydrophobicity index of amino acids (such as the Kyte-Doolittle hydrophobicity index) to assess the hydrophobicity trend within the amino acid molecule region. The Kyte-Doolittle score (Kyte & Doolittle, 1982) is a method for evaluating the hydrophilicity or hydrophobicity of amino acid side chains in water. By combining the structural sites of amino acids with the interaction forces, the hydrophobicity score of the amino acid molecule structure in the biological reagent is calculated, i.e. Where n is the total number of amino acid molecular structural sites, x i P represents the spatial position corresponding to the i-th amino acid molecular structural site. iThe interaction force of amino acids corresponding to the i-th amino acid molecular structure site reflects the hydrophobicity characteristics of this region, and finally the hydrophobicity score of the amino acid molecular structure of the biological reagent is obtained.

[0088] Step S133: Perform molecular structure functional charge analysis on the domain sequence of the amino acid molecule of the biological reagent to obtain the molecular structure functional charge characteristics of the amino acid molecule of the biological reagent, including the molecular structure charge distribution, the molecular structure functional charge coupling degree, and the molecular structure charge density of the amino acid molecule.

[0089] In this embodiment of the invention, the charge characteristics of the amino acid molecular domain region of the biological reagent are further explored by performing molecular structure-functional charge analysis on the sequence of the domain region. Molecular simulation software (such as APBS, DelPhi, etc.) is used to calculate the charge distribution of the amino acid molecular region to obtain the charge distribution of each amino acid residue. Furthermore, through charge coupling analysis and combined with molecular dynamics simulation data, the charge interaction strength between each amino acid in the amino acid molecular domain is calculated, that is, the charge coupling degree of the amino acid molecular structure-functional charge. This analysis also includes the evaluation of charge density, that is, analyzing the charge density per unit volume in the region, and then inferring its influence on the overall charge characteristics of the biological reagent. These charge characteristic data help to understand the charge behavior of the biological reagent in solution, and finally obtain the charge characteristics of the amino acid molecular structure of the biological reagent.

[0090] Step S134: Based on the hydrophobicity score and functional charge characteristics of the amino acid molecular structure of the biological reagent, the corresponding domain sequence of the amino acid molecular structure of the biological reagent is measured using the protein secondary structure measurement formula to obtain the secondary structure degree of the amino acid molecular region sequence; based on the secondary structure degree of the amino acid molecular region sequence, the corresponding domain sequence of the amino acid molecular structure of the biological reagent is subjected to secondary structure labeling to obtain the secondary structure domain sequence fragment corresponding to the biological reagent.

[0091] In this embodiment of the invention, a suitable secondary structure metric calculation formula is constructed by combining the total length of the amino acid molecular domain sequence, the hydrophobicity score of the amino acid molecular structure of the biological reagent, the hydrophobicity weighting factor associated with the amino acid, the functional charge distribution of the amino acid molecular structure of the biological reagent, the functional charge coupling degree of the amino acid molecular structure, the charge density of the amino acid, the charge characteristic weighting factor associated with the amino acid, the functional group characteristic parameters of the amino acid molecular structure, the functional group characteristic weighting factor associated with the amino acid, and related parameters. This formula is used to calculate the domain metric of the corresponding amino acid molecular domain sequence of the biological reagent, thereby quantifying the secondary structure degree of the region and obtaining the secondary structure degree of the amino acid molecular region sequence. Simultaneously, by combining these secondary structure metrics, an algorithm is used to perform secondary structure labeling to determine the final secondary structure domain sequence fragment, i.e., the secondary structure metric at the corresponding site is greater than or equal to a threshold. The goal of this operation is to provide the secondary structure information of the amino acid molecule of the biological reagent in space, ultimately obtaining the secondary structure domain sequence fragment corresponding to the biological reagent.

[0092] Step S135: Perform molecular tertiary structure deduction analysis on the molecular structural domain region sequence of the biological reagent gene to obtain the corresponding tertiary structural domain sequence fragment of the biological reagent.

[0093] In this embodiment of the invention, the tertiary structure of the gene molecular domain region sequence of the biological reagent is deduced and analyzed. The three-dimensional structure of the amino acid sequence translated from the gene sequence is predicted using structure prediction software (such as I-TASSER, Phyre2, Rosetta, etc.). This process first performs rapid docking of the amino acid sequence to construct its preliminary tertiary structure model, and then further optimizes its structure through molecular dynamics simulation to adjust the spatial position of amino acid residues and obtain a stable tertiary domain sequence fragment. In this process, simulation technology is used to evaluate the spatial arrangement of amino acids in the three-dimensional structure, molecular interactions, and corresponding binding sites to obtain a reliable tertiary structure sequence of the biological reagent, and finally obtains the tertiary domain sequence fragment corresponding to the biological reagent.

[0094] Furthermore, the specific formula for calculating the protein secondary structure measurement in step S134 is as follows:

[0095]

[0096] In the formula, S s Let L be the secondary structure degree of the amino acid molecule's domain sequence, r be the spatial position parameter of the domain, H(r) be the hydrophobicity score of the amino acid molecule's domain sequence at spatial position r, and α be the degree of secondary structure degree of the amino acid molecule. hHere, C(r) represents the hydrophobicity weighting factor associated with amino acids, C(r) is the functional charge distribution of the amino acid molecular region sequence at spatial position r, O(r) is the functional charge coupling degree of the amino acid molecular region sequence at spatial position r, M(r) is the structural charge density of the amino acid molecular region sequence at spatial position r, and β... c γ is the charge characteristic weighting factor associated with amino acids, F(r) is the functional group characteristic parameter of the amino acid molecular region sequence at spatial position r, and γ is the charge characteristic weighting factor associated with amino acids. f η is the weighting factor for the functional group characteristics associated with amino acids, and η is the correction coefficient for the secondary structure degree of the amino acid molecular region sequence.

[0097] This invention, through the use of a specific mathematical model and verification, derives a formula for calculating the secondary structure of proteins. This formula is used to calculate the domain structure of corresponding biological reagent amino acid molecular domain sequences. By introducing multiple parameters (such as hydrophobicity, functional charge, and functional group properties), it can evaluate the secondary structure characteristics of amino acid molecular regions from multiple dimensions. Specifically: hydrophobicity plays an important role in protein folding and structure formation, especially affecting the aggregation of internal nonpolar amino acid residues. Hydrophobicity has a significant impact on protein stability, folding kinetics, and function. The charge distribution of proteins is crucial to their stability, interactions, and conformational changes, especially in the interaction between proteins and other molecules (such as ligands, proteins, or DNA). The formula, by considering factors such as charge distribution, charge coupling degree, and charge density, helps to more accurately predict the spatial conformation of proteins. The properties of functional groups in proteins (such as amino, carboxyl, and amide groups) also have an important impact on their structural stability and function. This parameter helps to capture changes in the chemical environment of proteins, further enhancing the predictive ability of structure. Secondly, in the formula, α... h β c γ fAs weighting factors for hydrophobicity, charge properties, and functional group properties, respectively, this method can be adjusted according to different protein types or research needs. This flexibility ensures that the method can be customized and optimized according to different research objectives and data backgrounds, providing more accurate and detailed structure predictions. By adjusting these weighting factors, the calculation formula can be optimized to adapt to the properties of different proteins, especially those with special structures or functions. In the formula, 'r' serves as a spatial position parameter for the domains, representing the spatial position of the protein molecule's domains. This positional dependence allows for precise assessment of the protein's folding state in three-dimensional space and the influence of its local environment on secondary structure. Amino acid residues at different positions have different physicochemical properties; therefore, the formula considers the specific spatial position of each residue, making the assessment of secondary structure more detailed. Furthermore, the introduction of a correction coefficient enhances the adaptability and robustness of the calculation formula. This coefficient allows for necessary corrections to the formula to accommodate errors in actual data or experiments, further improving the accuracy and practicality of the predictions. By combining multiple factors such as hydrophobicity, charge properties, and functional groups in the calculation, the formula can provide a comprehensive measure of secondary structure. This comprehensive evaluation method goes beyond a single characteristic, fully considering how proteins fold, stabilize, and perform their functions in their natural environment, thus providing a more realistic measure of secondary structure. In summary, this formula fully considers the degree of secondary structure S of amino acid molecular region sequences. s The total length L of the amino acid molecular domain sequence, the spatial position parameter r of the domain, the hydrophobicity score H(r) of the amino acid molecular region sequence at spatial position r, and the hydrophobicity weighting factor α associated with the amino acid. h The functional charge distribution C(r) of an amino acid molecular region sequence at spatial position r, the functional charge coupling degree O(r) of an amino acid molecular region sequence at spatial position r, the charge density M(r) of an amino acid molecular region sequence at spatial position r, and the weighting factor β of the charge characteristics associated with amino acids. c The functional group characteristic parameter F(r) of the amino acid molecular region sequence at spatial position r, and the functional group characteristic weighting factor γ associated with the amino acid. f The correction coefficient η for the secondary structure degree of an amino acid molecular region sequence is based on the secondary structure degree S of the amino acid molecular region sequence. s The interrelationships between the above parameters constitute a functional relationship. This formula enables the calculation of the structural domain measurement of corresponding biological reagent amino acid molecular structural domain sequences. Furthermore, by introducing a correction coefficient η for the secondary structure degree of amino acid molecular regional sequences, adjustments can be made based on errors that occur during the calculation process, thereby improving the accuracy and applicability of the protein secondary structure measurement formula.

[0098] Furthermore, step S135 includes the following steps:

[0099] Gene molecular dynamics simulation analysis was performed on the molecular structural domain sequences of biological reagent genes to generate the biological reagent gene molecular dynamics simulation process.

[0100] In this embodiment of the invention, the domain region sequences of the biological reagent gene molecule are extracted to obtain the sequence information of the target gene, especially the relevant domain region sequences, through high-throughput sequencing technology or gene chip technology. Then, molecular dynamics simulation software (such as GROMACS, AMBER, or CHARMM) is used to perform dynamic simulation of the gene molecule. Specifically, an appropriate force field (e.g., CHARMM27 or OPLS-AA) is selected, and the gene molecule domain region sequences are mapped into the simulation space. A system model is constructed using these software programs, and appropriate temperature (e.g., 300K), pressure (e.g., 1 atm), and time step (e.g., 2 fs) are set. The simulation process needs to be carried out on a time scale of at least several hundred nanoseconds to ensure that the dynamic process can converge sufficiently. During the simulation, energy minimization techniques are used to remove errors introduced by the structure, and molecular behavior changes are analyzed through molecular dynamics trajectory data to generate a gene molecular dynamics simulation process file, which records the structural changes and energy fluctuations of the molecule during time evolution, and finally generates the biological reagent gene molecular dynamics simulation process.

[0101] Preferably, the gene molecular folding path is deduced from the molecular dynamics simulation process of the biological reagent gene to generate a gene molecular folding deduction path diagram of the biological reagent;

[0102] In this embodiment of the invention, the folding process of a gene molecule from its initial random conformation to its final stable conformation is analyzed based on trajectory data from a biomolecular dynamics simulation. To this end, structural biology tools (such as Rosetta or Folding@Home) are used to perform a detailed analysis of the dynamic simulation trajectory. By progressively analyzing the conformational changes of the molecule at various moments during the simulation, particularly changes in key structural domains, the folding path of the gene molecule is deduced. Typically, the analysis steps include identifying the dynamic processes of secondary structures (such as α-helices and β-sheets) formed during folding, and inferring energy changes during folding based on intermolecular forces (such as hydrogen bonds, van der Waals forces, electrostatic forces, etc.). By analyzing the deduced path, a folding path diagram is drawn, reflecting multiple possible pathways of molecular folding and the energy states of each pathway. This diagram represents the process of a biomolecular reagent gene molecule from its unfolded state to its final folded state, ultimately generating a folding path diagram of the biomolecular reagent gene molecule.

[0103] Preferably, based on the molecular folding deduction path map of biological reagent genes, the corresponding molecular structural domain sequence of biological reagent genes is analyzed for domain folding energy distribution to obtain the molecular structural domain folding energy distribution map of biological reagent genes.

[0104] In this embodiment of the invention, based on the folding path diagram, for each key structural domain region, the energy changes experienced by each region during the folding process are further analyzed using the domain folding energy analysis method. Energy calculation tools (such as FoldX, Rosetta Energy Minimization, etc.) are used to quantitatively analyze the folding energy of the biological reagent gene molecule. By calculating the total energy of the gene molecule in different conformations, including van der Waals energy, electrostatic energy, solvent accessibility, and other factors, the energy differences of different folding states are determined. This data is visualized as an energy distribution map, reflecting the concentrated energy distribution areas and lower-energy folding regions under different folding states. This map helps to further understand which structural domain regions are relatively stable in energy during the gene molecule folding process, and which regions are relatively unstable or bottleneck regions of folding, ultimately obtaining the domain folding energy distribution map of the biological reagent gene molecule.

[0105] Preferably, the corresponding gene molecular folding energy stability score and gene molecular folding energy distribution gradient are obtained by using the folding energy distribution map of the gene molecular structure domain of the biological reagent. Based on the gene molecular folding energy stability score and gene molecular folding energy distribution gradient, the corresponding biological reagent gene molecular structure domain region sequence is matched and calculated using the protein tertiary structure matching calculation formula to obtain the tertiary structure matching degree of the gene molecular region sequence.

[0106] In this embodiment of the invention, a folding stability score for each domain region is calculated based on the folding energy distribution map. The scoring criteria include the location of energy troughs, energy distribution, and the stability of the domain region during folding. The gradient value corresponding to the energy distribution is statistically analyzed to obtain the gene molecule folding energy stability score and the gene molecule folding energy distribution gradient. Simultaneously, a suitable tertiary structure matching calculation formula is constructed by combining the gene molecule folding spatial range, gene molecule folding spatial coordinates, folding energy Hamiltonian, gene molecule folding energy adjustment coefficient, gene molecule folding energy stability score, gene molecule folding energy stability influence parameter, gene molecule folding energy distribution gradient, folding energy distribution gradient influence coefficient, and related parameters. This formula is used to match the corresponding biological reagent gene molecule domain region sequences. By comparing the secondary structure information of the target gene molecule with known protein three-dimensional structure databases (such as PDB), the similarity between the target domain region and the known protein three-dimensional structures in the database is evaluated. This value represents the similarity of the tertiary structure of the gene molecule region sequence, ultimately yielding the tertiary structure matching degree of the gene molecule region sequence.

[0107] The specific formula for calculating protein tertiary structure matching is as follows:

[0108]

[0109] In the formula, M represents the tertiary structure matching degree of the gene molecular region sequence, Ω represents the gene molecular folding space range, and x represents the gene molecular folding space coordinates. Let δ be the Hamiltonian of the folding energy at coordinate x in the gene folding space, and δ be the adjustment coefficient for the gene folding energy. U is the gradient symbol. fold (x) represents the gene folding energy stability score at coordinate x in the gene folding space, and ε is the parameter affecting gene folding energy stability. ξ is the energy distribution gradient of gene molecule folding, θ is the influence coefficient of the energy distribution gradient of folding, and ξ is the correction coefficient of the tertiary structure matching degree of gene molecule region sequence.

[0110] This invention, through the use of a specific mathematical model and verification, yields a protein tertiary structure matching calculation formula. This formula is used to perform matching calculations on corresponding biological reagent gene molecular domain sequences. The formula can calculate and predict the tertiary structure of biological reagent gene molecular domain sequences using a precise mathematical model. This process not only enhances the understanding of gene folding behavior but also provides valuable references for downstream biological experiments. Multiple parameters in the formula (such as folding energy Hamiltonian, folding energy stability score, and folding energy distribution gradient) help to adjust and optimize the folding path and energy distribution in real time during the folding process. Through systematic deduction of folding path and energy stability, folding paths with high stability and reasonable structures can be effectively identified and selected, thus providing a scientific basis for the precise design of gene molecular structures. The folding energy adjustment coefficient, stability score, influencing parameters, and energy distribution gradient influence coefficient in the formula can all be used to quantify the folding stability and efficiency of proteins. By adjusting these coefficients, the folding trend, stability, and efficiency of proteins can be predicted under different folding conditions, helping to select appropriate experimental conditions and reduce experimental complexity. The matching degree value in the formula for calculating the tertiary structure of proteins provides a quantitative indicator for the final structure prediction, representing the similarity between the predicted tertiary structure and the actual structure. This helps assess the relative structural adaptability of gene molecular region sequences in biological reagents under different folding paths, thereby determining the sequence closest to its natural folding state. This formula is not only applicable to routine protein folding prediction but also allows for personalized design and optimization for specific biological reagent gene molecules. For example, when designing protein drugs or biomaterials with specific functions, this formula can accurately predict and optimize their tertiary structure based on features such as domain folding paths and energy distribution, improving their functional performance and stability. By calibrating the tertiary structure matching degree of gene molecular region sequences, the corresponding folding results can be deduced before actual experiments, allowing the selection of the optimal scheme for experimentation. This significantly improves the success rate of biological experiments and avoids numerous repeated experiments and inefficient attempts. In summary, this formula fully considers the tertiary structure matching degree M of gene molecular region sequences, the folding space range Ω of gene molecules, the folding space coordinate x of gene molecules, and the folding energy Hamiltonian at the folding space coordinate x of gene molecules. Gene molecular folding energy adjustment coefficient δ, gradient sign The gene fold energy stability score U at the gene fold space coordinate x. fold (x), parameter ε affecting gene folding energy stability, gradient of gene folding energy distribution. The folding energy distribution gradient influence coefficient θ, the correction coefficient ξ for the tertiary structure matching degree of the gene molecular region sequence, and the interrelationship between the tertiary structure matching degree M of the gene molecular region sequence and the above parameters constitute a functional relationship. This formula can perform matching calculations for the molecular structural domain regions of corresponding biological reagent genes. Furthermore, by introducing a correction coefficient ξ for the tertiary structure matching degree of gene molecular region sequences, adjustments can be made based on errors that occur during the calculation process, thereby improving the accuracy and applicability of the protein tertiary structure matching calculation formula.

[0111] Preferably, the tertiary structure of the corresponding biological reagent gene molecular domain region sequence is determined based on the tertiary structure matching degree of the gene molecular region sequence to obtain the tertiary domain sequence fragment corresponding to the biological reagent.

[0112] In this embodiment of the invention, based on the tertiary structure matching degree of the gene molecular region sequence obtained in the previous step, sequence fragments with matching degrees exceeding a threshold are selected as target regions for tertiary structure labeling. These regions are then visualized using structural biology software (such as PyMOL, Chimera, etc.), and their three-dimensional structures are further adjusted and optimized. Specifically, the target sequence region is compared with existing high-resolution protein three-dimensional structure models, and the spatial structure of the target sequence is adjusted according to the spatial arrangement and interaction forces of the domains. The most probable tertiary structure model is gradually generated. This process involves backtracking optimization of the model (such as energy minimization and constraint adjustment) until a stable structural state is reached, ultimately yielding the tertiary domain sequence fragment corresponding to the biological reagent.

[0113] Furthermore, step S15 includes the following steps:

[0114] Step S151: Perform functional domain segmentation on the corresponding secondary structure domain sequence fragments within the biological reagent molecule structural domain data to obtain the functional domain type segmentation of the secondary structure of the biological reagent molecule, including the enzyme functional domain fragments and the binding functional domain fragments of the biological reagent molecule.

[0115] In this embodiment of the invention, secondary structure domain sequence fragments corresponding to the structural domain data of biological reagent molecules are analyzed to identify the secondary structure domain fragments and classify them into functional domains. This process is based on the amino acid sequence of the biological reagent molecule and uses secondary structure prediction software (such as PSIPRED, JPred, or I-TASSER) to predict the secondary structure. This tool predicts the domains of the amino acid sequence and combines it with known biological functional domain databases (such as Pfam, SMART, and InterPro) to label and confirm the functional domains. Each secondary structure domain is determined based on its amino acid sequence and... The predicted results were assigned to specific functional domain types, including enzyme functional domains and binding functional domains. Enzyme functional domains include fragments with catalytic activity, such as those classified based on certain enzyme catalytic mechanisms (e.g., hydrolysis, redox, transfer, etc.). Binding functional domains are related to the binding of specific molecules or ions, such as receptor binding domains or DNA binding domains. Using the above tools and methods, the secondary structure functional domains of biological reagents were divided into multiple fragments and classified into different functional domain types, ultimately resulting in the segmentation of secondary structure functional domains of biological reagent molecules, including enzyme functional domain fragments and binding functional domain fragments.

[0116] Step S152: Obtain the known catalytic enzyme reaction mechanism and enzyme substrate characteristics, and perform catalytic potential prediction analysis on the enzyme functional domain fragments of biological reagent molecules based on the known catalytic enzyme reaction mechanism and enzyme substrate characteristics to generate a biological reagent molecule catalytic potential score; perform enzyme catalytic activity annotation processing on the corresponding biological reagent molecule enzyme functional domain fragments based on the biological reagent molecule catalytic potential score to obtain the enzyme catalytic activity functional domain fragments corresponding to each biological reagent.

[0117] In this embodiment of the invention, enzyme catalytic potential is predicted and analyzed by combining known catalytic enzyme reaction mechanisms and enzyme substrate characteristic data with enzyme functional domain sequence fragments based on biological reagents. This process involves consulting known enzyme reaction databases (such as BRENDA, KEGG, EC codes, etc.) to obtain the catalytic mechanisms and substrate characteristics of different enzyme reactions. Utilizing the regularity of enzyme catalytic reaction mechanisms and based on the similarity between enzyme substrate characteristics and catalytic mechanisms, machine learning algorithms (such as support vector machines, random forests, etc.) are used for model training and prediction to generate a catalytic potential score for each enzyme functional domain fragment. This score, based on the predicted catalytic activity, reflects the potential of the functional domain in the catalytic reaction. A higher score indicates a stronger catalytic potential for the enzyme functional domain, thereby predicting and generating a catalytic potential score for the biological reagent molecule. Simultaneously, by combining the catalytic potential scores of previously analyzed biological reagent molecules, the corresponding enzyme functional domain fragments of the biological reagent molecules are annotated with enzyme catalytic activity. Based on the previously obtained catalytic potential scores, each enzyme functional domain fragment is assigned a catalytic activity score, which reflects the actual enzyme catalytic ability of the functional domain. On this basis, by analyzing the enzyme's substrate characteristics, reaction mechanism, and experimental data (such as enzymatic parameters: Km, Vmax, etc.), the enzyme catalytic activity is further annotated in detail. This process can be assisted by enzymatic experimental data models (such as the Michaelis-Menten equation) to verify and correct the prediction results, ensuring the accuracy of catalytic activity annotation. Based on the catalytic potential scores and activity annotations, the activity evaluation of each enzyme functional domain fragment is completed, and the catalytic activity level of each enzyme functional domain is labeled, ultimately obtaining the enzyme catalytic activity functional domain fragment corresponding to each biological reagent.

[0118] Step S153: By simulating different environmental constraints, and performing dynamic annotation of binding affinity of the binding functional domain fragments of biological reagent molecules based on different environmental constraints, the binding affinity functional domain fragments corresponding to each biological reagent are obtained.

[0119] In this embodiment of the invention, the binding affinity of binding functional domain fragments of biological reagent molecules is dynamically annotated. First, based on known binding functional domain types (such as receptor binding domains, protein-protein interaction domains, etc.) and molecular structure data of binding ligands, molecular docking software (such as AutoDock, DOCK, GOLD, etc.) is used to simulate the influence of different environmental factors (such as pH, temperature, ion concentration, etc.) on binding affinity. By setting these environmental factors and performing dynamic simulation and energy calculation of the binding process, the predicted binding affinity values ​​of each binding functional domain fragment under different environmental conditions are obtained. The predicted binding affinity values ​​are based on parameters such as the molecular docking score function, binding constant (Kd), and binding free energy (ΔG), thereby providing binding affinity annotation data for each biological reagent molecule, and finally obtaining the binding affinity functional domain fragment corresponding to each biological reagent.

[0120] Step S154: Perform transmembrane transport capability prediction analysis on the corresponding tertiary domain sequence fragments within the molecular structural domain data of biological reagents to obtain the predicted value of transmembrane transport capability of biological reagent molecules.

[0121] In this embodiment of the invention, the transmembrane transport capacity of the corresponding tertiary domain sequence fragments within the structural domain data of biological reagent molecules is predicted and analyzed. This is achieved by using known transmembrane transporter databases (such as TCDB) and transport mechanisms, combined with the amino acid sequence and tertiary structure information of the biological reagent molecule, and using transmembrane protein prediction tools (such as TMHMM, Phobius, TOPCONS, etc.) to predict the transport potential. These tools analyze the transmembrane helical structure and hydrophilicity / hydrophobicity in the amino acid sequence to predict whether the biological reagent molecule possesses transmembrane transport capacity. Based on this, according to the characteristics of membrane proteins and experimental data, a transmembrane transport capacity score is further provided for each tertiary domain fragment, indicating the likelihood that the functional domain fragment will act as a transmembrane transporter protein. Finally, the predicted value of the transmembrane transport capacity of the biological reagent molecule structure is obtained.

[0122] Step S155: Based on the predicted transmembrane transport capacity of the biological reagent molecular structure, the corresponding tertiary domain sequence fragments are annotated with material transport and degradation data to obtain the material transport and degradation functional domain fragments corresponding to each biological reagent.

[0123] In this embodiment of the invention, by combining previously obtained predicted values ​​of transmembrane transport capacity, the tertiary domain fragments of biological reagent molecules are further annotated with material transport and degradation information. By simulating the effects of different environmental conditions (such as temperature, pH, ion concentration, etc.) on the material transport process, dynamic simulation tools (such as GROMACS, AMBER, etc.) are used to predict whether biological reagent molecules will perform material transport and degradation functions on the cell membrane under specific conditions. Combined with existing transport reaction kinetic models (such as Michaelis-Menten kinetics or Hill equations), the transport and degradation characteristics of each tertiary domain fragment are annotated. By predicting the transport capacity and degradation activity of the functional domain under different environmental conditions, each functional domain fragment is assigned an annotation label of material transport and degradation functional domain, and finally, the material transport and degradation functional domain fragment corresponding to each biological reagent is obtained.

[0124] Furthermore, step S2 includes the following steps:

[0125] Step S21: Perform inter-sequence alignment analysis on each protein sequence in the biological reagent protein sequence data to obtain the inter-sequence alignment matrix of the biological reagent protein sequence, which includes the alignment score, matching length and gap region between each protein sequence.

[0126] In this embodiment of the invention, multiple sequence alignment (MSA) is performed on each protein sequence within the biological reagent protein sequence data. Common alignment tools such as BLAST (Basic Local Alignment Search Tool) or Clustal Omega can be used to align the protein sequences. The input protein sequence data is first preprocessed to remove low-quality sequences and standardize the format of all sequences (e.g., FASTA format). During the alignment process, an appropriate alignment score matrix (e.g., BLOSUM or PAM) is set to calculate the score for amino acid substitutions, and appropriate gap penalty and gap extension are set. The alignment tool uses a penalty to handle insertion and deletion regions. Based on the set parameters, the alignment tool outputs an alignment matrix, which includes the alignment score, match length, and existing insertion and deletion regions (gap) between each pair of protein sequences. The alignment score reflects the similarity of the protein sequences, the match length indicates the number of consecutive amino acids that match between the two sequences, and insertions and deletions refer to the unmatched regions in the sequence. By analyzing this information, an alignment matrix between protein sequences can be obtained, and finally, the alignment matrix between biological reagent protein sequences can be obtained.

[0127] Step S22: Based on the alignment matrix between biological reagent protein sequences, perform sequence difference statistical analysis on the corresponding protein sequences within the biological reagent protein sequence data to obtain the alignment features of biological reagent protein sequences, including the differences in amino acid residue substitution, insertion, and deletion between protein sequences.

[0128] In this embodiment of the invention, after completing the protein sequence alignment matrix, sequence differential statistical analysis is performed on the corresponding protein sequences. The goal of this step is to identify and statistically analyze the differences between different protein sequences, mainly including amino acid residue substitutions, insertions, and deletions. Using the alignment results, the residue differences of each aligned sequence are first extracted. By calculating the changes in amino acids at each alignment position, the frequency of amino acid substitutions is statistically determined. At the same time, the insertion and deletion positions in the sequence are analyzed to identify which regions have undergone insertion or deletion. Specifically, the Biopython library in the Python programming language can be used to parse the alignment data and perform differential statistics. Through the alignment matrix, the amino acid changes of each pair of protein sequences are checked one by one, and the type of substitution (such as polar-polar substitution, acid-basic substitution, etc.) is recorded. Based on these difference information, a sequence differential statistical analysis table is established, which includes the substitution type at each amino acid position, the location of the insertion and deletion regions, and the statistical data of substitutions, ultimately obtaining the biological reagent protein sequence alignment characteristics.

[0129] Step S23: Perform protein structure feature prediction analysis on the molecular domain data of biological reagents to generate molecular domain features of biological reagents, including molecular domain stability, molecular domain hydrophobic interaction force and molecular domain binding energy features.

[0130] In this embodiment of the invention, protein structural features are predicted and analyzed by previously extracted biological reagent molecular domain data. This involves analyzing the domains of the protein sequence. First, protein structure prediction software (such as Phyre2 or I-TASSER) is used to predict the three-dimensional structure of each protein sequence. This process predicts the spatial structure of the protein sequence based on sequence information, obtaining the conformation and interaction information of the domains. Based on this, the stability of the domains can be further calculated, for example, by assessing the protein's free energy or by simulating the protein folding process. Hydrophobic interactions refer to the nonpolar interactions between hydrophobic amino acid residues within a protein molecule. Molecular dynamics simulations (such as GROMACS software) can be used to calculate the energy changes of these interactions. By combining the results of molecular dynamics simulations, the binding energy characteristics of the molecular domains are further calculated, predicting the binding mode and strength of the domains with other molecules. These analytical results help to understand the functionality and stability of the domains, ultimately generating the molecular domain features of the biological reagent, including molecular domain stability, molecular domain hydrophobic interactions, and molecular domain binding energy characteristics.

[0131] Step S24: Extract specific molecular functional features based on the functional domain data of biological reagents to obtain specific molecular functional domain features of biological reagents, including enzyme activity levels, binding affinity, and substance transport and degradation capabilities of biological reagents.

[0132] In this embodiment of the invention, specific functional characteristics of proteins are extracted using molecular functional domain data based on biological reagents. First, the protein sequence is mapped to a known functional domain database using bioinformatics tools (such as InterProScan and Pfam) to determine its included functional domains and their corresponding biological functions. For each determined functional domain, the enzyme activity level is predicted using existing experimental data or computational models (such as Enzyme Commission numbers or EC numbers), and the binding affinity of the protein is calculated. The binding affinity can be predicted by molecular docking analysis (using tools such as AutoDock and Dock) to determine the binding strength between the protein and its ligands. Furthermore, by combining the transport and degradation capacity characteristic data of biological reagents, the activity of proteins in the process of material transport or degradation can be predicted using bioprocess networks (such as the KEGG database) or molecular simulation tools. These specific molecular functional characteristics include enzyme activity, binding affinity, and transport and degradation capacity, ultimately yielding the specific molecular functional domain characteristics of the biological reagents.

[0133] Furthermore, step S3 includes the following steps:

[0134] Step S31: Construct a biological reagent function prediction network framework by connecting convolutional neural networks (CNN), recurrent neural networks (RNN), and transformer neural networks;

[0135] In this embodiment of the invention, a neural network framework for predicting the function of biological reagents is constructed by using a combination of convolutional neural networks (CNN), recurrent neural networks (RNN), and Transformer neural networks. Specifically, CNN is used for local feature extraction, RNN is used to process the temporal dependencies of feature sequence data, and Transformer captures global dependencies through its self-attention mechanism. First, the protein sequence features of the biological reagent are input into the CNN model. The CNN extracts local features in the protein sequence, such as the amino acid arrangement and local structural information, through convolution operations. Then, the output of the CNN is used as the input of the RNN to capture long-range dependencies in the protein sequence, such as the interaction between structural and functional domains. The output of the RNN is then input into the Transformer network. The multi-head self-attention mechanism in the Transformer is used to further extract complex global features in the sequence. These features are fused and output through fully connected layers, and finally, the network framework for predicting the function of biological reagents is constructed.

[0136] Step S32: Input the protein sequence alignment features of the biological reagent, the molecular structural domain features of the biological reagent, and the specific molecular functional domain features of the biological reagent into the biological reagent function prediction network framework to train and construct the function prediction model. At the same time, the activity, structural stability, and specific function prediction targets of the biological reagent are considered to generate the biological reagent function prediction model.

[0137] In this embodiment of the invention, the protein sequence alignment features, molecular structural domain features, and molecular functional domain features of the biological reagent are input into a pre-constructed functional prediction network framework. After processing by the aforementioned CNN, RNN, and Transformer networks, these features are fed into an ensemble learning model (such as a multilayer perceptron, MLP) for training the functional prediction model. During training, the network considers not only the basic molecular features of the biological reagent but also target information such as the activity, structural stability, and specific functions of the biological reagent. These targets serve as supervisory signals to help the network perform multi-task learning, thereby improving the model's prediction accuracy. Specifically, the loss function in the network will contain multiple sub-terms, each corresponding to a different target prediction task, such as activity prediction or stability assessment. The loss of these targets is optimized through backpropagation, enabling the model to learn multiple aspects of prediction information simultaneously, thereby generating an accurate functional prediction model and ultimately a biological reagent functional prediction model.

[0138] Step S33: Obtain the molecular features of the biological reagent to be predicted, including the protein sequence alignment features, molecular structural domain features, and molecular functional domain features corresponding to the biological reagent to be predicted.

[0139] In this embodiment of the invention, corresponding molecular features are obtained for the biological reagent to be predicted. These features include protein sequence alignment features, molecular domain features, and molecular functional domain features of the biological reagent to be predicted. First, the protein sequence alignment features can be obtained by comparing the sequence of the molecule to be predicted with the sequence in a known protein database (such as UniProt) to extract conserved regions, sequence similarity, and variation information. Second, the molecular domain features are obtained by predicting the functional domains of the protein sequence (such as the Pfam database). These domains reflect the functional regions of the protein. The molecular functional domain features can be predicted by functional annotation tools (such as InterPro) to provide features directly related to molecular function. These features are combined to form a complete feature set of the biological reagent to be predicted, and finally, the molecular features of the biological reagent to be predicted are obtained.

[0140] Step S34: Input the molecular features of the biological reagent to be predicted into the biological reagent function prediction model to predict the target function and obtain the target biological reagent function prediction result.

[0141] In this embodiment of the invention, the molecular features of the biological reagent to be predicted are input into a previously trained biological reagent function prediction model. Specifically, the protein sequence alignment features, molecular structural domain features, and molecular functional domain features of the biological reagent to be predicted are input into the model in a predetermined format. Based on these input data, the model uses the aforementioned deep neural network (a combination of CNN, RNN, and Transformer) combined with a multi-task learning framework to comprehensively predict the function of the biological reagent. The predicted targets typically include molecular activity, structural stability, and specific functions. After being processed by the model, this information is output in the form of probabilities, and each function is scored. The model output can provide researchers with the functional prediction results of the target biological reagent, ultimately yielding the functional prediction result of the target biological reagent.

[0142] Furthermore, the present invention also provides a biological reagent function prediction system for performing the biological reagent function prediction method described above, the biological reagent function prediction system comprising:

[0143] The biological reagent molecular information acquisition module is used to acquire biomolecular information from public databases of biological reagents to obtain a biological reagent biomolecular information set, which includes biological reagent protein sequence data, biological reagent molecular structural domain data, and biological reagent molecular functional domain data.

[0144] The biological reagent molecular feature analysis module is used to perform sequence alignment analysis on biological reagent protein sequence data to obtain biological reagent protein sequence alignment features; to perform protein structure feature prediction analysis on biological reagent molecular domain data to generate biological reagent molecular domain features; and to extract specific molecular functional features based on biological reagent molecular functional domain data to obtain specific molecular functional domain features of biological reagents.

[0145] The biological reagent target function prediction module is used to construct a biological reagent function prediction network framework. It inputs the biological reagent protein sequence alignment features, biological reagent molecular structural domain features, and biological reagent specific molecular functional domain features into the biological reagent function prediction network framework to train and construct the function prediction model, thereby generating the biological reagent function prediction model. It also acquires the molecular features of the biological reagent to be predicted and inputs these features into the biological reagent function prediction model to predict the target function, thereby obtaining the target biological reagent function prediction result.

[0146] The biological reagent function verification and optimization module is used to perform function verification and optimization on the corresponding biological reagent to be predicted based on the function prediction results of the target biological reagent, so as to generate the function optimization results of the target biological reagent.

[0147] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for predicting the function of a biological reagent, characterized in that, Includes the following steps: Step S1: Collect biomolecular information from a public database of biological reagents to obtain a biomolecular information set for the biological reagents. This biomolecular information set includes protein sequence data, molecular structural domain data, and functional domain data of the biological reagents. Step S1 includes the following steps: Step S11: Perform preliminary screening of raw biological reagent data in the public database of biological reagents to obtain a set of raw biological information of biological reagents; Step S12: Extract protein sequences from the original biological information set of biological reagents to obtain biological reagent protein sequence data, including the amino acid molecular sequence and gene molecular sequence corresponding to each biological reagent. Step S13: Perform molecular domain identification, labeling, and extraction processing on the corresponding amino acid molecular sequences and gene molecular sequences within the biological reagent protein sequence data to obtain biological reagent molecular domain data, including secondary domain sequence fragments and tertiary domain sequence fragments corresponding to each biological reagent; wherein, step S13 includes the following steps: Step S131: Perform molecular domain identification and prediction on the corresponding amino acid molecular sequences and gene molecular sequences within the biological reagent protein sequence data to obtain the amino acid molecular domain region sequences and gene molecular domain region sequences of the biological reagent. Step S132: Obtain the corresponding amino acid molecular structure sites and amino acid molecular interaction forces through the amino acid molecular structure domain region sequence of the biological reagent, and perform sequence structure hydrophobicity evaluation analysis on the corresponding amino acid molecular structure domain region sequence of the biological reagent based on the amino acid molecular structure sites and amino acid molecular interaction forces to obtain the hydrophobicity score of the amino acid molecular structure of the biological reagent. Step S133: Perform molecular structure functional charge analysis on the domain sequence of the amino acid molecule of the biological reagent to obtain the molecular structure functional charge characteristics of the amino acid molecule of the biological reagent, including the molecular structure charge distribution, the molecular structure functional charge coupling degree, and the molecular structure charge density of the amino acid molecule. Step S134: Based on the hydrophobicity score and functional charge characteristics of the amino acid molecular structure of the biological reagent, the corresponding domain sequence of the amino acid molecular structure of the biological reagent is measured using the protein secondary structure measurement formula to obtain the secondary structure degree of the amino acid molecular region sequence; based on the secondary structure degree of the amino acid molecular region sequence, the corresponding domain sequence of the amino acid molecular structure of the biological reagent is subjected to secondary structure labeling to obtain the secondary structure domain sequence fragment corresponding to the biological reagent. Step S135: Perform tertiary structure deduction analysis on the molecular structural domain region sequence of the biological reagent gene to obtain the corresponding tertiary structural domain sequence fragment; wherein, step S135 includes the following steps: Gene molecular dynamics simulation analysis was performed on the molecular structural domain sequences of biological reagent genes to generate the biological reagent gene molecular dynamics simulation process. Gene molecular folding path deduction was performed on the molecular dynamics simulation process of biological reagent genes to generate a gene molecular folding deduction path diagram of biological reagents; Based on the molecular folding deduction path map of biological reagent genes, the corresponding molecular structural domain sequence of biological reagent genes is analyzed for domain folding energy distribution to obtain the molecular structural domain folding energy distribution map of biological reagent genes. The folding energy distribution map of the gene molecular structure domain of the biological reagent is used to obtain the corresponding gene molecular folding energy stability score and gene molecular folding energy distribution gradient. Based on the gene molecular folding energy stability score and gene molecular folding energy distribution gradient, the protein tertiary structure matching calculation formula is used to match the corresponding biological reagent gene molecular structure domain region sequence to obtain the tertiary structure matching degree of the gene molecular region sequence. The specific formula for calculating protein tertiary structure matching is as follows: ; In the formula, This refers to the tertiary structure matching degree of gene molecular region sequences. The spatial range of gene molecule folding. These are the spatial coordinates for gene molecule folding. Coordinates in the gene molecular folding space The Hamiltonian of the folding energy at that point. This is the energy adjustment factor for gene molecule folding. The gradient symbol, Coordinates in the gene molecular folding space The gene molecular folding energy stability score at the location. The parameters affecting the energy stability of gene molecule folding. The gradient of energy distribution during gene molecule folding. The coefficient representing the influence of the energy distribution gradient during folding. This is a correction coefficient for the tertiary structure matching degree of gene molecular region sequences; Based on the tertiary structure matching degree of gene molecular region sequence, the tertiary structure labeling process is performed on the corresponding biological reagent gene molecular structural domain sequence to obtain the tertiary structural domain sequence fragment corresponding to the biological reagent. Step S14: Perform molecular functional domain annotation and analysis on the molecular structural domain data of biological reagents to obtain molecular functional domain data of biological reagents, including the enzyme catalytic activity, binding affinity and substance transport and degradation functional domain fragments corresponding to each biological reagent. Step S15: Integrate the biological reagent protein sequence data, biological reagent molecular structural domain data, and biological reagent molecular functional domain data to obtain a biological reagent biomolecular information set. Step S2: Perform sequence alignment analysis on the biological reagent protein sequence data to obtain the biological reagent protein sequence alignment features; perform protein structure feature prediction analysis on the biological reagent molecular domain data to generate biological reagent molecular domain features; extract specific molecular functional features based on the biological reagent molecular functional domain data to obtain specific molecular functional domain features of the biological reagent. Step S3: Construct a biological reagent function prediction network framework, and input the biological reagent protein sequence alignment features, biological reagent molecular structural domain features, and biological reagent specific molecular functional domain features into the biological reagent function prediction network framework to train and construct a function prediction model to generate a biological reagent function prediction model; obtain the molecular features of the biological reagent to be predicted, and input the molecular features of the biological reagent to be predicted into the biological reagent function prediction model to predict the target function to obtain the target biological reagent function prediction result; Step S4: Based on the function prediction results of the target biological reagent, perform functional verification and optimization on the corresponding biological reagent to be predicted, so as to generate the function optimization results of the target biological reagent.

2. The method for predicting the function of biological reagents according to claim 1, characterized in that, The original biological information set of biological reagents in step S11 includes original biological information data of enzymes, antibodies, and microbial types.

3. The method for predicting the function of biological reagents according to claim 1, characterized in that, The specific formula for calculating the protein secondary structure measurement in step S134 is as follows: ; In the formula, The degree of secondary structure of amino acid molecular region sequences. This represents the total length of the amino acid molecular domain sequence. For the spatial location parameters of the structural domain, The spatial location of amino acid molecular regions Hydrophobicity score of amino acid molecular structure of biological reagents Hydrophobicity weighting factors related to amino acids, The spatial location of amino acid molecular regions Biological reagents, amino acid molecular structure, functional charge distribution, The spatial location of amino acid molecular regions The functional-charge coupling degree of amino acid molecules. The spatial location of amino acid molecular regions The charge density of the amino acid molecule structure As a weighting factor for charge properties related to amino acids, The spatial location of amino acid molecular regions The functional group characteristics of amino acid molecules As a weighting factor for the functional group characteristics related to amino acids, This is a correction factor for the secondary structure degree of the amino acid molecular region sequence.

4. The method for predicting the function of biological reagents according to claim 1, characterized in that, Step S15 includes the following steps: Step S151: Perform functional domain segmentation on the corresponding secondary structure domain sequence fragments within the biological reagent molecule structural domain data to obtain the functional domain type segmentation of the secondary structure of the biological reagent molecule, including the enzyme functional domain fragments and the binding functional domain fragments of the biological reagent molecule. Step S152: Obtain the known catalytic enzyme reaction mechanism and enzyme substrate characteristics, and perform catalytic potential prediction analysis on the enzyme functional domain fragments of biological reagent molecules based on the known catalytic enzyme reaction mechanism and enzyme substrate characteristics to generate a biological reagent molecule catalytic potential score; perform enzyme catalytic activity annotation processing on the corresponding biological reagent molecule enzyme functional domain fragments based on the biological reagent molecule catalytic potential score to obtain the enzyme catalytic activity functional domain fragments corresponding to each biological reagent. Step S153: By simulating different environmental constraints, and performing dynamic annotation of binding affinity of the binding functional domain fragments of biological reagent molecules based on different environmental constraints, the binding affinity functional domain fragments corresponding to each biological reagent are obtained. Step S154: Perform transmembrane transport capability prediction analysis on the corresponding tertiary domain sequence fragments within the molecular structural domain data of biological reagents to obtain the predicted value of transmembrane transport capability of biological reagent molecules. Step S155: Based on the predicted transmembrane transport capacity of the biological reagent molecular structure, the corresponding tertiary domain sequence fragments are annotated with material transport and degradation data to obtain the material transport and degradation functional domain fragments corresponding to each biological reagent.

5. The method for predicting the function of biological reagents according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Perform inter-sequence alignment analysis on each protein sequence in the biological reagent protein sequence data to obtain the inter-sequence alignment matrix of the biological reagent protein sequence, which includes the alignment score, matching length and gap region between each protein sequence. Step S22: Based on the alignment matrix between biological reagent protein sequences, perform sequence difference statistical analysis on the corresponding protein sequences within the biological reagent protein sequence data to obtain the alignment features of biological reagent protein sequences, including the differences in amino acid residue substitution, insertion, and deletion between protein sequences. Step S23: Perform protein structure feature prediction analysis on the molecular domain data of biological reagents to generate molecular domain features of biological reagents, including molecular domain stability, molecular domain hydrophobic interaction force and molecular domain binding energy features. Step S24: Extract specific molecular functional features based on the functional domain data of biological reagents to obtain specific molecular functional domain features of biological reagents, including enzyme activity levels, binding affinity, and substance transport and degradation capabilities of biological reagents.

6. The method for predicting the function of biological reagents according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Construct a biological reagent function prediction network framework by connecting convolutional neural networks (CNN), recurrent neural networks (RNN), and transformer neural networks; Step S32: Input the protein sequence alignment features of the biological reagent, the molecular structural domain features of the biological reagent, and the specific molecular functional domain features of the biological reagent into the biological reagent function prediction network framework to train and construct the function prediction model. At the same time, the activity, structural stability, and specific function prediction targets of the biological reagent are considered to generate the biological reagent function prediction model. Step S33: Obtain the molecular features of the biological reagent to be predicted, including the protein sequence alignment features, molecular structural domain features, and molecular functional domain features corresponding to the biological reagent to be predicted. Step S34: Input the molecular features of the biological reagent to be predicted into the biological reagent function prediction model to predict the target function and obtain the target biological reagent function prediction result.

7. A biological reagent function prediction system, characterized in that, For performing the biological reagent function prediction method as described in claim 1, the biological reagent function prediction system comprises: The biological reagent molecular information acquisition module is used to acquire biomolecular information from public databases of biological reagents to obtain a biological reagent biomolecular information set, which includes biological reagent protein sequence data, biological reagent molecular structural domain data, and biological reagent molecular functional domain data. The biological reagent molecular feature analysis module is used to perform sequence alignment analysis on biological reagent protein sequence data to obtain biological reagent protein sequence alignment features; to perform protein structure feature prediction analysis on biological reagent molecular domain data to generate biological reagent molecular domain features; and to extract specific molecular functional features based on biological reagent molecular functional domain data to obtain specific molecular functional domain features of biological reagents. The biological reagent target function prediction module is used to construct a biological reagent function prediction network framework. It inputs the biological reagent protein sequence alignment features, biological reagent molecular structural domain features, and biological reagent specific molecular functional domain features into the biological reagent function prediction network framework to train and construct the function prediction model, thereby generating the biological reagent function prediction model. It also acquires the molecular features of the biological reagent to be predicted and inputs these features into the biological reagent function prediction model to predict the target function, thereby obtaining the target biological reagent function prediction result. The biological reagent function verification and optimization module is used to perform function verification and optimization on the corresponding biological reagent to be predicted based on the function prediction results of the target biological reagent, so as to generate the function optimization results of the target biological reagent.

Citation Information

Patent Citations

  • Multifunctional bioactive peptide prediction method and system

    CN116189798A

  • Artificial intelligence platform for protein engineering

    US20190259470A1

  • Predicting biological functions of proteins using dilated convolutional neural networks

    US20220172055A1

  • A method for identifying novel gene and the resulting novel genes

    WO2008000186A1