A method for predicting olfactory receptor function based on transfer learning and molecular docking

By combining two-step transfer learning with molecular docking and deep learning, the problem of rapid and accurate prediction of large-scale olfactory receptor function was solved. The constructed model has good applicability and predictive effect among different species.

CN119400236BActive Publication Date: 2025-09-09NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411350596.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-09-09
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

Existing technologies make it difficult to quickly and accurately predict the functions of large-scale olfactory receptors, especially due to the complexity and cost limitations of experimental methods, and molecular docking technology lacks prediction accuracy for olfactory receptors with high sequence variability.

Method used

A two-step transfer learning method combining molecular docking and deep learning was adopted. First, large-scale molecular docking data was used for pre-training, then fine-tuning was performed on a small-scale experimental dataset, and finally further adjustment was made based on the experimental data of the target species to construct an olfactory receptor function prediction model.

Benefits of technology

It achieves fast and accurate prediction of olfactory receptor function, improves the generalization ability and prediction accuracy of the model, and can obtain better results especially when data is scarce.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119400236B_ABST
    Figure CN119400236B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting olfactory receptor function based on transfer learning and molecular docking, comprising: obtaining a source domain dataset and a target domain dataset; using a feature extraction module to build a function prediction model by merging feature vectors and inputting them into a fully connected neural network; pre-training the function prediction model using the source domain dataset, and fine-tuning the function prediction model using the target domain dataset to obtain a first-step transfer learning model; determining whether the target species for which the olfactory receptor function is to be predicted is consistent with the species corresponding to the target domain dataset; if so, using the first-step transfer learning model to predict the olfactory receptor function; if not, performing one or more fine-tuning steps on the first-step transfer learning model until a second-step transfer learning model for predicting the olfactory receptor function is obtained. The method of the present invention effectively combines experimental data, molecular simulation, and deep learning technology to achieve rapid and accurate prediction of large-scale olfactory receptor function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and artificial intelligence technology, and in particular to a method for predicting olfactory receptor function based on transfer learning and molecular docking. Background Art

[0002] Olfactory receptors (ORs) are a class of proteins key to olfactory perception in organisms. They specifically recognize and bind to various volatile organic compounds (VOCs). Characterizing the function of ORs is crucial for understanding olfactory mechanisms, environmental adaptability, and social behavior. However, due to the high sequence variability of ORs and the complexity and time-consuming nature of experimental methods, only a few ORs have been clearly characterized, while the functions of most ORs remain unknown.

[0003] Traditional methods for predicting OR function rely primarily on experimental approaches, primarily involving constructing in vivo or in vitro protein expression systems for physiological and biochemical testing to determine the binding strength between the OR and VOC, i.e., the OR's function. While these experimental methods can provide accurate functional information, they are limited by experimental conditions, cost, and time, making them difficult to apply to large-scale OR function screening.

[0004] With the development of computational biology, structure-based protein function prediction methods have gradually become a research hotspot. Molecular docking technology, in particular, predicts the binding mode between proteins and ligands by simulating the spatial structural complementarity and energy minimization between them. Molecular docking technology plays an important role in the virtual screening process of drug discovery and has also been used to predict the function of ORs to improve the efficiency and success rate of experiments. However, molecular docking technology also has certain limitations. For example, the calculation results may contain false positives, and for ORs with high sequence variability, a single molecular docking method cannot meet the required prediction accuracy. In addition, due to the limitations of experimental data, it is challenging to construct a comprehensive and accurate OR function prediction model.

[0005] Deep learning (DL), a key branch of artificial intelligence, has shown great potential in addressing complex bioinformatics problems. DL techniques can improve the accuracy of protein function prediction by learning data representations at multiple levels of abstraction. However, training DL models requires a large amount of annotated data, while the current availability of experimental data for training is relatively limited, which limits the application of DL techniques in OR function prediction.

[0006] To overcome the data shortage, transfer learning has been proposed and applied to OR function prediction. Transfer learning allows a pre-trained model to be fine-tuned on a new task, addressing the data scarcity issue. By transferring the knowledge of a pre-trained model to a new dataset, the model's predictive performance can be significantly improved, achieving superior predictions even with limited data.

[0007] Although some methods have been used to predict OR function, how to effectively combine experimental data, molecular simulation and deep learning technology to achieve rapid and accurate prediction of large-scale ORs functions remains an urgent problem to be solved. Summary of the Invention

[0008] To address these challenges, the present invention proposes a method for predicting olfactory receptor function based on transfer learning and molecular docking. This method combines the advantages of abundant molecular docking data with limited experimental data. The method involves a two-step transfer learning process: first, a pre-trained model based on molecular docking data is trained and fine-tuned on a small-scale experimental dataset; then, the optimized model from one species is transferred to another species with limited experimental data. Compared with existing methods, this method demonstrates superior prediction accuracy and generalization.

[0009] The present invention provides a method for predicting olfactory receptor function based on transfer learning and molecular docking, which is characterized by comprising:

[0010] Step 1: Obtain source domain dataset and target domain dataset;

[0011] Step 2: Use the feature extraction module to build a function prediction model by merging the feature vectors and inputting them into a fully connected neural network;

[0012] Step 3: Pre-training the function prediction model using the source domain dataset, and fine-tuning the function prediction model using the target domain dataset so that it can predict the experimental results of functional identification, thereby obtaining the first step transfer learning model;

[0013] Step 4: Determine whether the target species for which the olfactory receptor function is to be predicted is consistent with the species corresponding to the target domain dataset. If they are consistent, use the first-step transfer learning model to predict the olfactory receptor function; if they are inconsistent, use the olfactory receptor function identification experimental data of the target species on the basis of the first-step transfer learning model to perform one or more steps of fine-tuning until a second-step transfer learning model that can be used to predict the olfactory receptor function of the target species is obtained, and then use the second-step transfer learning model to predict the olfactory receptor function.

[0014] Furthermore, large-scale molecular docking data is used as the source domain dataset.

[0015] Furthermore, the target domain dataset comes from existing batch function identification data.

[0016] Furthermore, a protein and small molecule language model based on a natural language processing method is used as the feature extraction module.

[0017] Furthermore, in step 2, the amino acid sequence of the olfactory receptor and the simplified molecular linear input specification of the small molecule are used as model inputs to obtain feature vector representations of the protein and the small molecule. The two feature vectors are then merged and input into a fully connected neural network, and the binding score between the protein and the small molecule is used as a label to build a functional prediction model.

[0018] Furthermore, a 5-layer deep fully connected neural network model is used to learn and fit the feature vector.

[0019] Furthermore, in the step three, the affinity between the protein and the small molecule is used as a label for the pre-training.

[0020] Furthermore, in step 4, the species corresponding to the experimental data has a close genetic relationship with the target species.

[0021] Preferably, for target species that do not yet have experimental data for olfactory receptor function identification, the experimental data sets of multiple different species are merged into a mixed data set and then the first step of transfer learning is performed, and the model obtained by the first step of transfer learning is used to predict the olfactory receptor function of the target species.

[0022] The olfactory receptor function prediction method based on transfer learning and molecular docking of the present invention effectively combines experimental data, molecular simulation and deep learning technology, and can achieve rapid and accurate prediction of large-scale olfactory receptor function. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Demonstrate the basic principle of multi-step transfer learning.

[0024] Figure 2 Shows the basic process of multi-step transfer learning.

[0025] Figure 3 The deep learning model architecture used by the method of the present invention is shown.

[0026] Figure 4 A flow chart showing the method of the present invention for predicting olfactory receptor function. DETAILED DESCRIPTION

[0027] In order to make those skilled in the art better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and the present invention is not limited thereto.

[0028] Figure 1 Show the basic principles of transfer learning, Figure 2 The basic process of multi-step transfer learning is shown. The function prediction model is pre-trained on a large molecular docking dataset to obtain the overall distribution trend of the interaction between OR proteins and VOC molecules. Further training is then performed on a larger experimental dataset to obtain a more accurate function prediction model. For species with insufficient experimental data, the existing function prediction model can be further optimized using the small amount of experimental data available for that species, thereby achieving OR function prediction for species with insufficient available experimental data.

[0029] The method of the present invention can be used to rapidly predict the function of olfactory receptors. Given sufficient functional identification experimental data, its prediction accuracy is superior to that of traditional molecular docking methods. Taking insect olfactory receptor function prediction as an example, the method is described for predicting olfactory receptor function. It includes the following steps:

[0030] (1) Obtain source domain datasets and target domain datasets. Prepare large-scale molecular docking data as the source domain dataset. The Insect Olfactory Receptor Structure and Function Prediction Database (iORbase, https: / / www.iorbase.com / ) provides large-scale molecular docking data between insect olfactory receptors and gas molecules, including molecular docking results between nearly 6,000 insect olfactory receptors from 59 insect species and nearly 2,000 gas molecules. It covers most types of insect olfactory receptors and compound types, and can better reflect the general laws of the interaction between insect olfactory receptors and gas molecules. Therefore, we choose to use the data in this database as the pre-training dataset. The target domain dataset comes from existing batch functional identification data, mainly collected from published literature.

[0031] For species without existing molecular docking datasets, molecular docking can be performed using various existing molecular docking software. Here, using VINA as an example, we briefly describe the steps involved. First, the structure of the olfactory receptor is predicted based on its amino acid sequence. Open-source protein structure prediction methods, such as alphafold2, can be used. Second, the three-dimensional structures of the protein and small molecule are used to calculate the binding free energy, which is the result of the molecular docking.

[0032] (2) Build a function prediction model. The function prediction model architecture used in this method is as follows Figure 3 As shown, the protein and small molecule language model based on natural language processing methods is first used as a feature extraction module, and the amino acid sequence of the olfactory receptor and the simplified molecular linear input specification (SMILES) of the small molecule are used as the input of the model to obtain the feature vector representation of the protein and small molecule. The two parts of the feature vector are then merged and input into a fully connected neural network, and the binding score between the protein and the small molecule is used as a label to train the model.

[0033] For feature extraction, we used established protein language models and small molecule language models. These models have been pre-trained on large-scale protein and small molecule datasets. They can extract features from protein and small molecule SMILES-formatted sequences, ultimately converting them into abstract feature vector representations. The language models used for feature extraction are from the open-source website hugging face (https: / / huggingface.co / ). Using these open-source language models, we can easily convert input sequences into array-based feature vectors for subsequent fully connected model training. The protein feature extraction module converts amino acid sequences of 300-400 amino acids into feature vectors of 1024 length, while the small molecule feature extraction module converts small molecule SMILES sequences of varying lengths into feature vectors of 768 length. These two vectors are concatenated end-to-end to produce the feature representation vectors for both the protein and small molecule. These concatenated feature vectors are then fed into the fully connected network model for subsequent binding score prediction.

[0034] We use a 5-layer deep fully connected neural network model to learn and fit the concatenated feature vectors. The input layer of the fully connected neural network is consistent with the length of the concatenated feature vector, which is 1024+768=1792. After the input layer, 3 layers of fully connected layers with 512 nodes are stacked, and the final output layer has a length of 1, which is used to output a certain value, that is, the binding score given by the model. In the fully connected model, we add a batch normalization layer after each fully connected layer to scale the data to the same range and eliminate the dimensional differences between the features, thereby increasing the convergence speed of the model. The feature vector passes through the neural network model and finally obtains a one-dimensional output score, which is the direct interaction score between the corresponding protein and the small molecule. The mean square error (MSE loss) is used as the loss function of the regression task:

[0035]

[0036] Where n is the number of samples, is the prediction result, y iThe difference between the label and the score output by the model is calculated, and the neural network weights are updated using backpropagation.

[0037] (3) Function prediction model pre-training and the first step of transfer learning. The first step of transfer learning is to migrate from the molecular docking dataset to the functional identification experimental dataset. First, the function prediction model is pre-trained using large-scale molecular docking data (source domain dataset). During training, the affinity (affinity score) between the protein and the small molecule is used as the label for model training, thereby training a model that predicts molecular docking. Afterwards, the model trained with the molecular docking data is trained again using the functional identification experimental data, that is, the model is further fine-tuned so that it can predict the experimental results of functional identification. The source domain dataset of the first step of transfer learning is the molecular docking dataset, and its source has been explained in step (1). Then, a certain amount of functional identification experimental data (target domain dataset) is used to fine-tune the model. The functional identification datasets involved are shown in Table 1.

[0038] Table 1 Functional identification experimental dataset information

[0039]

[0040]

[0041] If the target species for olfactory receptor function prediction matches the species corresponding to the experimental dataset, function prediction can be performed directly. However, since olfactory receptor sequences and functions can vary significantly between species, the model from the first transfer learning step is not suitable for directly predicting OR function across species. Therefore, a second transfer learning step, namely, transfer between species, is necessary.

[0042] (4) Second step transfer learning. If the target species is inconsistent with the species corresponding to the experimental dataset, the second step transfer learning can be used to further fine-tune the model. Based on the first step transfer learning model, a small amount of OR function identification experimental data of the target species is used for retraining, that is, the first step transfer learning model is fine-tuned until the olfactory receptor function prediction model of the target species is obtained. The first step transfer learning is the migration from molecular docking results (source domain) to experimental results (target domain), while the second step transfer learning is the migration from the large-scale experimental results of a certain species (source domain) to the individual experimental results of the target species (target domain). Generally speaking, the species corresponding to the experimental dataset and the target species should have a close relationship. For insects, transfer learning can be effectively carried out between two insect species belonging to the same order. If the relationship between the two species is far and the difference is large, resulting in a large difference between the source domain and the target domain of the second transfer learning, the reliability of the prediction results may be reduced.

[0043] Olfactory receptor function prediction in other situations. For target species without experimental data, the model generated using the first-step transfer learning can be directly used to predict olfactory receptor function for the target species. Comparisons have shown that combining experimental datasets from multiple species into a single mixed dataset for first-step transfer learning yields higher accuracy in predicting olfactory receptor function for target species without experimental data than using a single dataset for first-step transfer learning.

[0044] In summary, the applicable scenarios and corresponding operations of the olfactory receptor function prediction method proposed in the present invention are as follows: (1) There are a large number of functional identification results for the target species. The first step of transfer learning is performed using the existing experimental data, and the first step of transfer learning model is used to predict the olfactory receptor function of the target species; (2) There are a small number of functional identification results for the target species. The second step of transfer learning can be performed on the basis of the first step of transfer learning model, and the second step of transfer learning model is used to predict the olfactory receptor function of the target species. (3) For target species with no available experimental data, the first step of transfer learning is performed using a mixed data set obtained by merging multiple experimental data sets, and the corresponding model is used to predict the olfactory receptor function of the target species. The process of predicting the olfactory receptor function using the method of the present invention is as follows: Figure 4 shown.

[0045] The above are merely preferred embodiments of the present invention. It should be noted that the above preferred embodiments should not be construed as limiting the present invention, and the scope of protection of the present invention should be determined by the scope defined in the claims. Persons skilled in the art will appreciate that improvements and modifications may be made without departing from the spirit and scope of the present invention, and these improvements and modifications should be considered to fall within the scope of protection of the present invention.

Claims

1. A method for predicting olfactory receptor function based on transfer learning and molecular docking, characterized in that: include: Step 1: Obtain source domain dataset and target domain dataset; Step 2: Use the feature extraction module to build a function prediction model by merging the feature vectors and inputting them into a fully connected neural network; In step 2, the amino acid sequence of the olfactory receptor and the simplified molecular linear input specification of the small molecule are used as model inputs to obtain feature vector representations of the protein and small molecule. The two feature vectors are then combined and input into a fully connected neural network, and the binding score between the protein and the small molecule is used as a label to build a function prediction model. Step 3: Pre-training the function prediction model using the source domain dataset, and fine-tuning the function prediction model using the target domain dataset so that it can predict the experimental results of functional identification, thereby obtaining the first step transfer learning model; Step 4: Determine whether the target species for predicting the olfactory receptor function is consistent with the species corresponding to the target domain dataset. If they are consistent, use the first step transfer learning model to predict the olfactory receptor function; If there is any inconsistency, that is, based on the first-step transfer learning model, the experimental data of the olfactory receptor function identification of the target species is used to perform one or more steps of fine-tuning until a second-step transfer learning model that can be used to predict the olfactory receptor function of the target species is obtained, and then the second-step transfer learning model is used to predict the olfactory receptor function.

2. The method for predicting olfactory receptor function according to claim 1, wherein Large-scale molecular docking data were used as the source domain dataset.

3. The method for predicting olfactory receptor function according to claim 1, wherein The target domain dataset comes from existing batch function identification data.

4. The method for predicting olfactory receptor function according to claim 1, wherein A protein and small molecule language model based on a natural language processing method is used as the feature extraction module.

5. The method for predicting olfactory receptor function according to claim 1, wherein A 5-layer deep fully connected neural network model is used to learn and fit the feature vector.

6. The method for predicting olfactory receptor function according to claim 1, wherein In the step three, the affinity between the protein and the small molecule is used as a label for the pre-training.

7. The method for predicting olfactory receptor function according to claim 1, wherein In the step 4, the species corresponding to the experimental data has a close genetic relationship with the target species.

8. The method for predicting olfactory receptor function according to claim 1, wherein For target species that do not have experimental data for olfactory receptor function identification, the experimental datasets of multiple different species are merged into a mixed dataset and then the first step of transfer learning is performed. The model obtained by the first step of transfer learning is used to predict the olfactory receptor function of the target species.

Citation Information

Patent Citations

  • Drug molecule skeleton replacing and screening method based on deep transfer learning model

    CN115881244A

  • Machine learning model for sensory property prediction

    CN116670772A