Method for constructing a GPCR thermal stability mutation prediction model, prediction method and device
By constructing the thermal stability mutation prediction model of GPCR, using the pre-trained model for evolutionary fine-tuning, and combining the mutation thermal stability sample data with limited labeled data for supervision and training, the problem of time-consuming and labor-intensive and high R&D expenses of GPCR thermal stability site mutation experiments in the existing technology is solved, and efficient and accurate thermal stability mutation site prediction is achieved, reducing R&D costs.
Patent Information
- Application Number
- CN202210545127.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-05-19
AI Technical Summary
The prior art has problems such as time-consuming and labor-intensive and high R&D costs in the thermal stability site mutation experiment of GPCR, and it is difficult to efficiently and accurately predict thermal stability mutation sites.
By constructing the thermal stability mutation prediction model of GPCR, using the pre-trained model for evolutionary fine-tuning, and combining the mutation thermal stability sample data with limited labeled data for supervision and training, a prediction model that can efficiently and accurately predict thermal stability mutation sites was obtained.
It realizes efficient and accurate prediction of thermally stable mutation sites in GPCR, reducing experimental operations, improving R&D efficiency, and reducing R&D costs.
Smart Images

Figure CN114913914B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of G protein coupled receptor mutation, and in particular to a method for constructing a GPCR thermal stability mutation prediction model, a prediction method and a device. Background Art
[0002] The G protein-coupled receptor (GPCRs) family is one of the most important receptor families encoded by the human genome. It is widely expressed in the main organ and tissue systems of the human body. Its three-dimensional structure has seven transmembrane α helices, and there are binding sites at the C-terminus of its peptide chain and on the intracellular loop connecting the 5th and 6th transmembrane helices. Their main function is to transmit extracellular information to the cell through interaction with G proteins. GPCRs are involved in a large number of human diseases, physiological and pharmacological activities, and are currently the most studied drug targets. Studying the structural stability of GPCRs and their structural analysis technology is crucial for the development of new drugs. However, due to their relatively large structural flexibility, variable conformations, and folding errors in heterologous expression, it is still difficult to analyze the structure of some GPCRs. The prediction of thermal stability mutation sites of GPCRs is becoming increasingly important.
[0003] In the related art, two methods for modifying the thermal stability of GPCR mutants are generally used. One modification method is the systematic ALA scanning GPCR mutation method, which generally mutates the 7 TM region amino acids of the GPCR to ALA one by one (if the original amino acid in the TM region is ALA, the original amino acid is mutated to Leu), and then the thermal stability is tested experimentally. Another modification method is the directed protein evolution method, which generally designs a mutation library, then applies screening pressure, and finally obtains a mutant with optimized thermal stability. However, these methods can only explore part of the mutation space at present, which is time-consuming and labor-intensive and requires a large amount of R&D expenses.
[0004] Therefore, how to shorten the experimental trial process of GPCR mutation at the thermal stability site in order to save time and effort and reduce R&D costs is a problem that needs to be solved at present. Summary of the invention
[0005] In order to solve or partially solve the problems existing in the related art, the present application provides a method for constructing a GPCR thermal stability mutation prediction model, a GPCR thermal stability mutation prediction method and a device thereof, which can construct a prediction model to more efficiently and accurately predict which site mutations of the GPCR have the possibility of contributing to stability, thereby saving time and effort and reducing R&D costs.
[0006] The first aspect of the present application provides a method for constructing a GPCR thermal stability mutation prediction model, which comprises:
[0007] Get the pre-trained model;
[0008] Obtaining a homologous sequence of a sample GPCR, and performing self-supervised training on the pre-trained model according to the homologous sequence to obtain an evolutionary fine-tuning model;
[0009] Obtain mutation thermal stability sample data of the sample GPCR, and perform supervised training on the evolutionary fine-tuning model based on the mutation thermal stability sample data to obtain a prediction model.
[0010] In one embodiment, obtaining a homologous sequence of a sample GPCR comprises:
[0011] A sequence comparison is performed in a preset database according to the protein sequence of the sample GPCR to obtain a homologous sequence of the sample GPCR.
[0012] In one embodiment, the protein sequence of the sample GPCR is compared in a preset database, comprising:
[0013] According to the protein sequence of the sample GPCR, a sequence alignment is performed in the first protein database to obtain the first set of multiple sequence alignment results; wherein:
[0014] When the number of valid sequences in the first group of multiple sequence alignment results is greater than or equal to a preset number threshold, outputting the valid sequence as a homologous sequence of the sample GPCR;
[0015] When the number of valid sequences in the first group of multiple sequence alignment results is less than a preset number threshold, sequence alignment is performed in the Pth protein database according to the protein sequence of the sample GPCR, and the alignment result is combined with the previous P-1 group of multiple sequence alignment results to obtain the Pth group of multiple sequence alignment results, until the number of valid sequences in the Pth group of multiple sequence alignment results is greater than or equal to the preset number threshold, the sequence alignment is terminated, and the valid sequence is output as the homologous sequence of the sample GPCR, where P is a natural number greater than 1.
[0016] In one embodiment, the step of obtaining the mutation thermal stability sample data of the sample GPCR comprises:
[0017] Obtain the activity change of the sample GPCR after alanine scanning mutation before and after the first heating and the activity change of the corresponding wild-type sample GPCR before and after the second heating respectively;
[0018] The first activity change before and after heating is compared with a preset multiple threshold of the second activity change before and after heating to determine the positive sample data and the negative sample data corresponding to the sample GPCR.
[0019] In one embodiment, comparing the first activity change before and after heating with a preset multiple threshold of the second activity change before and after heating to determine the positive sample data and the negative sample data corresponding to the sample GPCR includes:
[0020] Comparing the first activity change before and after heating with K preset multiple thresholds of the second activity change before and after heating, respectively, to obtain training data consisting of positive sample data and negative sample data determined at each preset multiple threshold, wherein K is a positive integer greater than or equal to 2;
[0021] The step of performing supervised training on the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model comprises:
[0022] The evolutionary fine-tuning model is supervisedly trained using the training data under each preset multiple threshold to obtain K candidate prediction models.
[0023] In one embodiment, the method further comprises:
[0024] Use the preset evaluation index to score the K candidate prediction models respectively, and obtain the scoring result of each candidate prediction model;
[0025] The candidate prediction model whose scoring results meet the preset conditions is selected as the prediction model.
[0026] In one embodiment, the supervised training of the evolutionary fine-tuning model based on the mutation thermal stability sample data to obtain a prediction model includes:
[0027] Setting the evolutionary fine-tuning model according to M groups of random parameters to obtain M corresponding evolutionary fine-tuning sub-models; wherein M is a positive integer greater than or equal to 2;
[0028] The mutant thermal stability sample data of the sample GPCR is used as training data to train the M evolutionary fine-tuning sub-models respectively, to obtain the corresponding M trained fine-tuning sub-models, and the prediction model is constructed using the M fine-tuning sub-models.
[0029] In one embodiment, if the sample GPCRs are N, N is a positive integer greater than or equal to 2;
[0030] The method uses the mutant thermal stability sample data of the sample GPCR as training data to train M evolutionary fine-tuning sub-models respectively, obtains corresponding M trained fine-tuning sub-models, and constructs a prediction model using the M fine-tuning sub-models, including:
[0031] The mutant thermal stability sample data of N sample GPCRs are divided according to the N-fold cross-validation method to obtain N groups of training data, wherein each group of training data includes a training set consisting of the mutant thermal stability sample data of (N-1) sample GPCRs and a validation set consisting of the mutant thermal stability sample data of the remaining 1 sample GPCR;
[0032] Using the same set of training data to perform supervised training on M evolutionary fine-tuning sub-models respectively, to obtain corresponding M trained fine-tuning sub-models, and using the M fine-tuning sub-models to construct a corresponding prediction sub-model, to obtain N prediction sub-models;
[0033] A prediction model is constructed using the N prediction sub-models.
[0034] The second aspect of the present application provides a method for predicting thermal stability mutations of a GPCR, comprising:
[0035] Obtaining the protein sequence of the target GPCR;
[0036] The prediction model constructed according to the method for constructing a GPCR thermal stability mutation prediction model is used to predict the thermal stability mutation sites in the protein sequence of the target GPCR.
[0037] The third aspect of the present application provides a device for constructing a GPCR thermal stability mutation prediction model, comprising:
[0038] Pre-training module, used to obtain pre-training models;
[0039] A homologous sequence acquisition module is used to obtain the homologous sequence of the sample GPCR;
[0040] An evolution module, used for performing self-supervisory training on the pre-trained model according to the homologous sequence to obtain an evolutionary fine-tuning model;
[0041] The fine-tuning module is used to obtain the mutation thermal stability sample data of the sample GPCR, and supervise the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model.
[0042] The fourth aspect of the present application provides a device for predicting thermal stability mutations of a GPCR, comprising:
[0043] An acquisition module, used to obtain the protein sequence of the target GPCR;
[0044] A prediction module is used to predict the thermal stability mutation site in the protein sequence of the target GPCR using a prediction model constructed according to any of the above methods for constructing a thermal stability mutation prediction model for GPCR.
[0045] A fifth aspect of the present application provides an electronic device, including:
[0046] Processor; and
[0047] The memory stores executable codes thereon, and when the executable codes are executed by the processor, the processor is caused to execute the method as described above.
[0048] A sixth aspect of the present application provides a computer-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described above.
[0049] The technical solution provided by this application may have the following beneficial effects:
[0050] The technical solution of the present application is to perform evolutionary fine-tuning based on the pre-trained model, so that the model is more suitable for a specific GPCR family, and further utilize the mutation thermal stability sample data of GPCR with limited labeled data, so that the model can ultimately make directional predictions based on the sample data of specific functions, and the obtained trained prediction model can be applied to the prediction of thermal stability of site mutations in GPCRs. Such a design can effectively predict the thermal stability mutation sites in GPCRs, so that R&D personnel can conduct targeted experimental tests in the predicted thermal stability mutation sites, thereby helping to assist R&D personnel to reduce experimental operations, improve R&D efficiency, and reduce R&D costs.
[0051] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The above and other objects, features and advantages of the present application will become more apparent by describing in more detail exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.
[0053] Figure 1 It is a schematic diagram of the process of constructing a GPCR thermal stability mutation prediction model shown in the examples of the present application;
[0054] Figure 2 is another schematic flow chart of the method for constructing a GPCR thermal stability mutation prediction model shown in the examples of the present application;
[0055] Figure 3 is a schematic diagram of the process of predicting the thermal stability mutation of GPCR shown in the examples of the present application;
[0056] Figure 4Schematic diagram of the structure of a device for constructing a GPCR thermal stability mutation prediction model shown in the examples of the present application;
[0057] Figure 5 is another structural schematic diagram of a device for constructing a GPCR thermal stability mutation prediction model shown in an embodiment of the present application;
[0058] Figure 6 is a schematic diagram of the structure of a GPCR thermal stability mutation prediction device shown in an embodiment of the present application;
[0059] Figure 7 It is a schematic diagram of the structure of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0061] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0062] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.
[0063] Among the related technologies, studying the structural stability of GPCRs and their structural analysis technology is crucial for the development of new drugs. However, due to the limitations of structural biology and protein engineering technology in GPCR structural analysis (time-consuming and labor-intensive), the prediction of GPCR thermal stability mutation sites is becoming increasingly important. The current methods for modifying the thermal stability of GPCR mutants have certain limitations and require a lot of time and R&D costs.
[0064] In view of the above problems, the present application provides a method for constructing a GPCR thermal stability mutation prediction model, which can construct a prediction model to more efficiently and accurately predict which site mutations of the GPCR have the possibility of contributing to stability, thereby saving time and effort and reducing R&D costs. The technical solution of the present application embodiment is described in detail below in conjunction with the accompanying drawings.
[0065] Figure 1 It is a schematic diagram of the process of constructing a GPCR thermal stability mutation prediction model shown in the examples of the present application.
[0066] See also Figure 1 The method for constructing a GPCR thermal stability mutation prediction model shown in the examples of the present application comprises:
[0067] S110, obtaining a pre-trained model.
[0068] It can be understood that among the hundreds of GPCRs currently known, due to the insufficient amount of label data of GPCR-related mutation thermal stability sample data, in this step, the target model can be pre-trained using unlabeled pre-training sample data to obtain a pre-trained model. Based on the protein properties of the GPCR itself, the pre-training sample data can select sample sequences of natural proteins as training data to train the target model.
[0069] Furthermore, the sample sequence of natural proteins can be interpreted as sentences, and the constituent amino acids of the sample sequence can be interpreted as single words, thereby corresponding to the natural language processing (NLP) scenario. Therefore, optionally, the target model can be a machine learning model applied to the natural language processing scenario, such as the BERT model (Bidirectional Encoder Representations from Transformer). Of course, other types of pre-trained language models can also be used, which are not limited here. In order to improve data processing efficiency, in other embodiments, known open source pre-trained models can also be directly used.
[0070] S120, obtaining a homologous sequence of the sample GPCR, performing self-supervised training on the pre-trained model according to the homologous sequence, and obtaining an evolutionary fine-tuning model.
[0071] In this step, based on the protein sequence corresponding to each sample GPCR, evolutionary information can be searched in a known protein data sequence database using relevant technologies to obtain homologous sequences related to the evolution of each sample GPCR in the known protein sequence database.
[0072] Furthermore, by forming training data with homologous sequences corresponding to each sample GPCR, the pre-trained model is self-supervised trained to obtain a new evolutionary fine-tuning model that can be applied to the GPCR scenario. The training task of this step can be the same as the pre-training task of the above step. Through this step, the evolutionary information related to the sample GPCR is introduced into the pre-trained model for training, and an evolutionary fine-tuning model that has learned the GPCR family information is obtained, so that the evolutionary fine-tuning model can be applied to predictions in a specific field, namely GPCR.
[0073] S130, obtaining mutation thermal stability sample data of the sample GPCR, and performing supervised training on the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model.
[0074] In this step, the above-mentioned evolutionary fine-tuning model can be further fine-tuned by collecting the mutation thermal stability sample data of the labeled sample GPCR, performing targeted supervised training on the evolutionary fine-tuning model for specific functions, and obtaining a prediction model that can predict the mutation thermal stability of GPCR.
[0075] The method of the present application migrates sample GPCRs based on a trained pre-trained model, further trains the pre-trained model with homologous sequences related to the evolution of the sample GPCR, thereby introducing features related to the sample GPCR for application to a specific protein family, and finally, with a limited amount of mutation thermal stability sample data, uses sample data with real labels to perform targeted training in specific functions, and uses supervised training to fine-tune the evolutionary fine-tuning model, so that the obtained prediction model can improve the prediction accuracy in the thermal stability of GPCR mutation sites.
[0076] In summary, from this example, it can be seen that the solution of this application is based on evolutionary fine-tuning of the pre-trained model, making the model more suitable for a specific GPCR family, and further using the mutation thermal stability sample data of GPCR with limited labeled data, so that the model can ultimately make directional predictions based on sample data of specific functions, and the obtained trained prediction model can be applied to the prediction of thermal stability of site mutations in GPCRs. Such a design can effectively predict the thermal stability mutation sites in GPCRs, so that R&D personnel can conduct targeted experimental tests in the predicted thermal stability mutation sites, thereby helping to assist R&D personnel to reduce experimental operations, improve R&D efficiency, and reduce R&D costs.
[0077] Figure 2 1 is another schematic diagram of the process of constructing a GPCR thermal stability mutation prediction model shown in the examples of the present application.
[0078] See also Figure 2, a method for constructing a GPCR thermal stability mutation prediction model shown in another embodiment of the present application includes:
[0079] S210, obtaining a pre-trained model.
[0080] In this step, a large number of unlabeled sample sequences of natural proteins, for example tens of millions, can be obtained from known databases, where the known databases include protein databases such as Pfam database, Uniprot database and BFD database, which are not limited here.
[0081] In this embodiment, the language model may be a standard BERT model with 12 layers, 768 dimensions and 12 heads. Each sample sequence in the pre-training sample data may be input into the BERT model for self-supervised task training, and the self-supervised task training may include a character masking task and a next sentence prediction task.
[0082] In a specific embodiment, the character masking task is to randomly mask some characters in each sample sequence, for example, to mask 15% to 20% of the characters in a single sample sequence. Furthermore, the character masking method can be customized. For example, when it is necessary to mask 15% of the characters in the sample sequence, 80% of the characters to be masked can be replaced with masking characters, 10% of the characters to be masked can be replaced with random characters, and the remaining 10% of the characters are not processed. It can be understood that for the sample sequence of protein, a single character is a single amino acid in the sequence. The amino acids in the sample sequence are masked by the character masking task to achieve self-supervised training at the word level.
[0083] In a specific implementation, the next sentence prediction task is to randomly divide each sample sequence into two parts, each part is regarded as a sentence, and the two sentences are input into the BERT model at the same time, and the second sentence is replaced with a random sentence with a probability of 40% to 60%, for example, the second sentence is replaced with a sentence composed of random amino acid characters with a probability of 50%, and the model is allowed to predict whether the two sentences are coherent, so as to realize self-supervised training at the sentence level. Of course, when performing the next sentence prediction task, one sentence, that is, one sample sequence, can also be input.
[0084] Furthermore, in order to enable the language model to effectively perform self-supervised task training, in this embodiment, the parameters of the language model can be set according to preset weights. For example, the weight parameters in the BERT model can be assigned according to the weights trained in the TAPE (TasksAssessing Protein Embeddings) model published by Rao et al., thereby completing the architecture of the pre-trained model. Compared with directly using random parameters in the model, such a design can adjust the language model to be more suitable for the current task.
[0085] S220, performing sequence comparison in a preset database according to the protein sequence of the sample GPCR, and obtaining a homologous sequence of the sample GPCR.
[0086] In this step, all currently known GPCR types in the GPCR family can be used as sample GPCRs, or GPCRs with richer experimental data can be selected as sample GPCRs, for example, GPCRs such as the activated state conformation of the A2AR receptor bound to an activator (agonist 50-N-ethylcarboxamidoadenosine-bound human adenosine receptor, hereinafter referred to as 2YDV), the inactivated state conformation of the A2AR receptor bound to an antagonist (antagonist ZM-241385-bound A2AR, hereinafter referred to as 3PWH), the inactivated state conformation of the β1AR receptor bound to an antagonist (antagonist cyanopindolol-bound turkeyb1-adrenergic receptor, hereinafter referred to as 4BVN), the activated intermediate state of the NTSR receptor bound to an agonist (agonist NTS1-bound rat neurotensin receptor, hereinafter referred to as 4XEE), and the antagonist-bound AT1R (antagonist ZD7155-bound human angiotensinreceptor type 1, hereinafter referred to as 4YAY) can be selected. For the protein sequence of each sample GPCR, multiple sequence alignment can be performed using DeepMSA (deep multiple sequence alignment). For example, by combining multiple algorithms and based on the inherent characteristics of the protein sequence of the sample GPCR, sequences related to the protein sequence evolution of the sample GPCR can be searched in different protein sequence databases and sequence alignment can be performed. Finally, based on the sequence alignment results, homologous sequences related to the protein sequence evolution of each sample GPCR can be screened and obtained.
[0087] In order to obtain more abundant homologous sequences as possible, in some embodiments, a sequence alignment is performed in a first protein database according to the protein sequence of the sample GPCR to obtain a first group of multiple sequence alignment results; wherein: when the number of valid sequences in the first group of multiple sequence alignment results is greater than or equal to a preset number threshold, the valid sequence is output as the homologous sequence of the sample GPCR; when the number of valid sequences in the first group of multiple sequence alignment results is less than the preset number threshold, a sequence alignment is performed in the Pth protein database according to the protein sequence of the sample GPCR in sequence, and the alignment result is combined with the previous P-1 group of multiple sequence alignment results to obtain the Pth group of multiple sequence alignment results, until the number of valid sequences in the Pth group of multiple sequence alignment results is greater than or equal to the preset number threshold, the sequence alignment is terminated, and the valid sequence is output as the homologous sequence of the sample GPCR, wherein P is a natural number greater than 1. In other words, DeepMSA can be used to search in protein databases of different sources known at present, and obtain the homologous sequence of the sample GPCR in the corresponding protein database by using different sequence alignment methods.
[0088] In one example, for example, three different sequence alignment methods are integrated and sequence alignment is performed continuously. Specifically, for the protein sequence of one of the sample GPCRs, a sequence search is first performed in the Uniclust30 protein monomer sequence database using the HHblits software to obtain a first set of multiple sequence alignment results; the number of valid sequences Nf obtained in the first set of multiple sequence alignment results is determined, and the number of valid sequences Nf is compared with a preset number threshold, such as 128; if the number of valid sequences Nf in the first set of multiple sequence alignment results is greater than or equal to 128, then the valid sequence is output; if it is less than 128, then the Jackhmmer tool is used to search the sequence in the Uniclust30 protein monomer sequence database. f database, and combined with the results of the HHblits software search, the second set of multiple sequence alignment results are obtained; if the number of valid sequences Nf in the second set of multiple sequence alignment results is greater than or equal to 128, the valid sequences are output; if less than 128, the HMMsearch tool is used to search in the Metaclust database, and the results of the HHblits software and Jackhmmer tool search are combined to obtain the third set of multiple sequence alignments, and finally the valid sequences in the third group are output as the homologous sequences of the sample GPCR. For example, in this example, through the above three sequence alignment methods, a total of 121,787 evolutionarily related homologous sequences can be obtained for the five sample GPCRs, namely 2YDV, 3PWH, 4BVN, 4XEE and 4YAY.
[0089] It is to be understood that, in addition to the above sequence alignment methods, other methods may also be used, such as Blast sequence alignment, which is not limited here.
[0090] S230, performing self-supervised training on the pre-trained model for a character masking task and a next sentence prediction task according to the homologous sequence to obtain an evolutionary fine-tuning model.
[0091] Compared with the broad spectrum of natural protein sample sequences used in step S210, this step selectively selects homologous sequences related to GPCR evolution as training data to train the pre-trained model, so that the pre-trained model can incorporate the characteristics of the GPCR family. The same parameters and training tasks can be selected for training.
[0092] For example, on a pre-trained model such as the BERT model, self-supervised training is performed based on the above 121,787 homologous sequences. The training tasks are the same as those of the pre-trained model, namely, the character masking task and the next sentence prediction task, so as to obtain a trained evolutionary fine-tuning model.
[0093] S240, obtaining mutation thermal stability sample data of the sample GPCR, wherein the mutation thermal stability sample data includes positive sample data and negative sample data.
[0094] In order to improve the accuracy of the model prediction results, this embodiment uses the thermal stability test data of all sites of alanine scanning mutations of multiple sample GPCRs with more mature experimental data. In this embodiment, the mutation thermal stability sample data can include the thermal stability test data of all sites of alanine scanning mutations of five sample GPCRs such as 2YDV, 3PWH, 4BVN, 4XEE, and 4YAY as the mutation thermal stability sample data. These thermal stability test data are all test results of radioligand receptor binding assay. Of course, in other embodiments, the mutation thermal stability sample data of other types of sample GPCRs can also be used.
[0095] In one embodiment, the first activity change before and after heating of the sample GPCR after alanine scanning mutation and the second activity change before and after heating of the corresponding wild-type sample GPCR are obtained respectively; the first activity change before and after heating is compared with a preset multiple threshold of the second activity change before and after heating to determine the positive sample data and the negative sample data corresponding to the sample GPCR.
[0096] For ease of understanding, specifically, according to the thermal stability test data obtained above, it can be determined that after each site in each sample GPCR is subjected to alanine scanning mutation and heated, the activity change before and after heating can be obtained for each site mutation, that is, the first activity change before and after heating. Among them, the first activity change before and after heating can be expressed by the ratio of the activity after heating after the site mutation to the activity before heating, or it can be expressed by the difference between the activity after heating and the activity before heating, without limitation. The wild-type sample GPCR refers to each site of the same sample GPCR that has not been scanned for mutation. Similarly, the second activity change before and after heating of the wild-type sample GPCR can be expressed by the ratio or difference between its activity after heating and its activity before heating.
[0097] Further, the activity change before and after the first heating is compared with the preset multiple threshold of the activity change before and after the second heating, and the thermal stability change of the sample GPCR after mutation at one of the sites is determined, and the corresponding thermal stability prediction label, such as the label of "increase" or "decrease", is assigned, and the sample type is divided according to the label content. For ease of understanding, for example, the activity change before and after the first heating is expressed as a ratio of 90%, and the activity change before and after the second heating is expressed as a ratio of 80%, and the preset multiple threshold is 100%. According to 90% greater than 80% (80%*100%=80%), it means that the thermal stability of the sample GPCR after the mutation and heating at this site is improved, and the corresponding label is assigned as "increase", and the mutation thermal stability sample data of the mutation site is the positive sample data. On the contrary, if the activity change before and after the first heating is 70%, according to 70% less than 80% (80%*100%=80%), it means that the thermal stability of the sample GPCR after the mutation and heating at this site is reduced, and the corresponding label is assigned as "decrease", and the mutation thermal stability sample data of the mutation site is the negative sample data. As shown in Table 1 below, when the preset multiple threshold is set to 100%, based on the changes in thermal stability of all sites of the above five sample GPCRs after alanine scanning mutation and the changes in thermal stability of the corresponding wild-type sample GPCRs, the activity change before and after the first heating is compared with the preset multiple threshold of the activity change before and after the second heating, and the following data is obtained after summary, including the number of positive sample data and negative sample data corresponding to the five sample GPCRs.
[0098] Table 1
[0099] Sample GPCR Number of positive samples / Number of negative samples Number of positive samples: Number of negative samples 2YD 38 / 281 1:7 3PWH 69 / 222 1:3 4BV 52 / 265 1:5 4XEE 85 / 215 1:3 4YAY 11 / 296 1:27
[0100] Among them, the number of sequences after alanine scanning mutation at the five sample GPCR sites is 1534, with a total of 1534 positive sample data and negative sample data. Each sample GPCR contains positive sample data and negative sample data.
[0101] It should be understood that the numerical setting of the preset multiplier threshold will directly affect the division results of the positive sample data and the negative sample data. Furthermore, in one embodiment, the K preset multiplier thresholds of the activity change before and after the first heating are compared with the activity change before and after the second heating, respectively, and the training data consisting of the positive sample data and the negative sample data determined at each preset multiplier threshold are obtained, respectively, and K is a positive integer greater than or equal to 2. In other words, the preset multiplier threshold is not limited to the above-mentioned 100%, but can also be any value between 90% and 130%, such as 110%, 120%, etc. Obviously, when the numerical values of the preset multiplier thresholds are different, within the same batch of mutation thermal stability sample data, different positive sample data and negative sample data can be divided according to the aforementioned division method.
[0102] S250, supervised training is performed on the evolutionary fine-tuning model according to the positive sample data and negative sample data of each sample GPCR to obtain a prediction model.
[0103] It should be understood that, based on the above steps, different positive and negative sample data can be divided to form training data under different preset multiple thresholds, and the same training data includes a training set and a validation set. Using different training data to train the evolutionary fine-tuning model, the trained model obtained in the end will be different. Therefore, in one embodiment, the evolutionary fine-tuning model is supervised and trained using the training data under each preset multiple threshold to obtain K candidate prediction models.
[0104] In order to improve the prediction accuracy of the prediction model, optionally, in one embodiment, the mutation thermal stability sample data of N sample GPCRs are divided according to the N-fold cross-validation method to obtain N groups of training data, wherein each group of training data includes a training set composed of the mutation thermal stability sample data of (N-1) sample GPCRs and a validation set composed of the mutation thermal stability sample data of the remaining 1 sample GPCR.
[0105] For ease of understanding, for example, when the preset multiple thresholds include three types such as 100%, 110%, and 120%, three trained candidate prediction models can be obtained based on the corresponding training data. Among them, each candidate prediction model can be divided according to the N-fold cross-validation method to obtain N groups of training data during the training process. For example, when N is 5, the mutant thermal stability sample data of the above-mentioned 5 sample GPCRs are divided according to the five-fold cross-validation method at the same preset multiple threshold, and 5 groups of training data can be obtained. Each group of training data includes a training set consisting of mutant thermal stability sample data corresponding to 4 of the sample GPCRs and a validation set consisting of mutant thermal stability sample data of the remaining 1 sample GPCR. The mutant thermal stability sample data in each training set and validation set include positive sample data and negative sample data obtained by dividing at the corresponding preset multiple threshold.
[0106] Taking the preset multiple threshold of 100% as an example, the mutation thermal stability sample data of 5 sample GPCRs can be divided into positive sample data and negative sample data at the multiple threshold of 100%, and divided into 5 groups of training data according to the five-fold cross-validation method. The evolutionary fine-tuning model of the preset multiple threshold is trained using the 5 groups of training data to obtain the corresponding trained candidate prediction model. Similarly, the positive and negative sample data can be divided and trained according to the preset multiple thresholds of 110% and 120%, respectively, to obtain the corresponding candidate prediction models.
[0107] In one embodiment, K candidate prediction models are scored using preset evaluation indicators to obtain scoring results for each candidate prediction model; and the candidate prediction model whose scoring results meet preset conditions is selected as the prediction model. In other words, the candidate prediction models obtained under different preset multiple thresholds can be scored according to the preset evaluation indicators, so that the best candidate prediction model can be selected and put into use as the prediction model.
[0108] In order to quickly and objectively obtain the scoring results of different candidate prediction models, in a specific implementation, the prediction sample type of the candidate prediction model under different preset multiple thresholds of the validation set can be determined; the precision and recall of each candidate prediction model can be determined according to the prediction sample type; and the candidate prediction model with the largest precision and recall can be selected as the prediction model. That is to say, in this embodiment, the candidate prediction models are scored and screened using precision and recall. Of course, in other embodiments, evaluation indicators such as accuracy, auc, f1 score, determination coefficient R2, mean shared error MSE, square root error RMSE, etc. can also be used to score each candidate prediction model, without limitation here.
[0109] Furthermore, the predicted sample type is the type of positive or negative sample predicted by the candidate prediction model for the mutation thermal stability sample data of any sample GPCR in the validation set. The predicted sample type is compared with the corresponding real sample type, and the number of true positive results (TP), the number of false positive results (TN), the number of true negative results (FP), and the number of false negative results (FN) are counted respectively, so that the precision and recall are calculated according to the following formula.
[0110]
[0111]
[0112] Among them, the true positive result means that the predicted sample type output by the candidate prediction model is that a certain site mutation of the sample GPCR can improve thermal stability, and the actual sample type is that the site mutation can improve thermal stability; the true negative result means that the predicted sample type is that a certain site mutation cannot improve thermal stability, and the actual sample type is that the site mutation cannot improve thermal stability; the false positive result means that the predicted sample type is that a certain site mutation can improve thermal stability, and the actual sample type is that the site mutation cannot improve thermal stability; the false negative result means that the predicted sample type is that a certain site mutation cannot improve thermal stability, and the actual sample type is that the site mutation can improve thermal stability.
[0113] According to the above formula, the positive sample data and negative sample data in the above five sample GPCRs are divided according to the above three preset multiple thresholds of 100%, 110% and 120%, and the corresponding training data including the training set and the validation set are formed (the training data can be the training data divided by the five-fold cross validation method). The precision and recall of the candidate prediction model under different preset multiple thresholds are calculated when each sample GPCR is used as the validation set in turn, and the precision and recall of the traditional computational biology method based on Rosetta software are used as the reference standard. The specific data are shown in Table 2 below.
[0114] Table 2
[0115]
[0116] From the above table 2, under different multiple thresholds, compared with other multiple thresholds, the candidate prediction model has the best comprehensive performance when the multiple threshold is set to 110%, and the accuracy of the prediction results for different validation sets is relatively stable, and the recall rate is improved when the accuracy does not fluctuate much; in addition, the prediction difficulty of the validation set 4YAY is greater than that of the other 4 validation sets, and when the multiple threshold is 110%, the prediction effect of the validation set 4YAY can be on par with the prediction result of the traditional Rosseta model. Therefore, optionally, the multiple threshold of 110% can be set as the preset multiple threshold of the prediction model in actual application, and correspondingly, the candidate prediction model trained under this multiple threshold can be used as the final prediction model. The use of this prediction model for GPCR thermal stability prediction can improve the prediction accuracy.
[0117] Furthermore, in order to avoid the random influence of the neural network in the optimization process, in one embodiment, the evolutionary fine-tuning model is set according to M groups of random parameters to obtain M corresponding evolutionary fine-tuning sub-models; wherein M is a positive integer greater than or equal to 2; the mutant thermal stability sample data of the sample GPCR is used as training data to train the M evolutionary fine-tuning sub-models respectively to obtain M corresponding trained fine-tuning sub-models, and the M fine-tuning sub-models are used to construct a prediction model.
[0118] However, in order to avoid the random influence of the neural network in the optimization process as much as possible and improve the accuracy of the prediction results of the final prediction model, in this embodiment, M groups of random parameters can be set for the evolutionary fine-tuning model before training the evolutionary fine-tuning model.
[0119] For ease of understanding, taking M equal to 3 as an example, the evolutionary fine-tuning model is first set according to three different groups of random parameters, thereby obtaining three evolutionary fine-tuning sub-models with different random parameters, that is, the initial difference between the three evolutionary fine-tuning sub-models is only the difference in the initial random parameters. Specifically, the three groups of random parameters can be the random seeds of the neural network, such as random.seed() in python, which means the initial parameters of the deep learning neural network. By fixing the random seeds, the initial parameters of the neural network are fixed. Of course, the number of groups of random parameters is not limited here, and can also be 4 groups, 5 groups, 6 groups, 7 groups, etc., so as to obtain the corresponding number of evolutionary fine-tuning sub-models. After setting the random parameters of the evolutionary fine-tuning sub-model, the mutant thermal stability sample data of the sample GPCR can be used as training data to train the above three evolutionary fine-tuning sub-models respectively, and the three trained fine-tuning sub-models are used to jointly construct a prediction model.
[0120] In a specific embodiment, each fine-tuning sub-model independently predicts the protein sequence of a target GPCR input and outputs the probability value of the corresponding target GPCR site mutation having thermal stability after heating; the probability values output by each fine-tuning sub-model for the same site are averaged as the prediction result of the prediction model, or the probability values output by each fine-tuning sub-model for the same site are obtained according to the corresponding preset weights to obtain the final probability value and serve as the prediction result of the prediction model. It is predicted which sites in the protein sequence of the GPCR can improve thermal stability after mutation, and the probability value of the corresponding site mutation to improve thermal stability.
[0121] That is to say, when a protein sequence of a target GPCR to be predicted is input into the prediction model constructed by the above-mentioned multiple fine-tuning sub-models, during the calculation process of the prediction model, multiple fine-tuning sub-models respectively output preliminary prediction results for the protein sequence to be predicted, that is, each trained fine-tuning sub-model can predict which sites in the protein sequence of the target GPCR can improve thermal stability after mutation and heating, and the corresponding probability value of improving thermal stability. For the same mutation site predicted by multiple fine-tuning sub-models, the average value of the probability value predicted by each fine-tuning sub-model or the probability value after each probability value is weighted according to the corresponding preset weight is calculated, so as to obtain a unique probability value of the mutation site as the output result of the prediction model, and finally the prediction model can summarize and output all sites in the protein sequence of the GPCR that can improve thermal stability after mutation and the corresponding probability value.
[0122] In order to improve the prediction accuracy of each fine-tuning sub-model, when training the evolutionary fine-tuning sub-model, in a specific embodiment, when there are N sample GPCRs, the same set of training data is used to supervise the training of M evolutionary fine-tuning sub-models respectively, to obtain the corresponding M trained fine-tuning sub-models, and the M fine-tuning sub-models are used to construct a corresponding prediction sub-model to obtain N prediction sub-models; and the prediction model is constructed using the N prediction sub-models.
[0123] For ease of understanding, taking the above five sample GPCRs as an example, after dividing the training set and validation set according to the above five-fold cross-validation method and forming the corresponding five groups of training data, the first group of training data is used to train the three evolutionary fine-tuning sub-models to obtain the three trained fine-tuning sub-models under the first group of training data, and a prediction sub-model is constructed by the three trained fine-tuning sub-models. Similarly, the second group of training data is used to train the three evolutionary fine-tuning sub-models to obtain a prediction sub-model constructed by the three fine-tuning sub-models trained under the second group of training data. Similarly, the fifth group of training data is used to train the three evolutionary fine-tuning sub-models to obtain a prediction sub-model constructed by the three fine-tuning sub-models trained under the fifth group of training data, that is, five prediction sub-models are finally obtained.
[0124] Further, in one embodiment, a prediction model is constructed using N prediction sub-models, including: each prediction sub-model independently predicts a protein sequence of a target GPCR input and outputs a probability value that the corresponding site mutation in the target GPCR has thermal stability after heating; the probability values output by each prediction sub-model for the same site are averaged as the prediction result of the prediction model, or the probability values output by each prediction sub-model for the same site are obtained according to the corresponding preset weights to obtain a final probability value and used as the prediction result of the prediction model. In other words, when there are multiple prediction sub-models, each prediction sub-model can work independently and predict the protein sequence of the input target GPCR, and the prediction result of the prediction model is a combination of the prediction results of each prediction sub-model. By predicting together with multiple prediction sub-models, the robustness of the prediction result can be improved.
[0125] In summary, the technical solution of the present application, after a large number of natural protein sample sequences are used to establish a pre-training model through unlabeled self-supervised tasks, a large number of homologous sequences of GPCRs are further trained through unlabeled self-supervised tasks to obtain an evolutionary fine-tuning model, and finally a small amount of labeled and more accurate positive and negative sample data divided by a preset multiple threshold is used as training data for supervised learning, so that the evolutionary fine-tuning model can be adjusted in a targeted manner to obtain a prediction model, and the prediction model of the present application can be jointly predicted by multiple prediction sub-models or multiple fine-tuning sub-models to reduce the random effects in the optimization process and improve the accuracy of the prediction results. The method of the present application does not require a large amount of labeled real experimental data. It can gradually evolve and fine-tune the pre-training model obtained from a large amount of unlabeled data with the help of a deep learning model framework to obtain a prediction model with a high accuracy, thereby predicting the possible thermal stability mutation sites of GPCRs, reducing unnecessary experimental processes, improving research and development efficiency, and reducing research and development costs.
[0126] Figure 3 Schematic diagram of the process of predicting thermal stability mutations of GPCRs shown in the examples of the present application.
[0127] See also Figure 3 The method for predicting thermal stability mutations of GPCRs shown in the examples of the present application comprises:
[0128] S310, obtaining the protein sequence of the target GPCR.
[0129] In this step, the protein sequence of any type of GPCR can be obtained as input data for the prediction model.
[0130] S320, predicting the thermal stability mutation site in the protein sequence of the target GPCR according to the prediction model.
[0131] In this step, the prediction model constructed in any of the above embodiments can be used to predict the sites in the protein sequence of the target GPCR where the thermal stability is improved after the mutation. For example, the protein sequence of a site of the target GPCR after the mutation can be input into the prediction model to predict the thermal stability of the site after the mutation, such as improvement or decrease.
[0132] The prediction method of the present application can quickly and with high accuracy predict and evaluate thermal stability mutation sites based on the prediction model, which is convenient for R&D personnel to refer to and experiment, reduce R&D costs, and improve R&D efficiency.
[0133] Corresponding to the aforementioned application function implementation method embodiment, the present application also provides a device for constructing a GPCR thermal stability mutation prediction model, a GPCR thermal stability mutation prediction device, an electronic device and corresponding embodiments.
[0134] Figure 4 It is a schematic diagram of the structure of the device for constructing the thermal stability mutation prediction model of GPCR shown in the examples of the present application.
[0135] See also Figure 4 The present application embodiment provides a device for constructing a GPCR thermal stability mutation prediction model, which includes a pre-training module 410, a homologous sequence acquisition module 420, an evolution module 430 and a fine-tuning module 440, wherein:
[0136] The pre-training module 410 is used to obtain a pre-training model.
[0137] The homologous sequence 420 acquisition module is used to acquire the homologous sequence of the sample GPCR.
[0138] The evolution module 430 is used to perform self-supervisory training on the pre-trained model according to the homologous sequence to obtain an evolutionary fine-tuning model.
[0139] The fine-tuning module 440 is used to obtain the mutation thermal stability sample data of the sample GPCR, and perform supervised training on the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model.
[0140] Further, the homologous sequence acquisition module 420 is used to perform sequence alignment in a preset database according to the protein sequence of the sample GPCR, respectively, to obtain the homologous sequence of the sample GPCR. Specifically, the homologous sequence acquisition module 420 is used to perform sequence alignment in a first protein database according to the protein sequence of the sample GPCR, and obtain a first group of multiple sequence alignment results; wherein: when the number of valid sequences in the first group of multiple sequence alignment results is greater than or equal to the preset number threshold, the valid sequence is output as the homologous sequence of the sample GPCR; when the number of valid sequences in the first group of multiple sequence alignment results is less than the preset number threshold, the sequence alignment is performed in the Pth protein database according to the protein sequence of the sample GPCR, and the alignment result is combined with the previous P-1 group of multiple sequence alignment results to obtain the Pth group of multiple sequence alignment results, until the number of valid sequences in the Pth group of multiple sequence alignment is greater than or equal to the preset number threshold, the sequence alignment is terminated, and the valid sequence is output as the homologous sequence of the sample GPCR, wherein P is a natural number greater than 1.
[0141] The evolution module 430 is used to perform self-supervised training of the character masking task and the next sentence prediction task on the pre-trained model according to the homologous sequence to obtain an evolutionary fine-tuning model.
[0142] See also Figure 5 The fine-tuning module 440 of the present application includes a sample acquisition submodule 441 and a training submodule 442. The sample acquisition submodule 441 is used to respectively obtain the first activity change before and after heating of the sample GPCR after alanine scanning mutation and the second activity change before and after heating of the corresponding wild-type sample GPCR; compare the first activity change before and after heating with the preset multiple threshold of the second activity change before and after heating, and determine the positive sample data and negative sample data corresponding to the sample GPCR.
[0143] Among them, the sample acquisition submodule 441 is also used to compare the K preset multiplier thresholds of the activity change before and after the first heating with the activity change before and after the second heating, and obtain training data consisting of positive sample data and negative sample data determined under each preset multiplier threshold, where K is a positive integer greater than or equal to 2.
[0144] Furthermore, the sample acquisition submodule 441 is also used to divide the mutation thermal stability sample data of N sample GPCRs according to the N-fold cross-validation method to obtain N groups of training data, wherein each group of training data includes a training set consisting of mutation thermal stability sample data of (N-1) sample GPCRs and a validation set consisting of mutation thermal stability sample data of the remaining 1 sample GPCR.
[0145] The training submodule 442 is used to supervise the evolution fine-tuning model using the training data under each preset multiple threshold, and obtain K candidate prediction models. The training submodule 442 is used to score the K candidate prediction models using preset evaluation indicators, and obtain the scoring results of each candidate prediction model; the candidate prediction model whose scoring results meet the preset conditions is selected as the prediction model. In a specific embodiment, the training submodule 442 is used to determine the prediction sample type of the candidate prediction model of the verification set under different preset multiple thresholds; determine the precision and recall of each candidate prediction model according to the prediction sample type; select the candidate prediction model with the largest precision and recall as the prediction model.
[0146] In some embodiments, the training submodule 442 is used to set the evolutionary fine-tuning model according to M groups of random parameters to obtain M corresponding fine-tuning sub-models; wherein M is a positive integer greater than or equal to 2; the mutant thermal stability sample data of the sample GPCR is used as training data to train the M fine-tuning sub-models respectively to obtain M corresponding trained fine-tuning sub-models, and the M fine-tuning sub-models are used to construct a prediction model.
[0147] In some embodiments, the training submodule 442 is used to perform supervised training on M fine-tuning sub-models using the same set of training data, obtain corresponding M trained fine-tuning sub-models, and use the M fine-tuning sub-models to construct a corresponding prediction sub-model to obtain N prediction sub-models; and use the N prediction sub-models to construct a prediction model.
[0148] The device for constructing the thermal stability mutation prediction model of GPCR of the present application sequentially obtains the pre-trained model by training on a large amount of unlabeled data through the pre-training module. The homologous sequence acquisition module can obtain a large number of unlabeled homologous sequences limited to a specific GPCR field, so that the evolution module can be directed evolved into an evolution module applied in the GPCR sequence based on the pre-training module. Finally, the fine-tuning module performs precise fine-tuning training based on a small amount of labeled GPCR mutation thermal stability sample data to obtain a prediction model that can accurately predict the thermal stability mutation sites of GPCRs.
[0149] Figure 6 Schematic diagram of the structure of the thermal stability mutation prediction device for GPCR shown in the embodiment of the present application.
[0150] See also Figure 6 , an embodiment of the present application further provides a device for predicting thermal stability mutations of a GPCR, which includes an acquisition module 610 and a prediction module 620, wherein:
[0151] The acquisition module 610 is used to acquire the protein sequence of the target GPCR.
[0152] The prediction module 620 is used to predict the thermal stability mutation sites in the protein sequence of the target GPCR according to the prediction model constructed by the above-mentioned method for constructing the thermal stability mutation prediction model of GPCR.
[0153] The thermal stability mutation prediction device of the present application can quickly and accurately predict sites in the target GPCR that have thermal stability after mutation, helping to improve research and development efficiency and save time and research and development costs.
[0154] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.
[0155] Figure 7 It is a schematic diagram of the structure of an electronic device shown in an embodiment of the present application.
[0156] See also Figure 7 , the electronic device 1000 includes a memory 1010 and a processor 1020 .
[0157] The processor 1020 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.
[0158] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, ROM can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that does not lose the stored instructions and data even after the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as a permanent storage device. In some other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as a dynamic random access memory. The system memory may store some or all instructions and data required by the processor at run time. In addition, the memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a readable and / or writable removable storage device, such as a laser disc (CD), a read-only digital versatile disc (such as a DVD-ROM, a double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (such as an SD card, a mini SD card, a Micro-SD card, etc.), a magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0159] The memory 1010 stores executable codes, and when the executable codes are processed by the processor 1020 , the processor 1020 can execute part or all of the methods described above.
[0160] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0161] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium) on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present application.
[0162] The embodiments of the present application have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for constructing a GPCR thermal stability mutation prediction model. It is characterized in that include: Get the pre-trained model; Perform sequence alignment in a preset database according to the protein sequence of the sample GPCR to obtain a homologous sequence of the sample GPCR, and perform self-supervised training on the pre-trained model according to the homologous sequence to obtain an evolutionary fine-tuning model; Obtaining mutation thermal stability sample data of the sample GPCR, and performing supervised training on the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model; Wherein, the step of obtaining the mutation thermal stability sample data of the sample GPCR comprises: Obtaining respectively the first activity change before and after heating of the sample GPCR after alanine scanning mutation and the second activity change before and after heating of the corresponding wild-type sample GPCR; comparing respectively the first activity change before and after heating with K preset multiple thresholds of the second activity change before and after heating, and obtaining respectively training data consisting of positive sample data and negative sample data determined at each preset multiple threshold, wherein K is a positive integer greater than or equal to 1; The step of performing supervised training on the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model includes: The evolutionary fine-tuning model is supervisedly trained using the training data under each of the preset multiplier thresholds to obtain K candidate prediction models; the K candidate prediction models are scored using preset evaluation indicators to obtain a scoring result for each candidate prediction model; and the candidate prediction model whose scoring result meets the preset conditions is selected as the prediction model.
2. The method according to claim 1, It is characterized in that The method of performing sequence comparison in a preset database based on the protein sequence of the sample GPCR includes: According to the protein sequence of the sample GPCR, a sequence alignment is performed in the first protein database to obtain the first set of multiple sequence alignment results; wherein: When the number of valid sequences in the first group of multiple sequence alignment results is greater than or equal to a preset number threshold, outputting the valid sequence as a homologous sequence of the sample GPCR; When the number of valid sequences in the first group of multiple sequence alignment results is less than a preset number threshold, sequence alignment is performed in the Pth protein database according to the protein sequence of the sample GPCR, and the alignment result is combined with the previous P-1 group of multiple sequence alignment results to obtain the Pth group of multiple sequence alignment results, until the number of valid sequences in the Pth group of multiple sequence alignment results is greater than or equal to the preset number threshold, the sequence alignment is terminated, and the valid sequence is output as the homologous sequence of the sample GPCR, where P is a natural number greater than 1.
3. The method according to claim 1, It is characterized in that The preset magnification threshold is selected from 90% to 130%.
4. The method according to claim 1, It is characterized in that The step of performing supervised training on the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model comprises: Setting the evolutionary fine-tuning model according to M groups of random parameters to obtain M corresponding evolutionary fine-tuning sub-models; wherein M is a positive integer greater than or equal to 2; The mutant thermal stability sample data of the sample GPCR is used as training data to train the M evolutionary fine-tuning sub-models respectively, to obtain the corresponding M trained fine-tuning sub-models, and the prediction model is constructed using the M fine-tuning sub-models.
5. The method according to claim 4, It is characterized in that If the sample GPCRs are N, N is a positive integer greater than or equal to 2; The method uses the mutant thermal stability sample data of the sample GPCR as training data to train M evolutionary fine-tuning sub-models respectively, obtains corresponding M trained fine-tuning sub-models, and constructs a prediction model using the M fine-tuning sub-models, including: The mutant thermal stability sample data of N sample GPCRs are divided according to the N-fold cross-validation method to obtain N groups of training data, wherein each group of training data includes a training set consisting of the mutant thermal stability sample data of (N-1) sample GPCRs and a validation set consisting of the mutant thermal stability sample data of the remaining 1 sample GPCR; Using the same set of training data to perform supervised training on M evolutionary fine-tuning sub-models respectively, to obtain corresponding M trained fine-tuning sub-models, and using the M fine-tuning sub-models to construct a corresponding prediction sub-model, to obtain N prediction sub-models; A prediction model is constructed using the N prediction sub-models.
6. A method for predicting thermal stability mutations of GPCRs, It is characterized in that include: Obtaining the protein sequence of the target GPCR; The prediction model constructed by the method for constructing a thermal stability mutation prediction model of a GPCR according to any one of claims 1 to 5 predicts the thermal stability mutation site in the protein sequence of the target GPCR.
7. A device for constructing a GPCR thermal stability mutation prediction model, It is characterized in that include: Pre-training module, used to obtain pre-training models; A homologous sequence acquisition module is used to perform sequence comparison in a preset database according to the protein sequence of the sample GPCR to obtain the homologous sequence of the sample GPCR; An evolution module, used for performing self-supervisory training on the pre-trained model according to the homologous sequence to obtain an evolutionary fine-tuning model; A fine-tuning module, used to obtain the mutation thermal stability sample data of the sample GPCR, and perform supervised training on the evolutionary fine-tuning model according to the mutation thermal stability sample data to obtain a prediction model; Wherein, the fine-tuning module includes a sample acquisition submodule and a training submodule; The sample acquisition submodule is used to respectively obtain the first activity change before and after heating of the sample GPCR after alanine scanning mutation and the second activity change before and after heating of the corresponding wild-type sample GPCR; respectively compare the first activity change before and after heating with K preset multiple thresholds of the second activity change before and after heating, and respectively obtain training data consisting of positive sample data and negative sample data determined at each preset multiple threshold, wherein K is a positive integer greater than or equal to 1; The training submodule is used to use the training data under each of the preset multiplier thresholds to perform supervised training on the evolutionary fine-tuning model to obtain K candidate prediction models; use preset evaluation indicators to score the K candidate prediction models to obtain the scoring results of each candidate prediction model; and select the candidate prediction model whose scoring results meet the preset conditions as the prediction model.
8. A device for predicting thermal stability mutations of GPCRs, It is characterized in that include: An acquisition module, used to obtain the protein sequence of the target GPCR; A prediction module, for predicting the thermal stability mutation sites in the protein sequence of the target GPCR using the prediction model constructed according to the method for constructing a thermal stability mutation prediction model of GPCR according to any one of claims 1 to 5.
9. An electronic device, It is characterized in that include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having executable codes stored thereon, which, when executed by a processor of an electronic device, causes the processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Delaminating and classifying method for G-protein-coupled receptor family
CN103258146A
GPCR(G Protein-Coupled Receptor)-drug interaction prediction method based on postprocessing study
CN104239751A