Molecular Property Prediction Method, Device, Intelligent Device, and Terminal
By pre-training the initial prediction model and training and tuning the sample molecular set, the target attribute prediction model is constructed, which solves the problem of complex molecular attribute testing process and low accuracy in the process of new drug discovery, and achieves efficient and accurate molecular attribute prediction.
Patent Information
- Application Number
- CN202011374210.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-30
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-11-30
AI Technical Summary
In the process of discovering new drugs in the prior art, the molecular attribute testing process is complex and costly, especially the accuracy of predicting attributes such as toxicity under a small amount of sample data is low.
By pre-training the initial prediction model and training and tuning the sample molecule set, a target attribute prediction model is constructed, and multiple sets of reference attribute molecular data and target attribute molecular data with supervised information are used for model training to improve the prediction accuracy of the model.
The accuracy of molecular attribute analysis of models trained based on a small number of sample data is improved, the demand for supervised information molecules is reduced, and the efficiency of molecular attribute prediction is improved.
Smart Images

Figure CN112420125B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer application technologies, and in particular, to a method, device, intelligent device, and terminal for predicting molecular properties. Background Art
[0002] In the process of new drug discovery, after a hit compound is determined through virtual screening, it is necessary to conduct experimental analysis on the properties of the hit compound to determine a lead compound. In traditional experiments, the process of testing drug molecular properties is complex, requiring the preparation of compounds and consuming manpower and material resources. At the same time, for some sensitive properties, such as toxicity, before the safe dose is determined, conducting clinical tests on humans poses great risks.
[0003] Using machine learning methods to predict molecular properties can reduce the need for traditional experiments and lower the cost of molecular property testing. Since there are multiple molecular properties to be predicted, the prediction effect of a unified model is relatively poor. If each property is predicted separately, due to certain properties, such as various toxicities, there is little supervised molecular data, and the prediction effect of the trained model is poor, resulting in a low analysis accuracy of the model trained based on a small amount of sample data for the relevant properties (such as toxicity) of molecules. Summary of the Invention
[0004] Embodiments of the present invention provide a method, device, intelligent device, and terminal for predicting molecular properties, which can obtain a better molecular analysis model and can perform molecular analysis more accurately.
[0005] On the one hand, embodiments of the present invention provide a method for predicting molecular properties, the method comprising:
[0006] Obtaining molecular features of a target molecule to be analyzed for a target property, where the target molecule refers to a molecule for which a target property parameter value needs to be analyzed under the target property;
[0007] Inputting the molecular features into a target property prediction model associated with the target property, and analyzing the molecular features through the target property prediction model to obtain a parameter value of the target property corresponding to the target molecule;
[0008] Wherein, the target property prediction model is obtained by training a basic prediction model with a sample molecule set, and the sample molecule set includes: M sample groups composed of sample molecular features and sample parameter values, and the sample parameter values in the sample groups are training supervision values of the sample molecules under the target property;
[0009] The basic prediction model is pre-trained from an initial prediction model using N sets of training molecules. Each set of training molecules includes: P training groups composed of training molecule features and training parameter values. The training parameter values in the training groups are the training supervision values of the training molecules under a reference attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute.
[0010] On the one hand, an embodiment of the present invention provides a molecular property prediction device, which includes:
[0011] An acquisition module, configured to acquire the molecular features of a target molecule to be analyzed for a target property. The target molecule refers to a molecule for which the target property parameter value needs to be analyzed under the target property.
[0012] A processing module, configured to input the molecular features into a target property prediction model associated with the target property, and analyze the molecular features through the target property prediction model to obtain the parameter value of the target property corresponding to the target molecule.
[0013] Wherein, the target property prediction model is obtained by training a basic prediction model using a sample molecule set. The sample molecule set includes: M sample groups composed of sample molecule features and sample parameter values. The sample parameter values in the sample groups are the training supervision values of the sample molecules under the target property.
[0014] The basic prediction model is pre-trained from an initial prediction model using N sets of training molecules. Each set of training molecules includes: P training groups composed of training molecule features and training parameter values. The training parameter values in the training groups are the training supervision values of the training molecules under a reference attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute.
[0015] On the one hand, an embodiment of the present invention provides an intelligent device, including a processor, an input interface, an output interface, and a memory. The processor, input interface, output interface, and memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the molecular property prediction method.
[0016] On the one hand, an embodiment of the present invention provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the molecular property prediction method.
[0017] In the embodiments of the present invention, for a model capable of analyzing the attribute parameter values of molecules under target attributes, it is obtained by further learning on the basis of other pre-trained prediction models. Specifically, an initial prediction model is pre-trained with multiple groups of reference attribute molecule data with supervision information (N training molecule sets) to obtain a basic prediction model, and the model is trained and optimized with target attribute molecule data with supervision information (sample molecule sets) to obtain a target attribute prediction model. The above method solves the problem of less supervised information molecule data for target attributes during the model training process, and improves the accuracy of the model for molecule attribute analysis trained based on a small amount of sample data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic diagram of a platform structure provided by an embodiment of the present invention;
[0020] Figure 2 It is a schematic flowchart of a method for predicting molecular attributes provided by an embodiment of the present invention;
[0021] Figure 3 It is a schematic flowchart of a method for training a prediction model provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic flowchart of another method for training a prediction model provided by an embodiment of the present invention;
[0023] Figure 5 It is a schematic flowchart of a model pre-training process provided by an embodiment of the present invention;
[0024] Figure 6 It is a schematic diagram of the architecture of a molecular attribute prediction system provided by an embodiment of the present invention;
[0025] Figure 7 It is a schematic diagram of a display interface provided by an embodiment of the present invention;
[0026] Figure 8 It is a schematic diagram of the structure of a molecular attribute prediction device provided by an embodiment of the present invention;
[0027] Figure 9 It is a schematic diagram of the structure of a prediction model training device provided by an embodiment of the present invention;
[0028] Figure 10It is a schematic structural diagram of an intelligent device provided by an embodiment of the present invention;
[0029] Figure 11 It is a schematic structural diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners
[0030] The discovery of drugs usually includes the following steps: 1. Target identification and confirmation; 2. Hit compound discovery; 3. Lead compound discovery and optimization; 4. Candidate compound confirmation and development; 5. Clinical trials. Based on the above process, this solution provides a drug platform, including functions such as screening drug molecules, predicting properties, and generating molecules. As Figure 1 shown, it is a schematic structural diagram of the drug platform provided by an embodiment of the present invention. The drug platform may specifically include a hit compound discovery module, a property prediction module, a synthetic route planning module, a protein structure prediction module, and a molecule generation module. Among them, the hit compound discovery module discovers hit compounds by using a virtual screening method based on the target structure or a virtual screening method based on the molecular structure. The property prediction model is specifically used to predict the properties of hit compounds such as absorption, distribution, metabolism, excretion, and toxicity. The synthetic route planning module is used to plan the synthetic route of compounds. The protein structure prediction module is used to predict the structure of proteins generated based on molecules. The molecule generation module is used to generate corresponding molecules based on the planned route. This solution provides a method for predicting molecular properties, which is specifically applied to the property prediction module in the drug platform and is used to predict the properties of the screened molecules (hit compounds) to achieve intelligent analysis of molecular properties.
[0031] Specifically, the molecular property prediction method proposed by the present invention specifically applies an artificial intelligence solution (Artificial Intelligence, AI). Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. In the molecular property prediction solution proposed in the embodiments of the present invention, by using an attribute prediction model to learn the mapping relationship between molecular features and molecular properties, when predicting the properties of a new molecule, the molecular features of the target molecule to be analyzed for the target property are obtained, and through the target property prediction model, the parameter value of the target property corresponding to the target molecule is obtained. Among them, in the training process of the target property prediction model, this solution is specifically divided into two stages. The first stage is to pre-train an initial prediction model using N training molecule sets to obtain a basic prediction model. Each training molecule set includes P training groups composed of training molecule features and training parameter values. Through pre-training, the basic prediction model can be made to have the ability to learn the mapping relationship between molecular features and molecular properties. The second stage is to train the basic prediction model using a sample molecule set. The sample molecule set includes M sample groups composed of sample molecule features and sample parameter values, so that the basic mapping model learns the mapping relationship between molecular features and target properties to obtain a target property prediction model. The model training process of this solution is to re-train based on the pre-trained model, so only a small amount of molecular data with supervised information under the target property is required to complete the training of the target property prediction model. And in the process of pre-training the prediction model, the model needs to be made to have the ability to learn the mapping relationship. Therefore, only the molecular data with supervised information under other reference properties is used to pre-train the initial prediction model. The sample size is sufficient, enabling the model to have enough samples for pre-training. This solution can complete the training of the target property prediction model based on a small amount of molecules with supervised information, improve the prediction accuracy of the target property prediction model in the actual property prediction process, and thus achieve a high prediction accuracy for molecular properties.
[0032] In one embodiment, the general process of the molecular property prediction method provided by this solution is as follows: ① Model pre-training, specifically, pre-train an initial prediction model through N training molecular sets to obtain a basic prediction model. Each training molecular set includes P training groups composed of training molecular features and training parameter values. The training parameter values in the training groups are the training supervision values of the training molecules in the training molecular set under the reference property. ② Model training and optimization, specifically, train the basic prediction model through a sample molecular set to obtain a target property prediction model. The sample molecular set includes M sample groups composed of sample molecular features and sample parameter values. The sample parameter values in the sample groups are the training supervision values of the sample molecules under the target property. ③ Obtain the molecular features of the target molecule to be analyzed for the target property. The target molecule refers to the molecule for which the target property parameter value needs to be analyzed under the target property. ④ Input the molecular features into the target property prediction model associated with the target property, and analyze the molecular features through the target property prediction model to obtain the parameter value of the target property corresponding to the target molecule.
[0033] In the above solution, the model is pre-trained through multiple sets of reference property molecular data with supervision information (N training molecular sets), and the model is trained and optimized through target property molecular data with supervision information (sample molecular set), which solves the problem of less supervised information molecular data for the target property during the model training process. Further, the target property prediction model is used to predict the target property of the molecule, which improves the prediction accuracy of the prediction model during the actual prediction process, making the prediction accuracy of the target property of the molecule high.
[0034] Based on the above description, an embodiment of the present invention provides a molecular property prediction method. Please refer to Figure 2 , and the molecular property prediction process may include the following steps S201 - S202:
[0035] S201. Obtain the molecular features of the target molecule to be analyzed for the target property.
[0036] In the embodiment of the present invention, a molecule is an overall formed by atoms combined together according to a certain bonding order and spatial arrangement. Different molecules have different molecular properties, and the molecular properties can be toxicity, solubility, analgesic property, etc. The target molecule refers to the molecule for which the target property parameter value needs to be analyzed under the target property. The target property can specifically be any one of the molecular properties, used to distinguish from the reference property. In specific implementation, the molecular features can specifically be the vector representation corresponding to the molecule, used to reflect the structural features, composition element features, etc. of the molecule.
[0037] In one implementation, the specific way for the intelligent device to obtain the molecular features of the target molecule to be subjected to target property analysis may be to process the target molecule through a graph neural network algorithm to obtain the molecular features of the target molecule. In one embodiment, the atoms in the target molecule are represented as the node features of a graph, the chemical bonds are represented as the edges of the graph, and the hidden state of the nodes is updated by weighted summation of the features of the neighboring nodes, and the iteration is continuously performed to obtain the molecular features of the target molecule. Optionally, the specific way for the intelligent device to obtain the molecular features of the target molecule to be subjected to target property analysis may also be to process the target molecule through a message neural network algorithm to obtain the molecular features of the target molecule. In one embodiment, a graph is used to represent the target molecule, the features of the edges are introduced, and the forward propagation process in the algorithm is divided into two stages: the information transfer stage to update the features of the nodes; the readout stage: calculate the feature expression of the whole graph, and obtain the molecular features of the target molecule based on this feature expression. Optionally, the target molecule may also be processed through an extended connectivity fingerprint (Morgan fingerprints) algorithm to obtain the molecular features of the target molecule, and the molecular features of the target molecule may specifically be a vector with a preset dimension (such as 1024 dimensions).
[0038] In one implementation, the molecular features of each molecule may be stored in the database, and the intelligent device may obtain the molecular features of the target molecule from the database.
[0039] In one implementation scenario, if the intelligent device is a server, the user can input the target molecule to be subjected to property determination on the specified page provided by the terminal. Among them, the property determination is specifically used to determine multiple molecular properties of the target molecule, and the target property is any one of the multiple molecular properties. The terminal uploads the target molecule to the server so that the server processes the target molecule to obtain the vector representation of the target molecule, that is, the molecular features of the target molecule. Alternatively, if the intelligent device is a terminal, the user can input the target molecule to be subjected to property determination on the specified page provided by the terminal, and the terminal processes the target molecule to obtain the corresponding molecular features of the target molecule.
[0040] S202. Input the molecular features into the target property prediction model associated with the target property, and analyze the molecular features through the target property prediction model to obtain the parameter value of the target property corresponding to the target molecule.
[0041] In an embodiment of the present invention, after the intelligent device obtains the molecular characteristics of the target molecule, it can input the characteristics into the target attribute prediction model associated with the target attribute, so as to analyze the molecular characteristics through the target attribute prediction model and obtain the parameter value of the target attribute corresponding to the target molecule. Among them, the parameter value can specifically be used to reflect the classification to which the target attribute corresponding to the target molecule belongs. For example, when the target attribute is "whether it can pass through the blood-brain barrier", when the parameter value is 1, it means that the classification to which the target attribute belongs is that it can pass through the blood-brain barrier, and when the parameter value is 0, it means that the classification to which the target attribute belongs is that it cannot pass through the blood-brain barrier. Alternatively, the parameter value can also be used to reflect the magnitude of the target attribute corresponding to the target molecule. For example, when the target attribute is the hydration energy, the parameter value of the hydration energy is used to reflect the magnitude of the hydration energy of the target molecule. When the hydration energy range is 0-100, if the parameter value of the hydration energy is 80, it means that the hydration energy of the target molecule is relatively large, and if the parameter value of the hydration energy is 2, it means that the hydration energy of the target molecule is relatively small.
[0042] Among them, the target attribute prediction model is obtained by training a basic prediction model through a sample molecule set. The sample molecule set includes: M sample groups composed of sample molecule characteristics and sample parameter values, and the sample parameter values in the sample groups are the training supervision values of the sample molecules under the target attribute; the basic prediction model is pre-trained through N training molecule sets for the initial prediction model. Each training molecule set includes: P training groups composed of training molecule characteristics and training parameter values, and the training parameter values in the training groups are the training supervision values of the training molecules under the reference attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute.
[0043] In one embodiment, the specific training method of the target attribute prediction model is as Figure 3 shown in the embodiment shown.
[0044] In an embodiment of the present invention, using the target attribute prediction model to predict the target attribute of a molecule realizes the prediction of molecular attributes in an artificial intelligence-based manner, improving the efficiency of molecular attribute prediction. Moreover, by pre-training the initial prediction model through multiple groups of reference attribute molecule data with supervision information (N training molecule sets) to obtain a basic prediction model, and training and optimizing the model through the target attribute molecule data with supervision information (sample molecule set) to obtain the target attribute prediction model, it also solves the problem of less supervised information molecular data under the target attribute during the model training process, improving the accuracy of the model for molecular attribute analysis trained based on a small amount of sample data.
[0045] Based on the above description, an embodiment of the present invention provides a training method for a target attribute prediction model. Please refer to Figure 3 . The training process of the target attribute prediction model may include the following steps S301-S304:
[0046] S301. Obtain N sets of training molecules. Each set of training molecules includes P training groups composed of training molecule features and training parameter values. The training parameter values in the training groups are the training supervision values of the training molecules under the reference attribute.
[0047] In the embodiments of the present invention, the terminal can obtain N sets of training molecules. Each set of training molecules includes: P training groups composed of training molecule features and training parameter values. The training parameter values in the training groups are the training supervision values of the training molecules under the reference attribute, and P is a positive integer. The training supervision value under the reference attribute can be specifically set in advance by R & D personnel. It can be the actual value of the training molecule under the reference attribute. For example, if the reference attribute is anesthetic property, the actual value corresponding to the anesthetic property of the training molecule can be obtained through detection, and this actual value is used as the training supervision value of the training molecule under the reference attribute.
[0048] In specific implementation, the reference attribute can include N categories. Each of the N sets of training molecules can have a corresponding relationship with a category of reference attribute. For example, in the N sets of training molecules, training molecule set 1 has a corresponding relationship with the reference attribute of the first category, and training molecule set 2 has a corresponding relationship with the reference attribute of the second category. The reference attribute of the first category can be solubility, and the reference attribute of the second category can be stability. Then, the P training groups included in training molecule set 1 include the training molecule features of P training molecules and the training supervision values of the training molecules under solubility, and the P training groups included in training molecule set 2 include the training molecule features of P training molecules and the training supervision values of the training molecules under stability.
[0049] It should be noted that the reference attributes can be obtained by means of screening. The specific screening method can be as follows: obtain the target attributes to be analyzed for the target molecule, screen out at least one candidate attribute associated with the target attributes from the attribute set, and screen out the candidate attributes that meet the screening conditions from the at least one candidate attribute as the reference attributes. Among them, the association relationship between attributes can be specifically set by the R & D personnel in advance. For example, it can be preset in advance that there is an association relationship between absorbability, solubility, and stability. Then, when the target attribute is absorbability, the candidate attributes associated with the target attribute can be solubility and stability. Optionally, the association relationship between attributes can also be determined by the co-occurrence frequency between attributes. Then, the specific method for screening out at least one candidate attribute associated with the target attribute from the attribute set can be as follows: obtain the co-occurrence frequency between each candidate attribute in the attribute set and the target attribute. The co-occurrence frequency includes the frequency of the candidate attribute and the target attribute appearing in the historical attribute set at the same time. The historical attribute set includes the set of attributes analyzed for historical molecules. Determine the candidate attributes whose co-occurrence frequency with the target attribute is greater than the preset frequency as the candidate attributes associated with the target attribute. Among them, historical molecules include the molecules whose attributes have been predicted in the historical records. When testing the attributes of historical molecules, multiple attributes of the historical molecules can be tested simultaneously, and these multiple attributes constitute a historical attribute set. When a candidate attribute and the target attribute appear in a historical attribute set at the same time, it is determined that the candidate attribute and the target attribute co-occur once. In one embodiment, the screening condition can be that the number of training molecules with supervised information under the candidate attribute is greater than the preset number, that is, when the number of training molecules carrying training parameter values under the candidate attribute is greater than the preset number, it is determined that the candidate attribute meets the screening condition. By setting the screening condition, it can be avoided that attributes with insufficient samples are used as reference attributes, affecting subsequent model training.
[0050] S302. Pre-train the initial prediction model with N training molecule sets to obtain a basic prediction model.
[0051] In the embodiment of the present invention, after the terminal obtains N training molecule sets, it can pre-train the initial prediction model with the N training molecule sets. The specific pre-training method can be as follows: train the initial prediction model with each of the N training molecule sets in the N training molecule sets to obtain N intermediate parameters, and update the model parameters of the initial prediction model based on the N intermediate parameters to obtain the initial prediction model with updated model parameters. The terminal determines the initial prediction model with updated model parameters as the basic prediction model.
[0052] In a specific implementation, each intermediate parameter can enable the model to have good prediction ability for the reference attribute corresponding to the intermediate parameter. For example, if the reference attribute corresponding to intermediate parameter 1 is stability, when the model parameters of the initial prediction model are updated to intermediate parameter 1, the initial prediction model at this time has good prediction ability for the stability of the molecule. That is, when the molecular features are input into the initial prediction model with the model parameters updated to intermediate parameter 1, the matching degree between the predicted value of the model output for the stability of the molecule and the actual value of the stability of the molecule is relatively high. In one embodiment, the initial prediction model can be trained separately by each training molecule set. After one training molecule set trains the initial prediction model, an intermediate parameter can be obtained. Then, after N training molecule sets train the initial prediction model, N intermediate parameters can be obtained. The training method for each training molecule set to train the initial prediction model can be the same. In this solution, specifically, the training method of using any one target training molecule set in the N training molecule sets to train the initial prediction model is used as an example to illustrate the training method of training the initial prediction model based on the training molecule set. The specific training method is as Figure 5 shown in the embodiment described above.
[0053] Through Figure 5 the training method shown above, the terminal can obtain the intermediate parameters respectively corresponding to each training molecule set in the N training molecule sets. Further, the terminal updates the model parameters of the initial prediction model based on the N intermediate parameters. The specific update method can be that the terminal performs a summation process on the N intermediate parameters to obtain a target sum value, obtains the weighting coefficient corresponding to the target sum value, and performs a weighting process on the target sum value using the weighting coefficient to obtain a weighted sum value. Then the terminal updates the model parameters of the initial prediction model from the initial model parameters to the difference between the initial model parameters and the weighted sum value.
[0054] For example, the terminal obtains the intermediate parameters {θ1, θ2…θ N} respectively corresponding to each training molecule set in the N training molecule sets, where θ i is the intermediate parameter obtained after training the initial prediction model using the i-th training molecule set in the N training molecule sets, and i ∈ {1, 2…N}. The terminal performs a summation process on the N intermediate parameters to obtain a target sum value The weighting coefficient corresponding to the target sum value is β. Then the terminal updates the model parameters of the initial prediction model from the initial model parameters θ0 to the difference θ between the initial model parameters θ0 and the weighted sum value E . That is, the specific calculation method of θ E is
[0055] S303. Obtain a sample molecule set, where the sample molecule set includes M sample groups composed of sample molecule features and sample parameter values, and the sample parameter values in the sample groups are the training supervision values of the sample molecules under the target attribute.
[0056] In the embodiments of the present invention, the sample molecule set includes M sample groups composed of sample molecule features and sample parameter values, and the sample parameter values in the sample groups are the training supervision values of the sample molecules under the target attribute. The training supervision value under the target attribute can be specifically set in advance by R & D personnel, and it can be the actual value of the sample molecule under the target attribute. For example, if the target attribute is stability, the actual value corresponding to the stability of the training molecule can be obtained through detection, and this actual value is used as the training supervision value of the training molecule under the target attribute.
[0057] S304. Train the basic prediction model with the sample molecule set to obtain the target attribute prediction model.
[0058] In the embodiments of the present invention, after the terminal obtains the sample molecule set, it can train the initial prediction model with the sample molecule set to obtain the target attribute prediction model. Among them, the specific training method can be to screen the target training molecule set to obtain a first sample molecule set and a second sample molecule set. The first sample molecule set includes K sample groups composed of sample molecule features and sample parameter values, and the second sample molecule set includes L sample groups composed of sample molecule features and sample parameter values. The sample parameter values in the sample groups are the training supervision values of the sample molecules under the target attribute; the terminal updates the model parameters in the basic prediction model based on the first sample molecule set to obtain a first basic prediction model, and the model parameters in the first basic prediction model are updated from the initial parameters to the first parameters; the terminal tests the first basic prediction model based on the second sample molecule set to obtain a target test result; if the target test result meets the preset conditions, the terminal determines the first basic prediction model as the target attribute prediction model.
[0059] In a specific implementation, the specific way to update the model parameters in the basic prediction model based on the first sample molecule set may be to calculate the features of each sample molecule in the first sample molecule set through the basic prediction model to obtain K predicted sample values, and update the model parameters in the basic prediction model according to the K predicted sample values and K sample parameter values to obtain the first basic prediction model. Each predicted sample value is the predicted value of the target attribute corresponding to a sample molecule. The K sample parameter values include the sample parameter values in the K training groups included in the first sample molecule set. During the update process, a target loss function can be specifically used to calculate the K predicted sample values and K sample parameter values. After obtaining the sample target loss function value, the model parameters in the basic prediction model are updated based on the gradient value of the sample target loss function value.
[0060] For example, the initial parameter of the basic prediction model is θ E , and by calculating the K predicted sample values and K sample parameter values based on the target loss function, the sample target loss function value can be obtained Furthermore, a preset weighting coefficient α is used to weight the gradient value of the sample target loss function value , and the difference between the initial parameter θ E and the target weighted gradient value is determined as the first parameter θ E1 to achieve the update of the model parameters. The specific calculation method of θ E1 is as follows.
[0061]
[0062] Furthermore, the terminal tests the first basic prediction model based on the second sample molecule set to obtain a target test result. In one embodiment, the features of each sample molecule in the second sample molecule set are calculated through the first basic prediction model to obtain L predicted parameter values, and the target loss function value is obtained by calculating the L predicted parameter values and the L training parameter values in the second sample molecule set based on the target loss function and this target loss function value is used as the target test result. The target loss function can be a cross-entropy loss function, or the target loss function is a mean square error loss function, which can be specifically determined by the problem type corresponding to the target attribute. That is, when the target attribute corresponds to a classification problem (the attribute output value is 0 or 1), the target loss function is a cross-entropy loss function, and when the target attribute corresponds to a regression problem (the attribute output value is a real number), the target loss function is a mean square error loss function.
[0063] If the target test result meets the preset conditions, the first basic prediction model is determined as the target attribute prediction model. If the target test result does not meet the preset conditions, the data can be re-screened from the sample molecule bindings in the manner described in steps S303 - S304 to train the basic prediction model until the test result obtained from the training meets the prediction conditions, and then the basic prediction model at this time is determined as the target attribute prediction model.
[0064] As Figure 4 shown, it is a schematic flowchart of a process for training a model provided by this solution. Specifically, N training molecule sets include training molecule set 1, training molecule set 2... training molecule set N. By pre-training the initial prediction model with the N training molecule sets respectively, N intermediate parameters θ1, θ2... θ N can be obtained. Based on the N intermediate parameters, the model parameter θ0 of the initial prediction model is updated, and then the model parameter θ E of the basic prediction model can be obtained. By training θ E with the sample molecule set, the model parameter θ T can be obtained, and the model parameter of the basic prediction model is updated to θ T , and then the target attribute prediction model is obtained.
[0065] In the embodiments of the present invention, the target attribute prediction model is used to predict the target attribute of the molecule, realizing the prediction of the molecule attribute based on the artificial intelligence method, and improving the efficiency of the molecule attribute prediction. Moreover, by pre-training the initial prediction model with multiple groups of reference attribute molecule data with supervision information (N training molecule sets) to obtain the basic prediction model, and by training and optimizing the model with the target attribute molecule data with supervision information (sample molecule set) to obtain the target attribute prediction model, the problem of less molecule data with supervision information for the target attribute in the model training process is also solved, and the accuracy of the model for molecule attribute analysis obtained by training with a small amount of sample data is improved.
[0066] Based on the above description, the embodiments of the present invention provide a model pre-training method. Please refer to Figure 5 , and the model pre-training process may include the following steps S501 - S508:
[0067] S501. Screen the target training molecule set based on the first screening method to obtain the first training molecule set and the second training molecule set.
[0068] In an embodiment of the present invention, for any target training molecule set among N training molecule sets, the target training molecule set is screened based on a first screening method to obtain a first training molecule set and a second training molecule set. Among them, the first training molecule set includes K training groups composed of training molecule features and training parameter values, and the second training molecule set includes L training groups composed of training molecule features and training parameter values. Each training parameter value is the training supervision value of the training molecule in the training group under the reference attribute. K and L are positive integers less than P. The above first training molecule set is used to train the initial prediction model, and the second training molecule set is used to test the trained initial prediction model.
[0069] In one implementation, the first screening method is a random screening method, that is, randomly select K training groups composed of training molecule features and training parameter values from the target training molecule set as the first training molecule set, and then randomly select L training groups composed of training molecule features and training parameter values from the target training molecule set as the second training molecule set. Optionally, L is the difference between P and K, that is, the second training molecule set is the set constructed by the remaining training groups in the target training molecule set after the first training molecule set is screened out from the target training molecule set. For example, if the target training molecule set includes 15 (P) training groups, then randomly select 5 (K) training groups from the 15 training groups as the first training molecule set, and use the remaining 10 (L) training groups as the second training molecule set.
[0070] In one implementation, the first screening method can also be a regular screening method. For example, the first P training groups in the target training molecule set are used as the first training molecule set, and the last L training groups are used as the second training molecule set.
[0071] S502. Update the model parameters in the initial prediction model based on the first training molecule set to obtain a first initial prediction model.
[0072] In an embodiment of the present invention, after screening the target training molecule set to obtain the first training molecule set, the model parameters in the initial prediction model can be updated based on the first training molecule set to obtain a first initial prediction model. The model parameters in the first initial prediction model are updated from the initial model parameters to the first model parameters.
[0073] In a specific implementation, the specific manner of updating the model parameters in the initial prediction model based on the first training molecule set may be to perform operations on the training molecule features in the first training molecule set through the initial prediction model to obtain K predicted parameter values, and update the model parameters in the initial prediction model according to the K predicted parameter values and K training parameter values to obtain the first initial prediction model. Each predicted parameter value is the predicted value of the reference attribute corresponding to a training molecule, and the K training parameter values include the training parameter values in the K training groups included in the first training molecule set.
[0074] Specifically, the manner of updating the model parameters in the initial prediction model according to the K predicted parameter values and K training parameter values to obtain the first initial prediction model may specifically be to obtain the reference attributes corresponding to the training molecules in the first training molecule set, determine the target loss function corresponding to the reference attributes based on the correspondence between the attributes and the loss function, perform operations on the K predicted parameter values and K training parameter values based on the target loss function to obtain the target loss function value, and update the model parameters in the initial prediction model according to the target loss function value to obtain the first initial prediction model. The target loss function includes a cross-entropy loss function or a mean squared error loss function.
[0075] It should be noted that the specific manner of updating the model parameters in the initial prediction model according to the target loss function value to obtain the first initial prediction model may be to perform a gradient operation on the target loss function value to obtain the gradient value corresponding to the target loss function value, and update the model parameters of the initial prediction model based on the gradient value to obtain the first initial prediction model. The model parameters in the first initial prediction model are updated from the initial model parameters to the first model parameters.
[0076] For example, the initial model parameters of the initial prediction model are θ, and the K training molecule features in the first training molecule set include {x1, x2... x K}, then the K predicted parameter values obtained by performing operations on each training molecule feature through the initial prediction model include {f θ (x1), f θ (x2)... f θ (x K )}, and the training feature values corresponding to the respective training molecules are {y1, y2... y K}.
[0077] When the target loss function is a cross-entropy loss function, the specific manner of performing operations on the K predicted parameter values and K training parameter values based on the target loss function is:
[0078]
[0079] When the target loss function is the mean squared error loss function, the target loss function value is obtained by operating on the K predicted parameter values and the K training parameter values based on the target loss function. The specific method is as follows:
[0080]
[0081] where s ∈ {1, 2... K}. Through the above method, the target loss function value obtained by training the model based on the first training molecule set can be calculated. Furthermore, the terminal performs a gradient operation on the target loss function value to obtain the gradient value corresponding to the target loss function value. Then, the specific method for updating the model parameters of the initial prediction model based on the gradient value can be to perform a weighted processing on the gradient value using a preset weighting coefficient α to obtain a weighted gradient value. And the difference between the initial model parameter θ and the weighted gradient value is determined as the first model parameter θ k , so as to realize the update of the model parameters. The specific calculation method of θ k is as follows.
[0082]
[0083] It should also be noted that different reference attributes correspond to different loss functions. When the task type corresponding to the reference attribute is a classification task (the task output value is 0 or 1), the cross-entropy loss function is used as the target loss function. When the task type corresponding to the reference attribute is a regression task (the task output value is a real number), the mean squared error loss function is used as the target loss function.
[0084] S503. Test the first initial prediction model based on the second training molecule set to obtain a first test result.
[0085] In the embodiment of the present invention, after updating the parameters of the initial prediction model to the first model parameter to obtain the first initial prediction model, the first initial prediction model can be tested based on the second training molecule set to obtain a first test result. The second training molecule set includes L training groups composed of training molecule features and training parameter values.
[0086] In specific implementation, the specific method for testing the first initial prediction model based on the second training molecule set can be to operate on each training molecule feature in the second training molecule set through the first initial prediction model to obtain L predicted parameter values, operate on the L predicted parameter values and the L training parameter values in the second training molecule set based on the target loss function to obtain a first target loss function value, and use the first target loss function value as the first test result.
[0087] For example, the first initial model parameter of the first initial prediction model is θ k , and the training molecule features of K training molecules in the obtained first training molecule set include {x1, x2... x L}, and the corresponding training feature values of each training molecule are {y1, y2... y L}, and the K prediction parameter values obtained by operating on each training molecule feature through the initial prediction model include
[0088] When the target loss function is the cross-entropy loss function, based on the target loss function, L prediction parameter values and L training parameter values are operated to obtain the first target loss function value The specific method is as follows:
[0089]
[0090] When the target loss function is the mean square error loss function, based on the target loss function, L prediction parameter values and L training parameter values are operated to obtain the first target loss function value The specific method is as follows:
[0091]
[0092] where t ∈ {1, 2... L}, and through the above method, the first target loss function value obtained by training the model based on the second training molecule set can be calculated
[0093] S504. If the first test result meets the preset condition, determine the target intermediate parameter based on the first test result and the initial model parameter.
[0094] In the embodiment of the present invention, after obtaining the first test result, it is checked whether the first test result meets the preset condition. If it meets the preset condition, the target intermediate parameter is determined based on the first test result and the initial model parameter. In specific implementation, the terminal obtains the first target loss function value and uses the first target loss function value to perform gradient operation on the initial model parameter θ to obtain a target intermediate parameter θ i , then the specific calculation method of the target intermediate parameter θ i is as follows:
[0095]
[0096] where θ iDuring the calculation process, second-order gradient information about θ will be generated. Since the size of the second-order gradient information is the square of the model dimension, calculating this second-order gradient information will consume a large amount of computing resources and time costs. In an alternative implementation, to improve the calculation efficiency, the second-order gradient information can be reduced to first-order gradient information for operation. That is, the terminal determines a target intermediate parameter based on the first test result and the first model parameter. Specifically, the terminal uses the first target loss function value to perform gradient operation on the first model parameter to obtain a target intermediate parameter θ i , then the specific calculation method of the target intermediate parameter θ i is as follows:
[0097]
[0098] In the above method, since near the optimal value of θ, its second-order information is usually close to 0, this method will not cause a decrease in accuracy, and at the same time, it improves the training efficiency of the algorithm. Optionally, if the first test result does not meet the preset condition, then step S505 is executed. Among them, the preset condition can be greater than a preset threshold. Then, when the first test result is greater than the preset threshold, it is determined that the preset condition is met. When the first test result is less than or equal to the preset threshold, it is determined that the preset condition is not met.
[0099] S505. If the first test result does not meet the preset condition, then screen the target training molecule set based on the second screening method to obtain a third training molecule set and a fourth training molecule set.
[0100] In the embodiment of the present invention, after obtaining the first test result, if the first test result does not meet the preset condition, then screen the target training molecule set based on the second screening method to obtain a third training molecule set and a fourth training molecule set. Among them, the third training molecule set includes U training groups composed of training molecule features and training parameter values, and the fourth training molecule set includes V training groups composed of training molecule features and training parameter values. Each training parameter value is the training supervision value of the training molecule in the training group under the reference attribute. U and V are positive integers less than P. In an alternative implementation, the number of training groups in the third training molecule set can be the same as that of the first training molecule set, and the number of training groups in the fourth training molecule set can be the same as that of the second training molecule set, that is, U = K, V = L.
[0101] Among them, the second screening method can be a random screening method, that is, randomly select U training groups composed of training molecule features and training parameter values from the target training molecule set as the third training molecule set, and then randomly select V training groups composed of training molecule features and training parameter values from the target training molecule set as the fourth training molecule set. Or, the second screening method can also be a rule-based screening method, that is, based on a preset rule, select U training groups from the target training molecule set as the third training molecule set, and based on the preset rule, randomly select V training groups from the target training molecule set as the fourth training molecule set.
[0102] S506. Update the model parameters in the first initial prediction model based on the third training molecule set to obtain a second initial prediction model.
[0103] In the embodiments of the present invention, after obtaining the third training molecule set, the model parameters in the first initial prediction model can be updated based on the third training molecule set to obtain a second initial prediction model, and the model parameters in the second initial prediction model are updated from the first model parameters to the second model parameters.
[0104] For example, the initial model parameters of the first initial prediction model are θ k , and by operating on the U predicted parameter values and the U training parameter values based on the target loss function, the first target loss function value can be obtained. Further, the difference between the first initial model parameter θ k and the first weighted gradient value is determined as the second model parameter θ k1 to achieve the update of the model parameters. The specific calculation method of θ k is as follows.
[0105]
[0106] S507. Test the second initial prediction model based on the fourth training molecule set to obtain a second test result.
[0107] In the embodiments of the present invention, the specific method for testing the second initial prediction model based on the fourth training molecule set can be to operate on each training molecule feature in the fourth training molecule set through the second initial prediction model to obtain V predicted parameter values, operate on the V predicted parameter values and the V training parameter values in the fourth training molecule set based on the target loss function to obtain a second target loss function value, and use this second target loss function value as the second test result.
[0108] S508. If the second test result meets the preset conditions, determine the target intermediate parameter based on the second test result and the first model parameter.
[0109] In an embodiment of the present invention, after obtaining the second test result, it is checked whether the second test result meets a preset condition. If the preset condition is met, a target intermediate parameter is determined based on the second test result and the first model parameter, that is, the first model parameter θ is subjected to gradient calculation using the second target loss function value to obtain a target intermediate parameter. If the second test result does not meet the preset condition, the method described in steps S501 - S507 can be used to re - screen data to train the prediction model until the test result obtained from the training meets the prediction condition, and then the target intermediate parameter is determined based on the test result and the model parameter at this time. k In an embodiment of the present invention, by dividing the training molecule set into a training set and a test set to train the initial prediction model, a target intermediate parameter for updating the model parameter of the initial prediction model can be obtained, so as to subsequently update the model parameter of the initial prediction model based on each intermediate parameter and complete the pre - training of the initial prediction model.
[0110] In an embodiment of the present invention, by dividing the training molecule set into a training set and a test set to train the initial prediction model, a target intermediate parameter for updating the model parameter of the initial prediction model can be obtained, so as to subsequently update the model parameter of the initial prediction model based on each intermediate parameter and complete the pre - training of the initial prediction model.
[0111] In an implementation scenario, this technical solution can be used for predicting molecular properties in new drugs. Specifically, it can be applied to the property prediction module in a drug platform. Through the above - mentioned molecular prediction method, the properties of the drug can be obtained, and the prediction of drug properties based on artificial intelligence can be realized, accelerating the discovery and optimization process of lead compounds. Specifically, for a target molecule for which target property prediction is required, a target property prediction model for target property prediction is obtained, and the molecular characteristics are calculated through the target property prediction model to obtain the parameter value of the target property corresponding to the target molecule, thereby realizing the prediction of the target property of the target molecule. As Figure 6 shown, it is a schematic diagram of a molecular property prediction system architecture provided by an embodiment of the present invention. Figure 6 In it, the intelligent device for predicting the properties of molecules can specifically be a server 602. The user can input a target molecule to be predicted for properties and the corresponding target property in a terminal 601. The terminal 601 uploads the target molecule and target property input by the user to the server 602. After the server 602 obtains the target molecule and target property, it acquires the molecular characteristics of the target molecule, and calculates the molecular characteristics through the target property prediction model to obtain the parameter value of the target property corresponding to the target molecule. Further, the server 602 sends the parameter value of the target property to the terminal 601 so that the terminal 601 can display the parameter value of the target property. Optionally, the user can screen out multiple target properties for which the input target molecule needs to be predicted for properties from multiple property options provided by the terminal. The server respectively calls the corresponding target property prediction models to predict the target molecule and sends the property prediction results to the terminal for display. For example, asFigure 7 As shown, after the user inputs the target molecule (C2H8O2), the user can select the stability and toxicity options on the interactive page provided in the terminal. The parameter value obtained by the server processing the target molecule using the stability prediction model is 59, and the parameter value obtained by processing the target molecule using the toxicity prediction model is 2. These values are sent to Figure 7 the corresponding display box in
[0112] Based on the description of the above embodiments of the molecular property prediction method, an embodiment of the present invention also discloses a molecular property prediction device. The molecular property prediction device can be a computer program (including program code) running in a computer device, or an entity device included in the computer device. The molecular property prediction device can execute Figure 1 the method shown. Please refer to Figure 8 , the molecular property prediction device 80 includes: an acquisition module 801 and a processing module 802.
[0113] The acquisition module 801 is configured to acquire the molecular characteristics of the target molecule to be analyzed for the target property, where the target molecule refers to the molecule for which the target property parameter value is analyzed under the target property;
[0114] The processing module 802 is configured to input the molecular characteristics into the target property prediction model associated with the target property, and analyze the molecular characteristics through the target property prediction model to obtain the parameter value of the target property corresponding to the target molecule;
[0115] Among them, the target property prediction model is obtained by training a basic prediction model with a sample molecule set. The sample molecule set includes M sample groups composed of sample molecule characteristics and sample parameter values. The sample parameter value in the sample group is the training supervision value of the sample molecule under the target property;
[0116] The basic prediction model is pre-trained by an initial prediction model with N training molecule sets. Each training molecule set includes P training groups composed of training molecule characteristics and training parameter values. The training parameter value in the training group is the training supervision value of the training molecule under the reference property. M, N, and P are positive integers.
[0117] In an embodiment of the present invention, an acquisition module 801 acquires molecular features of a target molecule to be subjected to target attribute analysis, and a processing module 802 inputs the molecular features into a target attribute prediction model associated with the target attribute, analyzes the molecular features through the target attribute prediction model, and obtains a parameter value of the target attribute corresponding to the target molecule. By using the target attribute prediction model to predict the target attribute of the molecule, the prediction of the molecular attribute is realized in an artificial intelligence-based manner, the efficiency of the molecular attribute prediction is improved, and the attribute prediction model is trained by a combination of pre-training and training, reducing the requirement for sample molecules with supervised information and improving the accuracy of the model trained based on a small amount of sample data for molecular attribute analysis.
[0118] Based on the description of the foregoing embodiments of the prediction model training method, an embodiment of the present invention further discloses a prediction model training apparatus. The prediction model training apparatus may be a computer program (including program code) running on a computer device or an entity device included in the computer device. The molecular attribute prediction apparatus may execute Figures 2 - 5 the method shown in Figure 9 . Referring to
[0119] The acquisition module 901 is configured to acquire N training molecule sets, where each training molecule set includes P training groups composed of training molecular features and training parameter values, and the training parameter values in the training groups are training supervision values of the training molecules in the training molecule set under a reference attribute;
[0120] The training module 902 is configured to pre-train an initial prediction model through the N training molecule sets to obtain a basic prediction model;
[0121] The acquisition module 901 is further configured to acquire a sample molecule set, where the sample molecule set includes M sample groups composed of sample molecular features and sample parameter values, and the sample parameter values in the sample groups are training supervision values of the sample molecules under the target attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute;
[0122] The training module 902 is further configured to train the basic prediction model through the sample molecule set to obtain a target attribute prediction model.
[0123] In one implementation, the training module 902 is specifically configured to:
[0124] Train the initial prediction model through each of the N training molecule sets to obtain N intermediate parameters;
[0125] Updating the model parameters of the initial prediction model based on the N intermediate parameters to obtain the initial prediction model with updated model parameters;
[0126] Determining the initial prediction model with updated model parameters as the basic prediction model.
[0127] In one implementation, the training module 902 is specifically configured to:
[0128] Screening the target training molecule set based on the first screening method to obtain a first training molecule set and a second training molecule set. The first training molecule set includes K training groups composed of training molecule features and training parameter values, and the second training molecule set includes L training groups composed of training molecule features and training parameter values. Each training parameter value is the training supervision value of the training molecule under the reference attribute, and K and L are positive integers less than P;
[0129] Updating the model parameters in the initial prediction model based on the first training molecule set to obtain a first initial prediction model, where the model parameters in the first initial prediction model are updated from the initial model parameters to the first model parameters;
[0130] Testing the first initial prediction model based on the second training molecule set to obtain a first test result;
[0131] If the first test result meets the preset conditions, determining the target intermediate parameters based on the first test result and the initial model parameters.
[0132] In one implementation, the training module 902 is specifically configured to:
[0133] Performing operations on each training molecule feature in the first training molecule set through the initial prediction model to obtain K predicted parameter values, where each predicted parameter value is the predicted value of the reference attribute corresponding to a training molecule;
[0134] Updating the model parameters in the initial prediction model according to the K predicted parameter values and the K training parameter values to obtain a first initial prediction model, and the K training parameter values include the training parameter values in the K training groups included in the first training molecule set.
[0135] In one implementation, the training module 902 is specifically configured to:
[0136] Obtaining the reference attribute corresponding to each training molecule in the first training molecule set, and determining the target loss function corresponding to the reference attribute based on the correspondence between the attribute and the loss function. The target loss function includes a cross-entropy loss function or a mean squared error loss function;
[0137] Operate on the K predicted parameter values and K training parameter values based on the target loss function to obtain a target loss function value;
[0138] Update the model parameters in the initial prediction model according to the target loss function value to obtain a first initial prediction model.
[0139] In one implementation, the training module 902 is specifically configured to:
[0140] If the first test result does not meet the preset condition, screen the target training molecule set based on a second screening method to obtain a third training molecule set and a fourth training molecule set;
[0141] Update the model parameters in the first initial prediction model based on the third training molecule set to obtain a second initial prediction model, and the model parameters in the second initial prediction model are updated from the first model parameters to the second model parameters;
[0142] Test the second initial prediction model based on the fourth training molecule set to obtain a second test result;
[0143] If the second test result meets the preset condition, determine the target intermediate parameter based on the second test result and the first model parameters.
[0144] In one implementation, the training module 902 is specifically configured to:
[0145] Perform a summation process on the N intermediate parameters to obtain a target sum value;
[0146] Obtain the weighting coefficient corresponding to the target sum value, and perform a weighting process on the target sum value using the weighting coefficient to obtain a weighted sum value;
[0147] Update the model parameters of the initial prediction model from the initial model parameters to the difference between the initial model parameters and the weighted sum value.
[0148] In one implementation, the training module 902 is specifically configured to:
[0149] Obtain the target attribute for which the target molecule is to be analyzed for attributes;
[0150] Screen at least one candidate attribute associated with the target attribute from the attribute set;
[0151] Screen the candidate attributes that meet the screening conditions from the at least one candidate attribute as the reference attribute.
[0152] In one implementation, the training module 902 is specifically configured to:
[0153] Screen the target training molecule set to obtain a first sample molecule set and a second sample molecule set. The first sample molecule set includes K sample groups composed of sample molecule features and sample parameter values, and the second sample molecule set includes L sample groups composed of sample molecule features and sample parameter values. The sample parameter value in the sample group is the training supervision value of the sample molecule under the target attribute.
[0154] Update the model parameters in the basic prediction model based on the first sample molecule set to obtain a first basic prediction model. The model parameters in the first basic prediction model are updated from the initial parameters to the first parameters.
[0155] Test the first basic prediction model based on the second sample molecule set to obtain a target test result.
[0156] If the target test result meets the preset conditions, determine the first basic prediction model as the target attribute prediction model.
[0157] In one implementation, the training module 902 is specifically configured to:
[0158] Obtain the molecular features of the target molecule to be analyzed for the target attribute. The target molecule refers to the molecule for which the target attribute parameter value needs to be analyzed under the target attribute.
[0159] Input the molecular features into the target attribute prediction model associated with the target attribute, and analyze the molecular features through the target attribute prediction model to obtain the parameter value of the target attribute corresponding to the target molecule.
[0160] In the embodiment of the present invention, the acquisition module 901 acquires N training molecule sets, and the training module 902 pre-trains the initial prediction model through the N training molecule sets to obtain a basic prediction model; the acquisition module 901 acquires a sample molecule set, and the training module 902 trains the basic prediction model through the sample molecule set to obtain a target attribute prediction model. By training the attribute prediction model in a combined manner of pre-training and training, the demand for sample molecules with supervised information is reduced, and the accuracy of the model trained based on a small amount of sample data for molecular attribute analysis is improved.
[0161] Please refer to Figure 10 , which is a schematic structural diagram of an intelligent device provided by an embodiment of the present invention. As Figure 10As shown, the intelligent device includes: at least one processor 1001, an input device 1003, an output device 1004, a memory 1005, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the memory 1005 can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1005 can also be at least one storage device located far from the aforementioned processor 1001. Among them, the processor 1001 can be combined with Figure 8 the device described, a set of program codes are stored in the memory 1005, and the processor 1001, the input device 1003, and the output device 1004 call the program codes stored in the memory 1005 to perform the following operations:
[0162] The processor 1001 is used to obtain the molecular characteristics of the target molecule to be subjected to target attribute analysis, where the target molecule refers to the molecule for which the target attribute parameter value needs to be analyzed under the target attribute;
[0163] The processor 1001 is used to input the molecular characteristics into the target attribute prediction model associated with the target attribute, and analyze the molecular characteristics through the target attribute prediction model to obtain the parameter value of the target attribute corresponding to the target molecule;
[0164] Among them, the target attribute prediction model is obtained by training a basic prediction model with a sample molecule set, and the sample molecule set includes: M sample groups composed of sample molecular characteristics and sample parameter values, and the sample parameter values in the sample group are the training supervision values of the sample molecules under the target attribute;
[0165] The basic prediction model is pre-trained with an initial prediction model through N training molecule sets. Each training molecule set includes: P training groups composed of training molecular characteristics and training parameter values, and the training parameter values in the training group are the training supervision values of the training molecules under the reference attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute.
[0166] In an embodiment of the present invention, the processor 1001 obtains the molecular characteristics of a target molecule to be subjected to target attribute analysis. The processor 1001 inputs the molecular characteristics into a target attribute prediction model associated with the target attribute, and analyzes the molecular characteristics through the target attribute prediction model to obtain the parameter value of the target attribute corresponding to the target molecule. By using the target attribute prediction model to predict the target attribute of the molecule, the prediction of the molecular attribute is realized in an artificial intelligence-based manner, which improves the efficiency of the molecular attribute prediction. The attribute prediction model is trained by a combination of pre-training and training, which reduces the requirement for sample molecules with supervised information and improves the accuracy of the model trained based on a small amount of sample data for molecular attribute analysis.
[0167] Please refer to Figure 11 , which is a schematic structural diagram of a terminal provided by an embodiment of the present invention. As Figure 11 shown, the terminal includes: at least one processor 1111, an input device 1113, an output device 1114, a memory 1115, and at least one communication bus 1112. Among them, the communication bus 1112 is used to realize the connection and communication between these components. Among them, the memory 1105 can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Optionally, the memory 1105 can also be at least one storage device located far from the aforementioned processor 1101. The processor 1101 can be combined with Figure 8 the device described, a set of program codes are stored in the memory 1105, and the processor 1101, the input device 1103, and the output device 1104 call the program codes stored in the memory 1105 to perform the following operations:
[0168] The processor 1101 is used to obtain N training molecule sets, and each training molecule set includes: P training groups composed of training molecule characteristics and training parameter values, and the training parameter values in the training groups are the training supervision values of the training molecules under the reference attribute;
[0169] The processor 1101 is used to pre-train an initial prediction model through the N training molecule sets to obtain a basic prediction model;
[0170] The processor 1101 is used to obtain a sample molecule set, and the sample molecule set includes: M sample groups composed of sample molecule characteristics and sample parameter values, and the sample parameter values in the sample groups are the training supervision values of the sample molecules under the target attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute;
[0171] A processor 1101 for training a basic prediction model with the set of sample molecules to obtain a target attribute prediction model.
[0172] In one implementation, the processor 1101 is specifically configured to:
[0173] Train the initial prediction model with each training molecule set in the N training molecule sets to obtain N intermediate parameters;
[0174] Update the model parameters of the initial prediction model based on the N intermediate parameters to obtain the initial prediction model with updated model parameters;
[0175] Determine the initial prediction model with updated model parameters as the basic prediction model.
[0176] In one implementation, the processor 1101 is specifically configured to:
[0177] Screen the target training molecule set based on the first screening method to obtain a first training molecule set and a second training molecule set. The first training molecule set includes K training groups composed of training molecule features and training parameter values, and the second training molecule set includes L training groups composed of training molecule features and training parameter values. Each training parameter value is the training supervision value of the training molecule under the reference attribute. K and L are positive integers less than P;
[0178] Update the model parameters in the initial prediction model based on the first training molecule set to obtain a first initial prediction model, and the model parameters in the first initial prediction model are updated from the initial model parameters to the first model parameters;
[0179] Test the first initial prediction model based on the second training molecule set to obtain a first test result;
[0180] If the first test result meets the preset condition, determine the target intermediate parameter based on the first test result and the initial model parameters.
[0181] In one implementation, the processor 1101 is specifically configured to:
[0182] Perform operations on each training molecule feature in the first training molecule set through the initial prediction model to obtain K predicted parameter values, and each predicted parameter value is the predicted value of the reference attribute corresponding to a training molecule;
[0183] Update the model parameters in the initial prediction model according to the K predicted parameter values and the K training parameter values to obtain a first initial prediction model, and the K training parameter values include the training parameter values in the K training groups included in the first training molecule set.
[0184] In one implementation, the processor 1101 is specifically configured to:
[0185] Obtain the reference attributes corresponding to each training molecule in the first training molecule set, and determine the target loss function corresponding to the reference attributes based on the correspondence between the attributes and the loss function, where the target loss function includes a cross-entropy loss function or a mean squared error loss function;
[0186] Perform an operation on the K prediction parameter values and the K training parameter values based on the target loss function to obtain a target loss function value;
[0187] Update the model parameters in the initial prediction model according to the target loss function value to obtain a first initial prediction model.
[0188] In one implementation, the processor 1101 is specifically configured to:
[0189] If the first test result does not meet the preset condition, then screen the target training molecule set based on a second screening method to obtain a third training molecule set and a fourth training molecule set;
[0190] Update the model parameters in the first initial prediction model based on the third training molecule set to obtain a second initial prediction model, where the model parameters in the second initial prediction model are updated from the first model parameters to the second model parameters;
[0191] Test the second initial prediction model based on the fourth training molecule set to obtain a second test result;
[0192] If the second test result meets the preset condition, then determine the target intermediate parameter based on the second test result and the first model parameters.
[0193] In one implementation, the processor 1101 is specifically configured to:
[0194] Perform a summation process on the N intermediate parameters to obtain a target sum value;
[0195] Obtain the weighting coefficient corresponding to the target sum value, and perform a weighting process on the target sum value using the weighting coefficient to obtain a weighted sum value;
[0196] Update the model parameters of the initial prediction model from the initial model parameters to the difference between the initial model parameters and the weighted sum value.
[0197] In one implementation, the processor 1101 is specifically configured to:
[0198] Obtain the target attribute for which attribute analysis is to be performed on the target molecule;
[0199] Filter out at least one candidate attribute associated with the target attribute from the set of attributes.
[0200] Filter out the candidate attributes that meet the filtering conditions from the at least one candidate attribute as reference attributes.
[0201] In one implementation, the processor 1101 is specifically configured to:
[0202] Filter the target training molecule set to obtain a first sample molecule set and a second sample molecule set. The first sample molecule set includes K sample groups composed of sample molecule features and sample parameter values, and the second sample molecule set includes L sample groups composed of sample molecule features and sample parameter values. The sample parameter value in the sample group is the training supervision value of the sample molecule under the target attribute.
[0203] Update the model parameters in the basic prediction model based on the first sample molecule set to obtain a first basic prediction model. The model parameters in the first basic prediction model are updated from the initial parameters to the first parameters.
[0204] Test the first basic prediction model based on the second sample molecule set to obtain a target test result.
[0205] If the target test result meets the preset conditions, determine the first basic prediction model as the target attribute prediction model.
[0206] In one implementation, the processor 1101 is specifically configured to:
[0207] Obtain the molecular features of the target molecule to be analyzed for the target attribute. The target molecule refers to the molecule for which the target attribute parameter value needs to be analyzed under the target attribute.
[0208] Input the molecular features into the target attribute prediction model associated with the target attribute, and analyze the molecular features through the target attribute prediction model to obtain the parameter value of the target attribute corresponding to the target molecule.
[0209] In the embodiments of the present invention, the processor 1101 obtains N training molecule sets, and pre-trains the initial prediction model through the N training molecule sets to obtain a basic prediction model; the processor 1101 obtains a sample molecule set, and trains the basic prediction model through the sample molecule set to obtain a target attribute prediction model. By training the attribute prediction model in a combination of pre-training and training, the requirement for sample molecules with supervised information is reduced, and the accuracy of the model trained based on a small amount of sample data for molecular attribute analysis is improved.
[0210] In the embodiments of the present invention, the module can be implemented by a general integrated circuit, such as a CPU (Central Processing Unit), or by an ASIC (Application Specific Integrated Circuit).
[0211] It should be understood that in the embodiments of the present invention, the so-called processor may be a central processing module (Central Processing Unit, CPU), and this processor may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), field-programmable gate arrays (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0212] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 and Figure 11 only a thick line is used to represent it in [description], but it does not mean that there is only one bus or one type of bus.
[0213] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The described program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (Read-Only Memory, ROM), or a random access memory (Random Access Memory, RAM), etc.
[0214] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A method for predicting molecular properties, characterized in that, The method includes: Obtaining the molecular characteristics of a target molecule to be subjected to target attribute analysis, where the target molecule refers to a molecule for which the target attribute parameter value needs to be analyzed under the target attribute; Inputting the molecular characteristics into a target attribute prediction model associated with the target attribute, and analyzing the molecular characteristics through the target attribute prediction model to obtain the parameter value of the target attribute corresponding to the target molecule; Among them, the target attribute prediction model is obtained by training a basic prediction model with a sample molecule set, and the sample molecule set includes: M sample groups composed of sample molecular characteristics and sample parameter values, and the sample parameter values in the sample group are the training supervision values of the sample molecules under the target attribute; The basic prediction model is pre-trained with N training molecule sets for an initial prediction model. Each training molecule set includes: P training groups composed of training molecular characteristics and training parameter values, and the training parameter values in the training group are the training supervision values of the training molecules under a reference attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute; among them, the reference attribute includes multiple candidate attributes that meet the screening conditions selected from candidate attributes, and the candidate attributes are selected from an attribute set and have an association relationship with the target attribute; Among them, screening candidate attributes with an association relationship with the target attribute from the attribute set includes: obtaining the co-occurrence frequency between each candidate attribute in the attribute set and the target attribute. The co-occurrence frequency includes the frequency of the candidate attribute and the target attribute simultaneously appearing in the historical attribute set. The historical attribute set includes the set of attributes for which historical molecules are subjected to attribute analysis. The candidate attributes with a co-occurrence frequency greater than the preset frequency with the target attribute are determined as candidate attributes having an association relationship with the target attribute.
2. A method for training a prediction model, characterized in that, The method includes: Obtaining N training molecule sets and obtaining the target attribute for which the target molecule is to be subjected to attribute analysis. Each training molecule set includes: P training groups composed of training molecular characteristics and training parameter values, and the training parameter values in the training group are the training supervision values of the training molecules under a reference attribute; among them, the reference attribute includes multiple candidate attributes that meet the screening conditions selected from candidate attributes, and the candidate attributes are selected from an attribute set and have an association relationship with the target attribute; Pre-training the initial prediction model with the N training molecule sets to obtain a basic prediction model; Obtaining a sample molecule set, where the sample molecule set includes: M sample groups composed of sample molecular characteristics and sample parameter values, and the sample parameter values in the sample group are the training supervision values of the sample molecules under the target attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute; Training the basic prediction model with the sample molecule set to obtain a target attribute prediction model; Among them, screening candidate attributes associated with the target attribute from the attribute set includes: obtaining the co-occurrence frequency between each candidate attribute in the attribute set and the target attribute, where the co-occurrence frequency includes the frequency of the candidate attribute and the target attribute simultaneously appearing in the historical attribute set, and the historical attribute set includes the set of attributes obtained by analyzing the attributes of historical molecules. The candidate attributes with a co-occurrence frequency greater than the preset frequency with the target attribute are determined as the candidate attributes associated with the target attribute.
3. The method according to claim 2, wherein The pre-training of the initial prediction model with the N training molecule sets to obtain the basic prediction model includes: Training the initial prediction model with each training molecule set in the N training molecule sets to obtain N intermediate parameters; Updating the model parameters of the initial prediction model based on the N intermediate parameters to obtain the initial prediction model with updated model parameters; Determining the initial prediction model with updated model parameters as the basic prediction model.
4. The method according to claim 3, wherein The method of training the initial prediction model with any one of the N training molecule sets to obtain a target intermediate parameter includes: Screening the target training molecule set based on the first screening method to obtain a first training molecule set and a second training molecule set. The first training molecule set includes K training groups composed of training molecule features and training parameter values, and the second training molecule set includes L training groups composed of training molecule features and training parameter values. Each training parameter value is the training supervision value of the training molecule under the reference attribute, and K and L are positive integers less than P; Updating the model parameters in the initial prediction model based on the first training molecule set to obtain a first initial prediction model, where the model parameters in the first initial prediction model are updated from the initial model parameters to the first model parameters; Testing the first initial prediction model based on the second training molecule set to obtain a first test result; If the first test result meets the preset conditions, determining the target intermediate parameter based on the first test result and the initial model parameters.
5. The method according to claim 4, wherein The updating the model parameters in the initial prediction model based on the first training molecule set to obtain a first initial prediction model includes: Calculating, through the initial prediction model, each training molecule feature in the first training molecule set to obtain K predicted parameter values, where each predicted parameter value is the predicted value of the reference attribute corresponding to a training molecule; Updating the model parameters in the initial prediction model according to the K predicted parameter values and the K training parameter values to obtain a first initial prediction model, where the K training parameter values include the training parameter values in the K training groups included in the first training molecule set.
6. The method according to claim 5, wherein The updating the model parameters in the initial prediction model according to the K predicted parameter values and the K training parameter values to obtain a first initial prediction model includes: Obtain the reference attributes corresponding to each training molecule in the first training molecule set, and determine the target loss function corresponding to the reference attributes based on the correspondence between the attributes and the loss function. The target loss function includes a cross-entropy loss function or a mean squared error loss function; Perform an operation on the K predicted parameter values and the K training parameter values based on the target loss function to obtain a target loss function value; Update the model parameters in the initial prediction model according to the target loss function value to obtain a first initial prediction model.
7. The method according to claim 4, wherein After testing the first initial prediction model based on the second training molecule set to obtain a first test result, the method further includes: If the first test result does not meet the preset condition, then screen the target training molecule set based on a second screening method to obtain a third training molecule set and a fourth training molecule set; Update the model parameters in the first initial prediction model based on the third training molecule set to obtain a second initial prediction model, and the model parameters in the second initial prediction model are updated from the first model parameters to the second model parameters; Test the second initial prediction model based on the fourth training molecule set to obtain a second test result; If the second test result meets the preset condition, then determine the target intermediate parameter based on the second test result and the first model parameters.
8. The method according to claim 3, wherein The updating the model parameters of the initial prediction model based on the N intermediate parameters includes: Perform a summation process on the N intermediate parameters to obtain a target sum value; Obtain the weighting coefficient corresponding to the target sum value, and perform a weighting process on the target sum value using the weighting coefficient to obtain a weighted sum value; Update the model parameters of the initial prediction model from the initial model parameters to the difference between the initial model parameters and the weighted sum value.
9. The method according to claim 2, wherein The training the basic prediction model through the sample molecule set to obtain a target attribute prediction model includes: Screen the target training molecule set to obtain a first sample molecule set and a second sample molecule set. The first sample molecule set includes K sample groups composed of sample molecule features and sample parameter values, and the second sample molecule set includes L sample groups composed of sample molecule features and sample parameter values. The sample parameter values in the sample groups are the training supervision values of the sample molecules under the target attribute; Update the model parameters in the basic prediction model based on the first sample molecule set to obtain a first basic prediction model, and the model parameters in the first basic prediction model are updated from the initial parameters to the first parameters; Test the first basic prediction model based on the second sample molecule set to obtain a target test result; If the target test result meets the preset condition, then determine the first basic prediction model as the target attribute prediction model.
10. The method according to any one of claims 2-9, characterized in that, After training the basic prediction model through the sample molecule set to obtain a target attribute prediction model, the method further includes: Obtain the molecular characteristics of the target molecule to be subjected to the target attribute analysis, where the target molecule refers to the molecule for which the target attribute parameter value needs to be analyzed under the target attribute; Input the molecular characteristics into the target attribute prediction model associated with the target attribute, and analyze the molecular characteristics through the target attribute prediction model to obtain the parameter value of the target attribute corresponding to the target molecule.
11. A molecular property prediction device, characterized in that, The device includes: An acquisition module, configured to obtain the molecular characteristics of the target molecule to be subjected to the target attribute analysis, where the target molecule refers to the molecule for which the target attribute parameter value needs to be analyzed under the target attribute; A processing module, configured to input the molecular characteristics into the target attribute prediction model associated with the target attribute, and analyze the molecular characteristics through the target attribute prediction model to obtain the parameter value of the target attribute corresponding to the target molecule; Wherein, the target attribute prediction model is obtained by training a basic prediction model with a sample molecule set, and the sample molecule set includes: M sample groups composed of sample molecular characteristics and sample parameter values, and the sample parameter values in the sample group are the training supervision values of the sample molecules under the target attribute; The basic prediction model is pre-trained by N training molecule sets for the initial prediction model. Each training molecule set includes: P training groups composed of training molecular characteristics and training parameter values, and the training parameter values in the training group are the training supervision values of the training molecules in the training group under the reference attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute; wherein, the reference attribute includes multiple candidate attributes that meet the screening conditions selected from the candidate attributes, and the candidate attributes are selected from the attribute set and have an association relationship with the target attribute; Wherein, screening the candidate attributes having an association relationship with the target attribute from the attribute set includes: obtaining the co-occurrence frequency between each candidate attribute in the attribute set and the target attribute, where the co-occurrence frequency includes the frequency of the candidate attribute and the target attribute appearing in the historical attribute set at the same time, and the historical attribute set includes the set of attributes for analyzing the attributes of historical molecules, and determining the candidate attributes having an association relationship with the target attribute as the candidate attributes whose co-occurrence frequency with the target attribute is greater than the preset frequency.
12. A prediction model training device, characterized in that, The device includes: An acquisition module, configured to obtain N training molecule sets and obtain the target attribute to be subjected to the attribute analysis of the target molecule. Each training molecule set includes P training groups composed of training molecular characteristics and training parameter values, and the training parameter values in the training group are the training supervision values of the training molecules in the training group under the reference attribute; wherein, the acquisition module is further configured to screen out multiple candidate attributes that meet the screening conditions from the candidate attributes, and the candidate attributes are selected from the attribute set and have an association relationship with the target attribute; A training module, configured to pre-train the initial prediction model through the N training molecule sets to obtain a basic prediction model; The obtaining module is further configured to obtain a sample molecule set, where the sample molecule set includes M sample groups each composed of a sample molecule feature and a sample parameter value, and the sample parameter value in the sample group is a training supervision value of the sample molecule under a target attribute. M, N, and P are positive integers, and the reference attribute is different from the target attribute; The training module is further configured to train a basic prediction model through the sample molecule set to obtain a target attribute prediction model; Among them, when the obtaining module is used to screen out candidate attributes associated with the target attribute from the attribute set, it is configured to obtain the co-occurrence frequency between each candidate attribute in the attribute set and the target attribute. The co-occurrence frequency includes the frequency of the candidate attribute and the target attribute appearing simultaneously in the historical attribute set. The historical attribute set includes a set of attributes obtained by analyzing the attributes of historical molecules. The candidate attribute with a co-occurrence frequency greater than a preset frequency with the target attribute is determined as a candidate attribute associated with the target attribute.
13. An intelligent device, characterized in that, It includes a processor, an input interface, an output interface, and a memory. The processor, the input interface, the output interface, and the memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the method according to claim 1.
14. A terminal, characterized in that, It includes a processor, an input interface, an output interface, and a memory. The processor, the input interface, the output interface, and the memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the method according to any one of claims 2 - 10.
15. A computer storage medium, characterized in that, A computer program is stored in the computer storage medium. When the computer program is executed by a processor, it implements the method according to claim 1, or implements the method according to any one of claims 2 - 10.
Citation Information
Patent Citations
Sample recognition model generation method and device, computer equipment and storage medium
CN111444952A
Molecular attribute determination method and device, electronic equipment and storage medium
CN111724867A