Molecular property prediction methods, devices, equipment, media and products
By converting the scientific computing simulation process into a modeled data processing flow, combined with natural language text processing and Gaussian process regression models, the problem of low computational efficiency of massive small molecule data was solved, and efficient physicochemical property prediction and computational cycle reduction were achieved.
Patent Information
- Application Number
- CN202211738895.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-12-30
AI Technical Summary
When faced with scientific computing simulations of massive small molecule data, existing technologies suffer from low computing efficiency, long computing cycles, and high resource consumption.
The scientific computing simulation process is transformed into a modeled data processing flow, introducing feature extraction of molecular data and prediction of molecular physical and chemical properties based on feature extraction. Natural language text processing models such as BERT and Gaussian process regression models are used in combination with molecular property prediction models for feature extraction and physical and chemical property prediction.
It improves the efficiency of predicting the physical and chemical properties of massive molecular data, shortens the calculation cycle, and improves computing efficiency and resource utilization.
Smart Images

Figure CN115954062B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a method, device, equipment, medium and product for predicting molecular properties. Background Art
[0002] The physicochemical properties of small molecules used in energy materials have long relied on experimental testing, consuming significant experimental resources. With the development of high-performance computing technology, scientific computational simulation has become an important means of obtaining data on the physicochemical properties of small molecules. For example, physicochemical properties such as electrical conductivity can be determined through scientific computational simulation.
[0003] However, when conducting scientific computing simulations based on massive amounts of small molecule data, how to improve computing efficiency and shorten the computing cycle becomes a technical problem that needs to be solved at this stage. Summary of the Invention
[0004] The embodiments of the present application provide a molecular property prediction method, apparatus, device, medium and product to solve the problem of how to improve the computational efficiency of scientific computational simulation of massive small molecule data and shorten the computational cycle.
[0005] A first aspect of the embodiments of the present application provides a method for predicting molecular properties, comprising:
[0006] Performing molecular feature extraction based on the first molecular data set to obtain a first feature extraction result;
[0007] Based on the first feature extraction result, the molecular physicochemical property prediction is performed to obtain a first physicochemical property prediction result.
[0008] In the technical solution of the embodiment of the present application, the scientific computing simulation process is converted into a modeled data processing flow, and a processing method for feature extraction of molecular data in a data set and prediction of molecular physicochemical properties based on feature extraction is introduced to improve the efficiency of prediction of physicochemical properties of massive molecular data and shorten the calculation cycle.
[0009] In some embodiments, performing molecular feature extraction based on the first molecular dataset to obtain a first feature extraction result includes:
[0010] extracting information text corresponding to the molecular configuration information from the first molecular data set;
[0011] Perform feature extraction based on the information text to obtain a feature vector representation corresponding to the molecular configuration information;
[0012] The feature vector representation is determined as the first feature extraction result.
[0013] In this processing process, the information text corresponding to the molecular configuration information is used as the feature extraction object, and feature extraction processing is implemented to improve the feasibility and efficiency of molecular configuration feature extraction.
[0014] In some embodiments, the performing of molecular physicochemical property prediction based on the first feature extraction result to obtain a first physicochemical property prediction result includes:
[0015] Based on the first feature extraction result, the probability distribution of the feature extraction result in the set physicochemical property of the molecule is calculated to obtain the first physicochemical property prediction result.
[0016] This process, by calculating the probability distribution, facilitates confidence assessment of the prediction results of physical and chemical properties, and improves the interpretability and credibility of the prediction results.
[0017] In some embodiments, the step of calculating the probability distribution of the feature extraction result in the physicochemical property of the set molecule based on the first feature extraction result to obtain the first physicochemical property prediction result includes:
[0018] Based on the first feature extraction result, the probability distribution of the feature extraction result in the set molecular physicochemical property is calculated by Gaussian process regression to obtain the first physicochemical property prediction result.
[0019] This process introduces Gaussian process regression to implement physicochemical property prediction processing, thereby improving the interpretability and quantifiability of the physicochemical property prediction results.
[0020] In some embodiments, before performing molecular feature extraction based on the first molecular dataset to obtain a first feature extraction result, the method further includes:
[0021] Perform format conversion on the pre-built molecular configuration file to obtain the information text corresponding to the molecular configuration information;
[0022] Based on the information text, the first molecular data set is generated.
[0023] In this way, the textual conversion of molecular configuration data is achieved, ensuring the feasibility and efficiency of subsequent molecular configuration feature extraction.
[0024] In some embodiments, generating the first molecular dataset based on the information text includes:
[0025] Marking target molecules with repeated configurations in the molecular configuration file;
[0026] The information text is deduplicated according to the target molecule to generate the first molecule data set including the deduplicated information text.
[0027] In this way, configuration deduplication of the information text corresponding to the molecular configuration information is achieved.
[0028] In some embodiments, the molecular feature extraction is performed based on the first molecular dataset to obtain a first feature extraction result which is implemented by a first model portion of the molecular property prediction model;
[0029] The molecular physicochemical property prediction is performed based on the first feature extraction result, and the first physicochemical property prediction result is obtained by the second model part in the molecular property prediction model.
[0030] This process implements the molecular property prediction method through the molecular property prediction model, combines the molecular property prediction method with the artificial intelligence model, models the molecular structure-activity relationship, and automatically implements the processing flow of the molecular property prediction method, thereby improving the efficiency of physicochemical property prediction of massive molecular data and shortening the calculation cycle.
[0031] In some embodiments, the method further comprises:
[0032] Based on the target molecule training sample, the molecular property prediction model is trained until the molecular property prediction model meets the set convergence condition.
[0033] This process, by pre-training the model, ensures that the molecular property prediction model learns the mapping relationship between molecular configuration and molecular physicochemical properties, so that the model is constructed into a molecular structure-activity relationship model, ensuring the effectiveness of the model in implementing the molecular property prediction method.
[0034] In some embodiments, the training of the molecular property prediction model based on the target molecule training sample until the molecular property prediction model meets the set convergence condition includes:
[0035] Extracting features of the target molecule training sample based on the pre-trained first model portion to obtain molecular sample features;
[0036] The second model portion is trained based on the molecular sample characteristics until the second model portion meets the set convergence condition.
[0037] This ensures that the feature extraction results have reliable accuracy in the initial stage of model training, thereby improving the overall training accuracy and efficiency of the model.
[0038] In some embodiments, after training the second model part based on the molecular sample features, the method further includes:
[0039] When the second model partially satisfies the set convergence condition, selecting a plurality of target molecule data from the second molecule data set;
[0040] Based on the plurality of target molecule data, generating the target molecule training sample;
[0041] Return to executing the step of extracting features from the target molecule training sample based on the pre-trained first model portion to obtain molecular sample features, until the second model portion meets the set convergence condition.
[0042] This process implements multiple rounds of model training with fewer samples required for single training, improving model training efficiency and shortening the model training cycle.
[0043] In some embodiments, selecting a plurality of target molecule data from the second molecular dataset comprises:
[0044] Based on the first model part, performing molecular feature extraction on each of the multiple data sets divided from the second molecular data set to obtain a second feature extraction result;
[0045] Based on the second model part, performing molecular physicochemical property prediction on the second feature extraction result to obtain a second physicochemical property prediction result;
[0046] selecting a target set from the plurality of data sets based on the second molecular property prediction result;
[0047] The plurality of molecular data included in the target set is determined as the target molecular data.
[0048] In this implementation process, a small number of samples from different batches are continuously selected from the molecular dataset based on the trained model that meets the convergence conditions to continuously implement model optimization training, thereby improving the model training efficiency and shortening the model training cycle.
[0049] In some embodiments, selecting a target set from the plurality of data sets based on the second molecular property prediction result comprises:
[0050] Scoring each of the data sets based on the second molecular property prediction result to obtain a set score;
[0051] The target set is selected from the plurality of data sets based on the set scores.
[0052] In this process, a set scoring mechanism is introduced to ensure the quality of selected samples and improve the model training effect.
[0053] In some embodiments, after selecting a target set from the plurality of data sets based on the second molecular property prediction result, the method further includes:
[0054] The target set is removed from the second molecular dataset to obtain an updated second molecular dataset.
[0055] In this way, repeated selection of model training samples is avoided.
[0056] In some embodiments, generating the target molecule training sample based on the plurality of target molecule data comprises:
[0057] Performing scientific computational simulations on the plurality of target molecule data respectively to obtain the molecular physicochemical properties corresponding to each target molecule data;
[0058] The physicochemical properties of the molecules are used as labels of the corresponding target molecule data to obtain the target molecule training samples including the target molecule data and the labels.
[0059] This process enables the formed molecular training samples to have molecular structure-activity relationship characteristics corresponding to scientific computational simulations, thereby ensuring the effectiveness of subsequent model training.
[0060] A second aspect of an embodiment of the present application provides a molecular property prediction device, comprising:
[0061] A feature extraction module, configured to extract molecular features based on the first molecular data set to obtain a first feature extraction result;
[0062] The property prediction module is used to predict the physicochemical properties of the molecule based on the first feature extraction result to obtain a first physicochemical property prediction result.
[0063] A third aspect of an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the computer program.
[0064] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0065] A fifth aspect of the present application provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.
[0066] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference numerals are used throughout the drawings to represent the same components. In the drawings:
[0068] Figure 1 is a flow chart of a molecular property prediction method according to some embodiments of the present application;
[0069] Figure 2 is a flowchart of feature extraction in some embodiments of the present application;
[0070] Figure 3 is a flowchart of generating a molecular dataset in some embodiments of the present application;
[0071] Figure 4 is a schematic diagram of a molecular sample set according to some embodiments of the present application;
[0072] Figure 5 is a schematic diagram of a molecular property prediction model according to some embodiments of the present application;
[0073] Figure 6 This is a flowchart of model training in some embodiments of the present application;
[0074] Figure 7 This is a flowchart of model training in some embodiments of the present application;
[0075] Figure 8 is a flowchart of selecting molecular data in some embodiments of the present application;
[0076] Figure 9 is a flowchart of generating molecular training samples in some embodiments of the present application;
[0077] Figure 10 is a structural diagram of a molecular property prediction device provided in an embodiment of the present application;
[0078] Figure 11 This is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0079] The following embodiments of the technical solution of the present application will be described in detail with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present application and are therefore only examples and are not intended to limit the scope of protection of the present application.
[0080] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned figure descriptions are intended to cover non-exclusive inclusions.
[0081] In the description of the embodiments of this application, the technical terms "first" and "second" are used only to distinguish different objects and should not be understood to indicate or imply relative importance or implicitly specify the quantity, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "plurality" is more than two, unless otherwise clearly and specifically defined.
[0082] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0083] In the description of the embodiments of this application, the term "and / or" is simply a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0084] In the field of small molecules for energy materials, various physical and chemical properties have long relied on experimental testing, consuming a large amount of experimental resources.
[0085] With the development of high-performance computing technology, methods such as first-principles calculations and molecular dynamics calculations are often used in the technical field for scientific computing simulations, which has also become an important means of obtaining small molecule physicochemical performance data.
[0086] However, during the specific application process, the inventors found that the number of small molecules in energy materials is huge, and scientific calculations themselves often require a large amount of computing power resources. When faced with computational simulations of massive data, problems such as low computing efficiency, long computing cycles, and high resource consumption often arise.
[0087] Therefore, the inventors came up with the idea of modeling the scientific computing simulation process, introducing feature extraction of molecular data and implementing processing methods to predict the physicochemical properties of molecules based on feature extraction, so as to improve the efficiency of predicting the physicochemical properties of massive molecular data and shorten the calculation cycle.
[0088] Combine Figure 1 As shown, in some embodiments, a molecular property prediction method is proposed, comprising:
[0089] Step 101: Perform molecular feature extraction based on a first molecular data set to obtain a first feature extraction result.
[0090] In this step, when processing massive molecular data, it is necessary to extract features of the molecular data based on the molecular data set.
[0091] The molecular data set includes a set amount of molecular data, and the molecular data corresponds to a molecular configuration. Specifically, the molecular data includes, for example, a molecular configuration diagram, a molecular configuration description text, and the like.
[0092] During the processing, molecular features are extracted using molecular data in the molecular data set as processing objects.
[0093] In this process, a data set consisting of molecular data is taken as the processing object, and feature extraction is performed on a group of molecular features to ensure the processing efficiency of data calculation.
[0094] The extracted molecular features are specifically general features of the molecular configuration, which are specifically feature information corresponding to the geometric shapes of the spatial distribution of various groups or atoms in the molecule. Feature information includes, for example, spatial position information and spatial geometric distribution information.
[0095] For example, the spatial position information of oxygen atoms and hydrogen atoms in water molecules, the structural characteristics of their spatial geometric distribution, etc.
[0096] Step 102: Based on the first feature extraction result, the molecular physicochemical property prediction is performed to obtain a first physicochemical property prediction result.
[0097] Physicochemical properties include, for example, the electrical conductivity, binding energy, and other physical and chemical properties of molecules.
[0098] The processing process of the above steps corresponds to a modeled data processing flow. After first extracting molecular features based on the data set, it is necessary to predict the molecular physicochemical properties on this basis.
[0099] When predicting molecular physicochemical properties based on feature extraction results, the prediction results can be obtained by combining the mapping relationship between molecular features and molecular physicochemical properties. The mapping relationship can be learned based on data samples using a machine learning model.
[0100] In the technical solution of the embodiment of the present application, the scientific computing simulation process is converted into a modeled data processing flow, and a processing method for feature extraction of molecular data in a data set and prediction of molecular physicochemical properties based on feature extraction is introduced to improve the efficiency of prediction of physicochemical properties of massive molecular data and shorten the calculation cycle.
[0101] In some embodiments, combined Figure 2 As shown, step 101 performs molecular feature extraction based on the first molecular data set to obtain a first feature extraction result, which specifically includes the following processing steps:
[0102] Step 201: extracting information text corresponding to molecular configuration information from a first molecular data set;
[0103] Step 202: perform feature extraction based on the information text to obtain a feature vector representation corresponding to the molecular configuration information.
[0104] Step 203: Determine the feature vector representation as the first feature extraction result.
[0105] Specifically, when feature extraction is performed based on information text, a natural language text processing model is specifically used to implement it.
[0106] Since each small molecule normally presents the bonding relationship between atoms or a molecule as a whole, it is difficult to directly implement feature processing and feature extraction is difficult.
[0107] To overcome this problem, in the above implementation process, the molecular data in the molecular dataset is text data, and feature extraction is performed based on the information text corresponding to the molecular configuration information.
[0108] The natural language text processing model uses, for example, the BERT model. During processing, the information text needs to be segmented, and word embedding, segment embedding, and position embedding are performed based on the segmentation results to obtain word embedding vectors, segment embedding vectors, and position embedding vectors. Based on these embedding results, a feature vector representation corresponding to the molecular configuration information is obtained, i.e., a first feature extraction result is obtained. Alternatively, other natural language text processing models can be used to extract features from the information text, and the selection can be made based on specific needs.
[0109] In this embodiment, the information text corresponding to the molecular configuration information is used as a feature extraction object, and feature extraction processing is performed to improve the feasibility and efficiency of molecular configuration feature extraction.
[0110] In some embodiments, step 102 performs molecular physicochemical property prediction based on the first feature extraction result to obtain a first physicochemical property prediction result, which specifically includes the following processing steps:
[0111] Based on the first feature extraction result, the probability distribution of the feature extraction result in the set molecular physicochemical property is calculated to obtain the first physicochemical property prediction result.
[0112] The set molecular physicochemical property is at least one.
[0113] During this process, the probability distribution of each molecular configuration's corresponding feature extraction results for the given molecular physicochemical properties is predicted. By calculating the probability distribution of the feature extraction results for the given molecular physicochemical properties, the likelihood that the extracted features of the molecular data correspond to a given molecular physicochemical property is determined, and the predicted molecular physicochemical properties corresponding to the molecular data are then obtained.
[0114] In the process of calculating the probability distribution to obtain the first physicochemical property prediction result, the feature extraction result can be obtained based on the probability distribution to set the confidence interval of the molecular physicochemical property, and the confidence interval or the confidence level corresponding to the confidence interval can be used as the first physicochemical property prediction result.
[0115] The above process predicts the physicochemical properties of molecular data by calculating the probability distribution of the feature extraction results under the set molecular physicochemical properties, so as to facilitate the confidence assessment of the physicochemical property prediction results and improve the interpretability and credibility of the prediction results.
[0116] In some embodiments, the above-mentioned physicochemical property prediction processing operation may optionally be implemented using a Gaussian process regression model.
[0117] Correspondingly, in the process of calculating the probability distribution of the feature extraction result in the set molecular physicochemical property based on the first feature extraction result and obtaining the first physicochemical property prediction result, the process specifically includes:
[0118] Based on the first feature extraction result, the probability distribution of the feature extraction result in the set molecular physicochemical property is calculated through Gaussian process regression to obtain the first physicochemical property prediction result.
[0119] Specifically, Gaussian process regression is a nonparametric model that uses a Gaussian process prior to perform regression analysis on data. Compared to typical supervised learning models, Gaussian process regression models have a more streamlined structure and fewer parameters. In addition, compared to other non-Bayesian models, Gaussian process regression models have a clear probability formula, allowing model users to assess the uncertainty of the model's output predictions through confidence intervals or posterior probabilities.
[0120] This process introduces Gaussian process regression to implement physicochemical property prediction processing, thereby improving the interpretability and quantifiability of the physicochemical property prediction results.
[0121] Furthermore, in some embodiments, before performing molecular feature extraction on a molecular data set, it is also necessary to perform data pre-processing on the molecular data.
[0122] Specifically, combined Figure 3 As shown, step 101 performs molecular feature extraction based on the first molecular data set, and before obtaining the first feature extraction result, it also includes:
[0123] Step 301 : convert the format of the pre-built molecular configuration file to obtain information text corresponding to the molecular configuration information.
[0124] The initial sample space can be constructed for specific application scenarios. The sample space corresponds to the molecular configuration sample set. The molecular configuration file records the data of each molecular configuration sample in the molecular configuration sample set. Figure 4 As shown, in this embodiment, an electrolyte small molecule sample set is specifically constructed, wherein the content contained in the molecular configuration file is a molecular configuration diagram.
[0125] In the above steps, when converting the format of the pre-built molecular configuration file, the molecular configuration file can be processed by openbabel, RDKit, etc. to convert the format to obtain the corresponding SMILES (Simplified molecular input line entry system) code and the molecular configuration string as the information text.
[0126] Alternatively, when converting the format of a pre-built molecular configuration file, the molecular configuration diagram in the molecular configuration file can be directly subjected to image content recognition and converted into corresponding information text.
[0127] This is just an example description and is not intended to be limiting.
[0128] Step 302: Generate a first molecular data set based on the information text.
[0129] After obtaining the information text corresponding to the molecular configuration information, the first molecular data set in text form can be obtained.
[0130] In this way, the textual conversion of the molecular configuration data is achieved, so as to facilitate the subsequent feature extraction based on the information text corresponding to the molecular configuration information, thereby ensuring the feasibility and efficiency of the subsequent molecular configuration feature extraction.
[0131] Furthermore, in some embodiments, step 302 generates a first molecular dataset based on the information text, including:
[0132] Mark target molecules with repeated configurations in molecular configuration files;
[0133] The information text is deduplicated according to the target molecule to generate a first molecule data set including the deduplicated information text.
[0134] In this way, molecules with repeated configurations are marked, and configuration duplication of the information text corresponding to the molecular configuration information is achieved.
[0135] Furthermore, the above-mentioned molecular property prediction method in the embodiments of the present application needs to be implemented based on molecular property prediction.
[0136] Furthermore, in some embodiments, the aforementioned step 101 performs molecular feature extraction based on the first molecular dataset, and the step of obtaining the first feature extraction result is implemented by the first model portion of the molecular property prediction model. The aforementioned step 102 performs molecular physicochemical property prediction based on the first feature extraction result, and the step of obtaining the first physicochemical property prediction result is implemented by the second model portion of the molecular property prediction model.
[0137] As an artificial intelligence model, the molecular property prediction model corresponds to the purchase and sale relationship of molecules, so as to be able to implement the processing process of molecular feature extraction and molecular physicochemical property prediction.
[0138] The first model part and the second model part of the molecular property prediction model are two model components with an upstream and downstream relationship. The first model part is used to implement feature extraction; the second model part is used to implement property prediction.
[0139] This process implements the molecular property prediction method through the molecular property prediction model, combines the molecular property prediction method with the artificial intelligence model, models the molecular structure-activity relationship, and automatically implements the processing flow of the molecular property prediction method, thereby improving the efficiency of physicochemical property prediction of massive molecular data and shortening the calculation cycle.
[0140] Among them, molecular property prediction models can be searched based on open source model libraries such as HuggingFance.
[0141] When the first model part of the molecular property prediction model is a natural language text processing model, the sample data can be adjusted when constructing the molecular training sample to adapt to the word segmenter of the first model part.
[0142] In an alternative example, combined with Figure 5 As shown, the molecular property prediction model is a fusion model of BERT and Gaussian process regression. The first model part of the molecular property prediction model is the BERT model, and the second model part is the Gaussian process regression model. Based on the structure of this molecular property prediction model, the molecular dataset is used as the model input, and the BERT model performs feature extraction processing. The output is then fed into the Gaussian process regression model for physicochemical property prediction, and the physicochemical property prediction results are finally used as the model output.
[0143] Furthermore, based on the above-mentioned model structure, in the embodiment of the present application, there is also a model training process before the model is applied.
[0144] In some embodiments, the method of the present application further comprises:
[0145] Based on the target molecule training samples, the molecular property prediction model is trained until the molecular property prediction model meets the set convergence conditions.
[0146] The set model convergence conditions include, for example, that the physical and chemical property prediction results of the model meet a set performance indicator, or that the number of cycles of the iterative optimization of the model during model training reaches the target value.
[0147] During the model training process, the model is optimized to converge, for example, by adjusting the hyperparameters in the kernel function, loss function, or likelihood estimation in the model.
[0148] This process, by pre-training the model, ensures that the molecular property prediction model learns the mapping relationship between molecular configuration and molecular physicochemical properties, making the model a molecular structure-activity relationship construction, ensuring the effectiveness of the model in implementing the molecular property prediction method.
[0149] In some embodiments, combined Figure 6 As shown, the aforementioned step of training the molecular property prediction model based on the target molecule training sample until the molecular property prediction model meets the set convergence conditions includes:
[0150] Step 601, extracting features of a target molecule training sample based on a pre-trained first model portion to obtain molecular sample features;
[0151] Step 602: training the second model part based on the molecular sample characteristics.
[0152] The cutoff condition for training the second model part based on the molecular sample features is until the second model part meets the set convergence condition.
[0153] During the training of the second model part based on the characteristics of the molecular sample, each time the model iterates, it is necessary to determine whether the second model part meets the set convergence conditions. If it is judged to be yes, the model training is completed. If not, the model parameters are adjusted and the iterative process is continued until the second model part meets the set convergence conditions.
[0154] Specifically, in the above process, in this embodiment, the first model part is set as a pre-trained component of the molecular property prediction model, and the second model part is an untrained component of the molecular property prediction model.
[0155] During the model training process for the molecular property prediction model, the second model portion is primarily trained. When training the model based on molecular training samples, whether the molecular property prediction model meets the convergence conditions is determined based on whether the second model portion of the molecular property prediction model meets the set convergence conditions.
[0156] In this process, pre-trained model components are used in the molecular property prediction model to implement feature extraction of molecular training samples, and untrained model components in the model are trained to ensure that the feature extraction results have reliable accuracy in the initial stage of model training, thereby improving the overall training accuracy and training efficiency of the model.
[0157] In some embodiments, combined Figure 7 As shown, after the second model part is trained based on the molecular sample features in step 602, the following steps are further included:
[0158] Step 603: When the second model partially satisfies the set convergence condition, a plurality of target molecule data are selected from the second molecule data set;
[0159] Step 604: generating target molecule training samples based on the multiple target molecule data;
[0160] Return to step 601 to extract features of the target molecule training sample based on the pre-trained first model part to obtain molecular sample features, until the second model part meets the set convergence condition.
[0161] In this process, it is necessary to select multiple target molecule data from the second molecule data set to form target molecule training samples for the current model training.
[0162] Based on the target molecule training sample currently selected, a round of model training for the molecular property prediction model is performed, and the second model part in the molecular property prediction model satisfies the set convergence condition.
[0163] Subsequently, after the second model portion of the molecular property prediction model in the model meets the set convergence conditions, multiple target molecule data are continuously selected from the second molecular data set to form another target molecule training sample, and the next round of model training processing flow is carried out, and the second model portion of the molecular property prediction model meets the set convergence conditions.
[0164] The termination conditions of the multi-round model training process may be, but are not limited to:
[0165] It is detected that all the molecular data in the second molecular data set have been selected, or it is detected that the amount of molecular data in the second molecular data set that has been selected has reached a set proportion, or it is detected that the model prediction performance of the second model part in the molecular property prediction model has reached a set standard value after meeting the set convergence condition, or it is detected that the number of times the second model part in the molecular property prediction model meets the set convergence condition has reached a number threshold, or a model training termination instruction is received.
[0166] In this process, during multiple rounds of model training, molecular training samples are continuously selected from the molecular dataset to optimize the model training, and the molecular data in the molecular dataset are effectively used for batch model training. Multiple rounds of model training are implemented with fewer samples required for single training, thereby improving the model training efficiency and shortening the model training cycle.
[0167] In some embodiments, combined Figure 8 As shown, step 603 selects a plurality of target molecule data from the second molecule data set, including:
[0168] Step 801: Based on the first model part, molecular feature extraction is performed on each of the multiple data sets divided from the second molecular data set to obtain a second feature extraction result;
[0169] Step 802: Based on the second model part, perform molecular physicochemical property prediction on the second feature extraction result to obtain a second physicochemical property prediction result;
[0170] Step 803: selecting a target set from multiple data sets based on the second molecular property prediction result;
[0171] Step 804: Determine the multiple molecular data included in the target set as target molecular data.
[0172] In the above process, multiple data sets are divided from the second molecular data set, and a target set is selected from these data sets as data samples for the current round of model training.
[0173] By dividing the molecular data set into data sets and inputting the divided data sets into a molecular property prediction model that meets the convergence conditions, the corresponding physicochemical property prediction results are obtained respectively. In the process of multiple rounds of model training, the next data set is continuously selected from the molecular data set as a valid molecular training sample based on the model that has already met the convergence conditions.
[0174] Among them, when carrying out the current round of model training, it is necessary to select the target molecule data required for the current round from the second molecular data set based on the molecular property prediction model that reached the model convergence conditions in the previous round to obtain the target molecule training samples for this round.
[0175] In this implementation process, the model is trained based on a small number of samples in the molecular dataset, and based on the trained model that meets the convergence conditions, a small number of samples from different batches are continuously selected from the molecular dataset to continuously implement model optimization training, thereby improving the model training efficiency and shortening the model training cycle.
[0176] In some embodiments, step 803 selects a target set from multiple data sets based on the second molecular property prediction result, including:
[0177] Based on the second molecular property prediction result, each data set is scored to obtain a set score;
[0178] Based on the set scores, a target set is selected from multiple data sets.
[0179] When scoring each data set, the predicted value corresponding to the molecular property prediction result can be substituted into the scoring calculation formula to obtain the set score. The predicted value can be, for example, a probability distribution value, a confidence interval value, etc.
[0180] When selecting a target set from multiple data sets based on the set scores, the set with the highest score can be selected from the multiple data sets as the target set based on the set scores. Specific settings can be made as needed and are not limited to this.
[0181] This process introduces a set scoring mechanism to enable the model-based molecular property prediction results to select reasonable and effective data samples from the molecular data set, ensuring the quality of sample selection and improving the model training effect.
[0182] In some embodiments, after selecting a target set from multiple data sets based on the second molecular property prediction result in step 803, the method further includes:
[0183] The target set is removed from the second molecular dataset to obtain an updated second molecular dataset.
[0184] In this way, it can be ensured that when selecting a target set from the molecular dataset, the selection is achieved from the remaining data sets of the molecular dataset, thus avoiding repeated selection of model training samples.
[0185] In some embodiments, combined Figure 9 As shown, step 604 generates target molecule training samples based on multiple target molecule data, including:
[0186] Step 901 : Perform scientific computational simulation on multiple target molecule data to obtain molecular physicochemical properties corresponding to each target molecule data.
[0187] Here, the scientific computational simulation of the target molecular data may be a first-principles computational simulation or a molecular dynamics computational simulation.
[0188] By performing scientific computational simulation on the target molecule data, the molecular physicochemical properties of the target molecule data are obtained to obtain a label.
[0189] Step 902 : Using the physicochemical properties of the molecules as labels of the corresponding target molecule data, a target molecule training sample including the target molecule data and the labels is obtained.
[0190] This step generates molecular training samples containing target molecule data and their labels.
[0191] In specific applications, it is necessary to perform model training on the molecular property prediction model based on the target molecule training sample. The process can be:
[0192] The target molecule training sample is input into the molecular property prediction model. The first model part of the molecular property prediction model performs feature processing based on the target molecule data in the molecular property prediction model to obtain feature processing results, and the feature processing results are input into the second model part to predict the molecular physicochemical properties.
[0193] The molecular physicochemical property prediction results are used to adjust the parameters of the second model part 5 times based on the labels and the molecular physicochemical property prediction results, and the model is optimized iteratively until the second model part meets the set convergence conditions.
[0194] In this process, scientific computational simulation is performed on the target molecular data in advance, and the physicochemical properties simulated by scientific computation are used as data labels, so that the formed molecular training samples have the molecular structure-activity relationship characteristics corresponding to the scientific computational simulation, so as to ensure the effect of subsequent model training.
[0195] It should be understood that although the steps in the flowcharts of the embodiments described above are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders.
[0196] At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0197] Based on the same inventive concept, the embodiments of the present application further provide a data processing device. The data processing device provided in the embodiments of the present application can implement each process of the embodiments of the above-mentioned data processing method and can achieve the same technical effects. Therefore, the specific limitations in one or more data processing device embodiments provided below can refer to the limitations of the data processing method above. To avoid repetition, they are not repeated here.
[0198] In one embodiment, Figure 10 As shown, a molecular property prediction device 10 is provided, including: a feature extraction module and a property prediction module.
[0199] A feature extraction module 1001 is configured to extract molecular features based on a first molecular dataset to obtain a first feature extraction result;
[0200] The property prediction module 1002 is used to predict the physicochemical properties of the molecule based on the first feature extraction result to obtain a first physicochemical property prediction result.
[0201] In some embodiments, the feature extraction module 1001 is specifically configured to:
[0202] extracting information text corresponding to the molecular configuration information from the first molecular data set;
[0203] Perform feature extraction based on the information text to obtain a feature vector representation corresponding to the molecular configuration information;
[0204] The feature vector representation is determined as the first feature extraction result.
[0205] In some embodiments, the property prediction module 1002 is specifically configured to:
[0206] Based on the first feature extraction result, the probability distribution of the feature extraction result in the set physicochemical property of the molecule is calculated to obtain the first physicochemical property prediction result.
[0207] In some embodiments, the property prediction module 1002 is further specifically configured to:
[0208] Based on the first feature extraction result, the probability distribution of the feature extraction result in the set molecular physicochemical property is calculated by Gaussian process regression to obtain the first physicochemical property prediction result.
[0209] In some embodiments, the apparatus further comprises:
[0210] Data collection generation module, used to:
[0211] Perform format conversion on the pre-built molecular configuration file to obtain the information text corresponding to the molecular configuration information;
[0212] Based on the information text, the first molecular data set is generated.
[0213] In some embodiments, the data set generation module is specifically configured to:
[0214] Marking target molecules with repeated configurations in the molecular configuration file;
[0215] The information text is deduplicated according to the target molecule to generate the first molecule data set including the deduplicated information text.
[0216] In some embodiments, the molecular feature extraction is performed based on the first molecular data set, and the first feature extraction result is obtained by the first model part in the molecular property prediction model; the molecular physicochemical property prediction is performed based on the first feature extraction result, and the first physicochemical property prediction result is obtained by the second model part in the molecular property prediction model.
[0217] In some embodiments, the apparatus further comprises:
[0218] Model training module, used to:
[0219] Based on the target molecule training sample, the molecular property prediction model is trained until the molecular property prediction model meets the set convergence condition.
[0220] In some embodiments, the model training module is specifically used to:
[0221] Extracting features of the target molecule training sample based on the pre-trained first model portion to obtain molecular sample features;
[0222] The second model portion is trained based on the molecular sample characteristics until the second model portion meets the set convergence condition.
[0223] In some embodiments, the model training module is further specifically configured to:
[0224] When the second model partially satisfies the set convergence condition, selecting a plurality of target molecule data from the second molecule data set;
[0225] Based on the plurality of target molecule data, generating the target molecule training sample;
[0226] Return to executing the step of extracting features from the target molecule training sample based on the pre-trained first model portion to obtain molecular sample features, until the second model portion meets the set convergence condition.
[0227] In some embodiments, the model training module is further specifically configured to:
[0228] Based on the first model part, performing molecular feature extraction on each of the multiple data sets divided from the second molecular data set to obtain a second feature extraction result;
[0229] Based on the second model part, performing molecular physicochemical property prediction on the second feature extraction result to obtain a second physicochemical property prediction result;
[0230] selecting a target set from the plurality of data sets based on the second molecular property prediction result;
[0231] The plurality of molecular data included in the target set is determined as the target molecular data.
[0232] In some embodiments, the model training module is further specifically configured to:
[0233] Scoring each of the data sets based on the second molecular property prediction result to obtain a set score;
[0234] The target set is selected from the plurality of data sets based on the set scores.
[0235] In some embodiments, the model training module is further configured to:
[0236] The target set is removed from the second molecular dataset to obtain an updated second molecular dataset.
[0237] In some embodiments, the model training module is further specifically configured to:
[0238] Performing scientific computational simulations on the plurality of target molecule data respectively to obtain the molecular physicochemical properties corresponding to each target molecule data;
[0239] The physicochemical properties of the molecules are used as labels of the corresponding target molecule data to obtain the target molecule training samples including the target molecule data and the labels.
[0240] Each module in the molecular property prediction device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0241] In one embodiment, Figure 11 As shown, a computer device is provided. The computer device 11 of this embodiment includes: at least one processor 1100 ( Figure 11Only one is shown), a memory 1101 and a computer program 1102 stored in the memory 1101 and executable on the at least one processor 1100, wherein the processor 1100 implements the steps of any of the above-mentioned method embodiments when executing the computer program 1102.
[0242] The computer device 11 may be a desktop computer, a notebook computer, a PDA, a cloud server or other computing devices. The computer device 11 may include, but is not limited to, a processor 1100 and a memory 1101. Those skilled in the art will understand that Figure 11 This is merely an example of the computer device 11 and does not constitute a limitation on the computer device 11. The computer device 11 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.
[0243] The processor 1100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0244] The memory 1101 may be an internal storage unit of the computer device 11, such as a hard disk or memory of the computer device 11. The memory 1101 may also be an external storage device of the computer device 11, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 11. Furthermore, the memory 1101 may include both an internal storage unit of the computer device 11 and an external storage device. The memory 1101 is used to store the computer program and other programs and data required by the computer device. The memory 1101 may also be used to temporarily store data that has been output or is about to be output.
[0245] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0246] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0247] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0248] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which can be electrical, mechanical or other forms.
[0249] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0250] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0251] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0252] The present application implements all or part of the processes in the above-mentioned embodiment methods, and may also be implemented through a computer program product. When the computer program product runs on a computer device, the computer device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0253] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for predicting molecular properties, characterized in that: include: Performing molecular feature extraction based on the first molecular data set to obtain a first feature extraction result; implemented by a first model portion of a molecular property prediction model; Based on the first feature extraction result, a probability distribution of the feature extraction result in the set molecular physicochemical property is calculated by Gaussian process regression to obtain a first physicochemical property prediction result; which is implemented by the second model part of the molecular property prediction model; The method further includes selecting a plurality of target molecule data from the second molecule data set when the second model partially satisfies the set convergence condition; generating a target molecule training sample based on the plurality of target molecule data; Returning to the step of extracting features from the target molecule training sample based on the pre-trained first model portion to obtain molecular sample features, until the second model portion satisfies the set convergence condition; The selecting of the plurality of target molecular data from the second molecular data set comprises: performing molecular feature extraction on each of the plurality of data sets divided from the second molecular data set based on the first model portion to obtain a second feature extraction result; Based on the second model part, the second feature extraction result is predicted for molecular physicochemical properties to obtain a second physicochemical property prediction result; based on the second molecular property prediction result, a target set is selected from the multiple data sets; and the multiple molecular data contained in the target set are determined as the target molecular data.
2. The method according to claim 1, characterized in that The performing molecular feature extraction based on the first molecular data set to obtain a first feature extraction result includes: extracting information text corresponding to the molecular configuration information from the first molecular data set; Perform feature extraction based on the information text to obtain a feature vector representation corresponding to the molecular configuration information; The feature vector representation is determined as the first feature extraction result.
3. The method according to claim 1, characterized in that The step of predicting the molecular physicochemical properties based on the first feature extraction result to obtain the first physicochemical property prediction result includes: Based on the first feature extraction result, the probability distribution of the feature extraction result in the set physicochemical property of the molecule is calculated to obtain the first physicochemical property prediction result.
4. The method according to claim 1, wherein Before extracting molecular features based on the first molecular dataset and obtaining a first feature extraction result, the method further includes: Perform format conversion on the pre-built molecular configuration file to obtain the information text corresponding to the molecular configuration information; Based on the information text, the first molecular data set is generated.
5. The method according to claim 4, characterized in that The step of generating the first molecular data set based on the information text includes: Marking target molecules with repeated configurations in the molecular configuration file; The information text is deduplicated according to the target molecule to generate the first molecule data set including the deduplicated information text.
6. The method according to claim 1, characterized in that The method further comprises: Based on the target molecule training sample, the molecular property prediction model is trained until the molecular property prediction model meets the set convergence condition.
7. The method according to claim 6, characterized in that The molecular property prediction model is trained based on the target molecule training sample until the molecular property prediction model meets the set convergence condition, including: Extracting features of the target molecule training sample based on the pre-trained first model portion to obtain molecular sample features; The second model portion is trained based on the molecular sample characteristics until the second model portion meets the set convergence condition.
8. The method according to claim 1, characterized in that The selecting a target set from the plurality of data sets based on the second molecular property prediction result includes: Scoring each of the data sets based on the second molecular property prediction result to obtain a set score; The target set is selected from the plurality of data sets based on the set scores.
9. The method according to claim 1, characterized in that After selecting the target set from the plurality of data sets based on the second molecular property prediction result, the method further includes: removing the target set from the second molecular data set to obtain an updated second molecular data set.
10. The method according to claim 1, characterized in that The generating the target molecule training sample based on the plurality of target molecule data comprises: Performing scientific computational simulations on the plurality of target molecule data respectively to obtain the molecular physicochemical properties corresponding to each target molecule data; The physicochemical properties of the molecules are used as labels of the corresponding target molecule data to obtain the target molecule training samples including the target molecule data and the labels.
11. A molecular property prediction device, characterized in that: include: A feature extraction module, configured to extract molecular features based on the first molecular data set to obtain a first feature extraction result; implemented by a first model portion of a molecular property prediction model; A property prediction module is configured to calculate the probability distribution of the feature extraction results in the physicochemical properties of the set molecule through Gaussian process regression based on the first feature extraction results, thereby obtaining a first physicochemical property prediction result; the method is implemented by the second model portion of the molecular property prediction model; The method further includes selecting a plurality of target molecule data from the second molecule data set when the second model partially satisfies the set convergence condition; generating a target molecule training sample based on the plurality of target molecule data; Returning to the step of extracting features from the target molecule training sample based on the pre-trained first model portion to obtain molecular sample features, until the second model portion satisfies the set convergence condition; The selecting of the plurality of target molecular data from the second molecular data set comprises: performing molecular feature extraction on each of the plurality of data sets divided from the second molecular data set based on the first model portion to obtain a second feature extraction result; Based on the second model part, the second feature extraction result is predicted for molecular physicochemical properties to obtain a second physicochemical property prediction result; based on the second molecular property prediction result, a target set is selected from the multiple data sets; and the multiple molecular data contained in the target set are determined as the target molecular data.
12. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
File cross-modal data feature fusion method based on main affinity representation
CN112784017A
Small sample character and freehand sketch recognition method and device
CN113111803A