Molecular property prediction method and device based on chemical bond characterization learning
Through the molecular property prediction method learned by chemical bond characterization, the problem of high time-consuming and insufficient accuracy of molecular property prediction in drug research is solved, and an efficient and accurate molecular property prediction tool is provided, suitable for drug design and environmental pollutant evaluation.
Patent Information
- Application Number
- CN202510562957.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art predicts molecular properties in drug research time-consuming and costly, and traditional molecular representation methods are time-consuming and error-prone, resulting in insufficient accuracy of ML models and lack of efficient molecular representation methods.
The molecular property prediction method based on chemical bond characterization learning is adopted. By standardizing the simplified molecular linear input standard data of the target molecule, it is converted into chemical bond characterization data, and the chemical bond characteristics are classified using predetermined bond classification standards, combining pre-training and fine-tuning models for prediction.
Improves the accuracy and efficiency of molecular properties prediction, simplifies the calculation process, reduces costs, and provides interpretable prediction tools suitable for drug design and environmental pollutant toxicity assessment.
Smart Images

Figure CN120473011A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular property prediction, and in particular to a molecular property prediction method and device based on chemical bond characterization learning. Background Art
[0002] Despite significant progress in drug design and discovery, understanding molecular properties in the early stages of drug research still faces significant challenges. Traditional research methods are time-consuming and costly, particularly when animal experiments are involved. These methods pose ethical and practical challenges. For example, testing a molecule for carcinogenicity requires not only significant financial investment but also histopathological analysis of over 40 animal tissues and organs. Against this backdrop, the application of computational methods, such as machine learning (ML) and deep learning (DL), to molecular property prediction is highly attractive. These methods not only reduce reliance on time-consuming and laborious experiments, significantly reducing costs and time consumption, but also offer high predictive accuracy. Consequently, high-quality molecular property prediction models have become indispensable tools at all stages of drug research, including identifying drug targets, predicting the absorption, distribution, metabolism, excretion, and toxicity (ADMET) properties of new drugs, and evaluating molecular bioactivity.
[0003] In the process of realizing the concept of the present invention, it was found that the accuracy of ML-based models depends largely on how to select appropriate molecular representations. However, the commonly used molecular fingerprint and descriptor design and selection process is time-consuming and error-prone. Therefore, how to quickly select appropriate molecular representations and accurately predict the properties of molecules is crucial. Summary of the Invention
[0004] In view of this, the present invention provides a molecular property prediction method and device based on chemical bond characterization learning.
[0005] One aspect of the present invention provides a molecular property prediction method based on chemical bond representation learning, comprising: standardizing the simplified molecular linear input standard data of the target molecule to obtain standardized data of the target molecule; converting the standardized data into chemical bond representation data based on a predetermined bond classification standard, the predetermined bond classification standard being used to classify the characteristics of the chemical bond; and predicting the chemical properties of the target molecule based on the chemical bond representation data to obtain a prediction result.
[0006] According to an embodiment of the present invention, the characteristics of a chemical bond include the starting / ending atom, bond type, conjugation, stereoisomerism, and ring size of the chemical bond; based on a predetermined bond classification standard, the standardized data is converted into chemical bond characterization data, including: completing the missing hydrogen atoms in the standardized data to obtain standardized data with hydrogen atoms; based on the predetermined bond classification standard, classifying the characteristics of each of the multiple chemical bonds in the order of the multiple chemical bonds existing in the standardized data with hydrogen atoms to obtain the category of the characteristics of each of the multiple chemical bonds; based on a predetermined coding form, determining the coded data that matches the category of the characteristics of each of the multiple chemical bonds in turn; and determining the chemical bond characterization data based on the coded data that matches the category of the characteristics of each of the multiple chemical bonds.
[0007] According to an embodiment of the present invention, chemical bond representation data is determined based on encoding data that matches the categories of characteristics of each of the multiple chemical bonds, including: encoding the multiple chemical bonds in sequence according to the encoding data that matches the categories of characteristics of the multiple chemical bonds to obtain the encoding data of each of the multiple chemical bonds; determining target character strings that match the encoding data of each of the multiple chemical bonds in sequence based on a predetermined dictionary, wherein the predetermined dictionary represents a mapping relationship between the encoding data of the chemical bonds and the character strings; and combining the target character strings that match the encoding data of each of the multiple chemical bonds in sequence to obtain the chemical bond representation data.
[0008] According to an embodiment of the present invention, when the characteristics of the chemical bonds are different, the predetermined coding forms are different.
[0009] According to an embodiment of the present invention, the chemical bond characterization data includes at least a starting concatenated string and a target string that matches the encoding data of each of the multiple chemical bonds, and the starting concatenated string is used to connect the target strings; based on the chemical bond characterization data, the chemical properties of the target molecule are predicted to obtain a prediction result, including: inputting the chemical bond characterization data into a prediction model, and outputting the prediction result at the starting concatenated string, wherein the prediction model is pre-trained by a pre-trained model and a fine-tuning model, and the pre-trained parameters of the pre-trained model are used for fine-tuning the fine-tuning model.
[0010] According to an embodiment of the present invention, a pre-training method includes: obtaining a pre-training data set and a chemical property data set, wherein the pre-training data set includes a plurality of unlabeled sample simplified molecular linear input standard data, and the chemical property data set includes a plurality of sample simplified molecular linear input standard data with molecular chemical property labels; performing standardization processing on the sample data in the pre-training data set and the chemical property data set, respectively, to obtain a standardized pre-training data set and a standardized chemical property data set; based on a predetermined bond classification standard, converting the sample data in the standardized pre-training data set and the standardized chemical property data set into sample chemical bond representation data, respectively, to obtain a converted pre-training data set and a converted chemical property data set;
[0011] A predetermined proportion of sample chemical bond representation data is randomly selected from the converted pre-training data set, and the pre-training model is pre-trained to obtain pre-training parameters, wherein a first proportion of the sample chemical bond representation data of the predetermined proportion is masked, a second proportion of the sample chemical bond representation data is replaced, and a third proportion of the sample chemical bond representation data remains unchanged, and the sum of the first proportion, the second proportion and the third proportion is 100%; the pre-training parameters are determined as the initial parameters of the fine-tuning model; based on the converted chemical property data set, the initial parameters of the fine-tuning model are fine-tuned to obtain a prediction model.
[0012] According to an embodiment of the present invention, the sample chemical bond characterization data includes at least a starting splicing node symbol and sample coding data that matches the category of the sample characteristics of each of the multiple sample chemical bonds of the sample molecule, and the starting splicing node symbol is used to connect the sample coding data; based on the converted chemical property data set, the initial parameters of the fine-tuning model are fine-tuned to obtain a prediction model, including: determining a training set and a validation set from the converted chemical property data set according to a preset ratio; based on the validation set, screening multiple groups of preset hyperparameters to obtain target hyperparameters; based on the training set and the target hyperparameters, fine-tuning the initial parameters of the fine-tuning model to obtain a prediction model.
[0013] According to an embodiment of the present invention, after fine-tuning the initial parameters of the fine-tuning model, the method further includes: performing a performance evaluation on the fine-tuning model based on a test set to obtain an evaluation result, wherein the test set is determined from the converted chemical property data set; and when it is determined that the evaluation result meets the preset conditions, determining the fine-tuning model after fine-tuning as a prediction model.
[0014] According to an embodiment of the present invention, the molecular property prediction method based on chemical bond characterization learning also includes: removing sample data of predetermined chemical substances in the pre-training data set and the chemical property data set and then performing standardization processing, wherein the predetermined chemical substances include at least one of the following: mixtures, inorganic substances and molecules whose number of heavy atoms does not meet the threshold.
[0015] Another aspect of the present invention provides a molecular property prediction device based on chemical bond representation learning, including: a processing module for standardizing the simplified molecular linear input standard data of the target molecule to obtain standardized data of the target molecule; a conversion module for converting the standardized data into chemical bond representation data based on a predetermined bond classification standard, and the predetermined bond classification standard is used to classify the characteristics of the chemical bond; a prediction module for predicting the chemical properties of the target molecule based on the chemical bond representation data to obtain a prediction result.
[0016] According to an embodiment of the present invention, the SMILES data of a target molecule is converted into data represented by chemical bonds. This in-depth representation of the molecule, using chemical bond representation, fully explores the chemical information of the molecule from the perspective of chemical bonds. Furthermore, the chemical bond representation is used to interpret the mechanism of a molecule's properties at a non-atomic level, thereby improving the accuracy of molecular property prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0018] Figure 1 A flowchart of a molecular property prediction method based on chemical bond characterization learning according to an embodiment of the present invention is shown;
[0019] Figure 2 A schematic diagram of encoding multiple chemical bonds according to an embodiment of the present invention is shown;
[0020] Figure 3 A schematic diagram of the structure of a prediction model according to an embodiment of the present invention is shown;
[0021] Figure 4 A flowchart of a pre-training method according to an embodiment of the present invention is shown;
[0022] Figure 5 A block diagram of a molecular property prediction device based on chemical bond representation learning according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0023] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0024] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0025] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0026] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0027] During the implementation of the present invention, we discovered that molecular representation is the cornerstone of molecular property prediction, and that a reasonable molecular representation and an appropriate computational model are crucial for prediction. Molecular representation methods primarily fall into three categories: those based on descriptors / molecular fingerprints, molecular graphs, and Simplified Molecular Input Line Entry System (SMILES) sequences.
[0028] Molecular representations based on descriptors / molecular fingerprints encode structural information, pharmacophore characteristics, and physicochemical properties of a molecule into fixed-length feature vectors. Examples include Extended Connectivity Fingerprint (ECFP), Molecular Access System (MACCS), RDKit, and Morgan. Descriptors / molecular fingerprints, such as these, are widely used in conjunction with machine learning methods, such as molecular fingerprints and gradient boosting ensemble learning algorithms (ECFPs–XGBoost) and molecular fingerprints and random forests (Morgan–Random Forest). However, these expert-developed descriptors / molecular fingerprints are subjective, containing molecular information that may only be applicable to a single or a few specific tasks, and their generation relies heavily on the expert's prior knowledge.
[0029] Molecular graph methods differ from descriptors and molecular fingerprints in that they represent chemical bonds as edges and atoms as nodes, forming a topological structure that is well-suited for DL methods such as graph neural networks (GNNs). GNNs automatically extract atomic and bond features from this topological structure, avoiding the incomplete representation issues that can result from handcrafted descriptors and molecular fingerprints. However, due to overfitting and oversmoothing issues, most existing GNNs have only two or three layers, limiting their ability to extract deep molecular features. Furthermore, GNNs rely on large datasets to obtain optimal weight parameters. If the dataset is too small, their performance may be inferior to traditional descriptor and molecular fingerprint-based prediction methods.
[0030] Generally speaking, a common challenge faced by DL in molecular property prediction is the scarcity of labeled data. Models typically require large amounts of labeled data to achieve efficiency and generalization. In fields such as chemistry, biology, and medicine, obtaining labeled data requires expensive and time-consuming laboratory experiments, resulting in extremely limited labeled data and severely impacting model performance.
[0031] In the field of natural language processing (NLP), similar problems have been addressed through the introduction of pre-trained language models (LMs), such as BERT (Bidirectional Encoder Representations from Transformer) and Generative Pretrained Transformer (GPT). For molecules, SMILES representation, which is essentially a sequence of characters, has laid the foundation for NLP research on molecular property prediction. Furthermore, large public chemical datasets such as PubChem, ChEMBL, and ZINC provide SMILES representations of a large number of molecules, ensuring the feasibility of NLP applications in this field. However, existing models mostly focus on the characters (atoms) that make up the SMILES and lack a deep representation of the SMILES, resulting in information loss.
[0032] The accuracy of conventional ML-based QSAR / QSPR (Quantitative Structure-Activity Relationship / Quantitative Structure-Property Relationship) models depends heavily on the selection of appropriate molecular representations. Specifically, conventional ML methods require chemists to manually develop a set of rules to encode relevant structural information, pharmacophore features, and / or physicochemical properties of a molecule into fixed-length vectors, such as commonly used molecular fingerprints and descriptors. However, the design and selection process is time-consuming and error-prone, resulting in poor scalability and versatility of molecular descriptors, which in turn leads to certain limitations in the application of descriptor-based ML models in the field of molecular property prediction.
[0033] Therefore, a molecular property prediction method based on chemical bond representation learning is proposed. This method improves prediction accuracy while maintaining simplicity and saving computational time and resources. Considering that molecular structure determines its properties, molecules are composed of chemical bonds and atoms, and chemical bonds contain the atoms that make up the molecule. Representing molecules with chemical bonds may contain more chemical information. Therefore, using deep chemical bond representation provides another perspective to interpret molecular properties and increase model interpretability.
[0034] Embodiments of the present invention provide a molecular property prediction method and apparatus based on chemical bond representation learning. The method comprises: applying standardized processing to simplified molecular linear input standard data of a target molecule to obtain standardized data of the target molecule; converting the standardized data into chemical bond representation data based on a predetermined bond classification standard, wherein the predetermined bond classification standard is used to classify chemical bond characteristics; and predicting the chemical properties of the target molecule based on the chemical bond representation data to obtain a prediction result.
[0035] The following will be passed Figures 1 to 3 The molecular property prediction method based on chemical bond representation learning according to an embodiment of the present invention is described in detail.
[0036] Figure 1 A flowchart of a molecular property prediction method based on chemical bond characterization learning according to an embodiment of the present invention is shown.
[0037] like Figure 1 As shown, the molecular property prediction method 100 based on chemical bond representation learning includes operations S110 to S130.
[0038] In operation S110 , the simplified molecular linearity of the target molecule is input into the standard data for standardization processing to obtain the standardized data of the target molecule.
[0039] In operation S120 , the normalized data is converted into chemical bond characterization data based on a predetermined bond classification standard.
[0040] In operation S130 , the chemical properties of the target molecule are predicted based on the chemical bond characterization data to obtain a prediction result.
[0041] According to an embodiment of the present invention, the target molecule may be a chemical molecule for which molecular property prediction is to be performed. A cheminformatics toolkit may be used to normalize the simplified molecular linear input normalization data of the target molecule to obtain normalized data of the target molecule that uniquely corresponds to a molecular structure.
[0042] According to an embodiment of the present invention, the normalization process may include, but is not limited to: removing charges, removing small fragments, processing tautomerism, etc.
[0043] According to an embodiment of the present invention, a predetermined bond classification standard may be used to classify characteristics of chemical bonds.
[0044] For example, the features of chemical bonds in the standardized data may be classified according to a predetermined bond classification standard, and then the categories of the classified features of the chemical bonds in the standardized data may be represented to obtain chemical bond characterization data.
[0045] According to embodiments of the present invention, chemical bond characterization data can be vectorized and then, using a multi-head attention mechanism, segmented into multiple subspaces. Attention weights are independently calculated for each subspace, and the weighted summation of the subspace value matrices is performed to obtain feature information of the chemical bond characterization data. Based on this feature information, a prediction result is obtained. The prediction result represents the chemical properties of the target molecule, which may include, but are not limited to, respiratory toxicity and activity.
[0046] According to an embodiment of the present invention, the SMILES data of a target molecule is converted into data represented by chemical bonds. This in-depth representation of the molecule, using chemical bond representation, fully explores the chemical information of the molecule from the perspective of chemical bonds. Furthermore, the chemical bond representation is used to interpret the mechanism of a molecule's properties at a non-atomic level, thereby improving the accuracy of molecular property prediction.
[0047] According to an embodiment of the present invention, the characteristics of a chemical bond may include the starting / ending atom of the chemical bond, bond type, conjugation, stereoisomerism, and ring size. Figure 1 The operation S120 shown converts the standardized data into chemical bond representation data based on a predetermined bond classification standard, and may include: completing the missing hydrogen atoms in the standardized data to obtain standardized data with hydrogen atoms; based on the predetermined bond classification standard, classifying the respective characteristics of the multiple chemical bonds in sequence according to the order of the multiple chemical bonds existing in the standardized data with hydrogen atoms to obtain the categories of the respective characteristics of the multiple chemical bonds; based on a predetermined coding form, determining the coded data that matches the categories of the respective characteristics of the multiple chemical bonds in sequence; and determining the chemical bond representation data based on the coded data that matches the categories of the respective characteristics of the multiple chemical bonds.
[0048] According to an embodiment of the present invention, the chemical molecules in the standardized data are represented by carbon atoms and non-hydrogen atoms. The bond types may include single bonds, double bonds, triple bonds, aromatic bonds, chemical bonds that exist objectively but are not defined in the target molecule, etc. For example, the chemical bonds that exist objectively but are not defined in the target molecule can be represented by <unk>express.
[0049] According to an embodiment of the present invention, when the characteristics of the chemical bonds are different, the predetermined coding forms are different.
[0050] For example, when determining that the characteristic of a chemical bond is the start / end atom, the predetermined coding form can be the coding form of the atom of the start / end atom, for example, the coding form of the start / end atom of a C-H bond can be "C-H". When determining that the characteristic of a chemical bond is the bond type or conjugation, the predetermined coding form is a one-hot coding form. For example, when the chemical bond type is a single bond, the coding for the chemical bond type can be "1000", "1" indicates that the chemical bond type is a single bond, and "0" indicates that the chemical bond type is not a double bond, triple bond, or aromatic bond. The coding with conjugation can be "1". The coding without conjugation can be "0". When determining that the characteristic of a chemical bond is stereoisomerism, the predetermined coding form is a character coding form, for example, the coding without stereoisomerism can be "STEREONONE". When determining that the characteristic of a chemical bond is ring size, it can be coded in a digital form representing the ring size, for example, the maximum value of the ring size can be 7, a ring greater than 7 is represented as the maximum value, and no ring can be represented as empty.
[0051] According to an embodiment of the present invention, the categories of the characteristics of each of a plurality of chemical bonds are determined by a predetermined bond classification standard; and then, the characteristics of chemical bonds of different categories are encoded in different forms, which is conducive to clearly distinguishing the characteristics of the chemical bonds in the target molecule and fully mining the chemical information of the molecule.
[0052] According to an embodiment of the present invention, determining chemical bond characterization data based on encoding data that matches the categories of characteristics of each of the multiple chemical bonds may include: encoding the multiple chemical bonds in sequence according to the encoding data that matches the categories of characteristics of the multiple chemical bonds to obtain the encoding data of each of the multiple chemical bonds; determining target character strings that match the encoding data of each of the multiple chemical bonds in sequence based on a predetermined dictionary; and combining the target characterization strings that match the encoding data of each of the multiple chemical bonds in sequence to obtain the chemical bond characterization data.
[0053] Figure 2 A schematic diagram of encoding multiple chemical bonds according to an embodiment of the present invention is shown.
[0054] For example, Figure 2 As shown in the figure, taking the SMILES data of methyl butyrate, CCCC(=O)OC, as an example, CCCC(=O)OC is a standardized SMILES. CCCC(=O)OC has no hydrogen atoms and requires hydrogen addition. All chemical bonds in methyl butyrate can then be numbered in a specific order. For example, the chemical bonds in CCCC(=O)OC, which do not contain hydrogen atoms, are first numbered: the bond between the first and second C atoms is numbered 0, the bond between the second and third C atoms is numbered 1, and so on up to 5. Next, the first C atoms and the connected hydrogen atoms are numbered. There are three hydrogen atoms connected to form three C-H bonds, numbered 6 / 7 / 8. Similarly, the C-H bond formed by the second C atoms is numbered 9 / 10. Thus, the number list for the chemical bonds formed by the first C atoms is [0, 6, 7, 8], and the number list for the chemical bonds formed by the second C atoms is [0, 1, 9, 10]. Concatenating the bond number lists for all atoms in methyl butyrate gives the complete bond number list: [0, 6, 7, 8, 0, 1, 9, 10, 1, 2, 11, 12, 2, 3, 4,3, 4, 5, 5, 13, 14, 15, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15]. Multiple chemical bonds can be encoded in sequence according to the chemical bond number list. For example, the first C atom has four chemical bonds 0 / 6 / 7 / 8. According to the predetermined bond classification standard, 0 / 6 / 7 / 8 will be represented as 'CC, [1, 0, 0, 0], 0, STEREONONE,','CH, [1, 0, 0, 0], 0, STEREONONE,','CH, [1, 0, 0, 0], 0, STEREONONE,', 'CH, [1, 0,0, 0], 0, STEREONONE,' and their starting atom is the first C. When traversing to the hydrogen atom connected to the first C atom, 6 / 7 / 8 will be represented as 'HC, [1, 0, 0, 0], 0, STEREONONE,', and it can be seen that the hydrogen atom is used as the starting atom. Although 'HC, [1, 0, 0, 0], 0, STEREONONE,' and 'CH, [1, 0,0, 0], 0, STEREONONE,' are actually the same chemical bond, they are distinguished and considered to be different chemical bonds when generating the encoding data for each of the multiple chemical bonds.
[0055] According to an embodiment of the present invention, a predetermined dictionary can represent a mapping relationship between chemical bond encoding data and character strings. For example, based on an unlabeled SMILES dataset, SMILES data in the SMILES dataset can be converted to obtain chemical bond encoding data. The chemical bond encoding data is used as a key, and a character string that has a mapping relationship with the chemical bond encoding data is used as a value, which is stored in the predetermined dictionary.
[0056] According to embodiments of the present invention, the encoding data for multiple chemical bonds can accurately reflect the connection patterns and interaction characteristics between atoms within a molecule. For example, properties such as the activity and toxicity of a molecule are often closely related to the type, number, and spatial arrangement of its chemical bonds. By encoding the chemical bonds and matching them with target strings, the prediction model can accurately extract chemical bond information related to these properties, thereby improving the accuracy of the prediction model.
[0057] According to an embodiment of the present invention, the chemical bond representation data may include at least a starting concatenated string and target strings matching respective encoding data of a plurality of chemical bonds, wherein the starting concatenated string is used to connect the target strings.
[0058] According to an embodiment of the present invention, the starting concatenation string may be located before the start / end atom of the chemical bond. For example, the starting concatenation string may be: <glo>.
[0059] According to an embodiment of the present invention, the chemical bond characterization data may further include a placeholder string, which is used to fill the chemical bond characterization data of different lengths to the same length. In the case where the lengths of the chemical bond characterization data are inconsistent, {' <pad>': 0} to fill chemical bond representation data of different lengths to the same length. <pad>Used to pad chemical bond representation data of different lengths to the same length.
[0060] Figure 3 A schematic structural diagram of a prediction model according to an embodiment of the present invention is shown.
[0061] According to an embodiment of the present invention, Figure 1 Operation S130 shown, predicting the chemical properties of the target molecule based on the chemical bond characterization data to obtain a prediction result, may include the operations of inputting the chemical bond characterization data into a prediction model and outputting a prediction result at the starting concatenated string.
[0062] According to an embodiment of the present invention, the prediction model is pre-trained by a pre-trained model and a fine-tuning model, and the pre-trained parameters of the pre-trained model are used for fine-tuning the fine-tuning model.
[0063] For example, Figure 3 As shown, the structures of the pre-trained model and the fine-tuning model are similar, and both can include an embedding layer and a transformer layer. The difference is that the pre-trained model has a pre-training layer after the transformer layer, while the fine-tuning model has a prediction layer after the transformer layer. That is, the embedding layer and the transformer layer are shared in the pre-training and fine-tuning stages. The pre-training / prediction layers in the pre-training and fine-tuning stages are not shared. The pre-training layer can be, for example, a multi-classification fully connected layer that can perform pre-training and output prediction results at multiple locations of randomly selected chemical bond representation data. The prediction layer can be, for example, a fully connected layer, or a one-dimensional convolutional layer, an adaptive pooling layer, and a fully connected layer. When the prediction layer is a fully connected layer, it can be used to perform fine-tuning regression tasks and output prediction results at the starting concatenation string of the chemical bond representation data. When the prediction layer is a one-dimensional convolutional layer, an adaptive pooling layer, and a fully connected layer, it can be used to perform fine-tuning classification tasks and output prediction results at the starting concatenation string of the fine-tuning input. The transformer layer can be composed of a multi-head attention layer, normalization and residual connections, a fully connected layer, and normalization and residual connections. The number of heads in the multi-head attention layer can be 8.
[0064] For example, in the embedding layer, we can use the embedding dictionary , maps the pre-trained input or fine-tuning input A={a1,a2,a3,…,aN} to the embedding vector w={w1,w2,w3,…,wN}, where the embedding vector , V is the dictionary size, d model is the embedding vector size. Since pre-trained models such as the BERT model cannot automatically learn position information, it is necessary to add a predefined position encoding vector to each embedding vector in the embedding layer, and fuse the predefined position encoding vector and the embedding vector to obtain the input vector X. In the multi-head self-attention layer, X is mapped to the query vector , keyword vector , value vector , i=1,2,3,…,8, represents the number of heads in the multi-head attention layer, W i It is necessary to learn parameters, and the output of each self-attention layer can be , T is transpose, d k Therefore, d k =d model / 8. The information from different heads is mixed by connecting and linearly projecting again. After the attention output, a fully connected neural network is superimposed with an activation function in the middle. In addition, all self-attention layers and fully connected layers are followed by a normalization layer and a residual connection, which can improve generalization ability and better utilize chemical bond information.
[0065] For example, you can use cross entropy as the loss function for pre-training, use binary cross entropy with Sigmoid as the loss function for fine-tuning classification tasks, and use mean square error MSE as the loss function for fine-tuning regression tasks. The activation function used for the pre-trained model can be GELU, and the activation function used for fine-tuning classification and regression tasks is LeakyReLU. The pre-trained model can obtain prediction results at randomly selected chemical bonds, while the fine-tuned model can obtain prediction results at randomly selected chemical bonds. <glo>Get the prediction results.
[0066] According to embodiments of the present invention, a molecular representation approach distinct from molecular fingerprints and descriptors is combined with a model framework to improve the performance of molecular chemical property prediction. This predictive model can rapidly predict the chemical properties of molecules, providing an out-of-the-box, effective, and interpretable prediction tool with broad application prospects in areas such as drug design and development and environmental pollutant toxicity assessment.
[0067] Figure 4 A flowchart of a pre-training method according to an embodiment of the present invention is shown.
[0068] like Figure 4 As shown, the pre-training method includes operations S410 to S460.
[0069] In operation S410 , a pre-training dataset and a chemical property dataset are acquired.
[0070] In operation S420 , the sample data in the pre-training dataset and the chemical property dataset are standardized to obtain a standardized pre-training dataset and a standardized chemical property dataset.
[0071] In operation S430 , based on a predetermined bond classification standard, sample data in the standardized pre-training dataset and the standardized chemical property dataset are respectively converted into sample chemical bond representation data to obtain a converted pre-training dataset and a converted chemical property dataset.
[0072] In operation S440 , a predetermined proportion of sample chemical bond representation data is randomly selected from the converted pre-training data set, and a pre-training model is pre-trained to obtain pre-training parameters.
[0073] In operation S450 , the pre-trained parameters are determined as initial parameters of the fine-tuning model.
[0074] In operation S460 , the initial parameters of the fine-tuning model are fine-tuned based on the converted chemical property dataset to obtain a prediction model.
[0075] According to an embodiment of the present invention, the pre-training data set may include a plurality of unlabeled sample simplified molecular linear input normative data, and the chemical property data set may include a plurality of sample simplified molecular linear input normative data with molecular chemical property labels.
[0076] For example, to classify whether a molecule has respiratory toxicity, a large amount of unlabeled sample simplified molecular linear input canonical data can be collected from datasets such as ChEMBL and PubChem as a pre-training dataset, totaling approximately 1.84 million data points. Chemical property datasets, such as the respiratory toxicity dataset, can be downloaded from target websites. This dataset consists of sample simplified molecular linear input canonical data for 1,041 compounds and molecular chemical property labels, such as binary classification labels, where label '0' can indicate non-toxicity and label '1' can indicate toxicity.
[0077] The simplified molecular linear input normative data of the unlabeled samples in the pre-training dataset can be normalized to obtain a standardized pre-training dataset. The simplified molecular linear input normative data of the samples in the chemical property dataset can be normalized to obtain a standardized chemical property dataset.
[0078] Based on the predetermined bond classification standard, tools such as RDKit can be used to convert the sample data in the standardized pre-training dataset and the standardized chemical property dataset into sample chemical bond representation data, thereby obtaining the converted pre-training dataset and the converted chemical property dataset. Figure 1 The implementation method of operation S120 is not described in detail here.
[0079] A predetermined proportion of sample chemical bond representation data can be randomly selected from the converted pre-training dataset, and the pre-training model can be pre-trained to obtain pre-training parameters. Among the predetermined proportions of sample chemical bond representation data, a first proportion of sample chemical bond representation data is masked, a second proportion of sample chemical bond representation data is replaced, and a third proportion of sample chemical bond representation data remains unchanged, and the sum of the first proportion, the second proportion, and the third proportion is 100%.
[0080] For example, 15% of the sample chemical bond representation data can be randomly selected from the converted pre-training data set for prediction, and 80% of the sample chemical bond representation data in the selected 15% sample chemical bond representation data are predicted using mask symbols such as <mask>Masking is performed, 10% of the sample chemical bond characterization data are randomly replaced with other sample chemical bond characterization data, and the remaining 10% of the sample chemical bond characterization data remain unchanged.
[0081] A predetermined percentage of sample chemical bond representation data can be randomly selected and divided into a training set and a test set for a pretrained model, such as a BERT model, in a ratio of, for example, 9:1. The training set is used to pretrain the BERT model to obtain parameters, while the test set is used during the pretraining process to mask and restore chemical bonds in the randomly selected sample chemical bond representation data, with the restoration rate used as an evaluation metric. If all masked bonds are correctly predicted, the loss is minimized and reaches zero. If only some masked bonds are correctly restored, a loss is incurred. The training set is used to train the pretrained model using the batch gradient descent algorithm and optimizer (AdamW). For example, the learning rate of the pretrained model can be set to 1e-4, the batch size to 150, and the model is pretrained for five epochs, with a model pretraining restoration rate of 99.2%. The pretraining parameters from the fifth epoch are saved.
[0082] The pre-trained parameters can be used as the initial parameters of the fine-tuning model. In order to obtain the optimal fine-tuning model, the prediction results can be obtained at the starting splicing string based on the pre-trained parameters, and then the hyperparameter search can be performed based on the prediction results and the molecular chemical property labels. The model is retrained using the optimal hyperparameter combination, and after verification, the optimal model is saved to obtain the prediction model.
[0083] According to an embodiment of the present invention, a pre-trained model is trained by converting a plurality of unlabeled sample simplified molecular linear input specification data into sample chemical bond representation data. On the basis of the pre-trained model parameters, the simplified molecular linear input specification data of a plurality of samples labeled with molecular chemical properties are further converted into sample chemical bond representation data, so that the fine-tuning model can obtain more accurate prediction results at the starting splicing string, thereby improving the training speed and the accuracy of the model.
[0084] According to an embodiment of the present invention, the sample chemical bond characterization data may include at least sample encoding data matching the category of the sample characteristics of each of the multiple sample chemical bonds of the sample molecule, and the starting splicing node symbol is used to connect the sample encoding data. Based on the converted chemical property data set, fine-tuning the initial parameters of the fine-tuning model to obtain a prediction model may include the following operations: determining a training set and a validation set from the converted chemical property data set according to a preset ratio; based on the validation set, screening multiple sets of preset hyperparameters to obtain target hyperparameters; and fine-tuning the initial parameters of the fine-tuning model based on the training set and the target hyperparameters to obtain a prediction model.
[0085] According to an embodiment of the present invention, before training a fine-tuning model, multiple sets of preset hyperparameters can be preset for the fine-tuning model. This allows these sets of preset hyperparameters to be optimized using a validation set to obtain the optimal hyperparameters, i.e., the target hyperparameters. These multiple sets of preset hyperparameters can include the learning rate (learning_rate), weight decay rate (weight_decay), number of frozen layers (num_layers), dropout rate (dropout_rate), and random seed (seed) for validation set partitioning during training. During model training, parameters for frozen layers in the model are not updated. Taking the dropout rate into account can reduce model overfitting.
[0086] According to an embodiment of the present invention, after fine-tuning the initial parameters of the fine-tuning model, the method further includes: performing a performance evaluation on the fine-tuning model based on a test set to obtain an evaluation result, wherein the test set is determined from the converted chemical property data set; and when it is determined that the evaluation result meets the preset conditions, determining the fine-tuning model after fine-tuning as a prediction model.
[0087] For example, the converted chemical property dataset can be divided into a training set, a validation set, and a test set in a ratio of 8:1:1. The training set is used to train the fine-tuned model, and the validation set is used to optimize the preset hyperparameters of the fine-tuned model to obtain the optimal fine-tuned model. During the hyperparameter optimization process, for each set of hyperparameters, the fine-tuned model is trained using the training set and validated using the validation set. The optimal hyperparameters are then determined based on statistical parameters. For example, the optimal hyperparameter combination may be: learning_rate = 0.00092, weight_decay = 3.585e-05, num_layer = 97, dropout_rate = 0.1, and seed = 7. The selected statistical parameter may be the area under the curve (ROC-AUC) of the receiver operating characteristic curve. The fine-tuned model is retrained on the training set using the optimal hyperparameter combination, and the optimal model is saved based on the statistical parameters of the validation set to serve as the predictive model for determining whether a molecule has respiratory toxicity. In addition, the fine-tuning model also uses an early stopping strategy to avoid overfitting during training, where the tolerance value setting can be, for example, 30 and the maximum iteration round setting can be, for example, 300.
[0088] The test set can be input into the prediction model for respiratory toxicity prediction. Based on the chemical bond characterization data and molecular chemical property markers of the test set samples, the ROC-AUC value was calculated to be 0.898. The validation set was input into the prediction model for respiratory toxicity prediction, and the resulting ROC-AUC value was 0.910. It can be seen that the ROC-AUC value of the prediction model on the test set is comparable to that on the validation set. Other prediction models such as the extended connectivity fingerprint-extreme gradient boosting (ECFP-XGBoost) model, the improved graph neural network model (HRGCN+), Attentive FP, Aug-BERT, Cano-BERT, K-BERT, etc. obtained ROC-AUC values of 0.790, 0.855, 0.847, 0.878, 0.843, and 0.873 on the data set, respectively. The prediction model obtained by the present invention has a higher ROC-AUC value than other prediction models, indicating that the prediction model of the present invention has excellent generalization performance and prediction performance.
[0089] It should be noted that for regression tasks, the statistical parameter can be the root mean square error (RMSE).
[0090] According to an embodiment of the present invention, the present invention is based on a molecular representation method characterized by chemical bonds, combined with a prediction model obtained by training a pre-trained model and a fine-tuned model, which can self-supervise molecular representation learning and property prediction, thereby improving the performance of molecular property prediction.
[0091] According to another embodiment of the present invention, the molecular property prediction method based on chemical bond characterization learning may include the following: Figure 1 In addition to the operations S110 to S130 shown, the operation may also include: removing sample data of predetermined chemical substances from the pre-training dataset and the chemical property dataset and then performing standardization processing.
[0092] According to an embodiment of the present invention, the predetermined chemical substance includes at least one of the following: a mixture, an inorganic substance, and a molecule whose number of heavy atoms does not meet a threshold.
[0093] According to an embodiment of the present invention, by removing sample data of predetermined chemical substances from the pre-training dataset and the chemical property dataset, the influence of predetermined chemical substances on the property prediction of sample molecules can be reduced, thereby improving the accuracy of the training model.
[0094] Figure 5 A block diagram of a molecular property prediction device based on chemical bond representation learning according to an embodiment of the present invention is shown.
[0095] like Figure 5 As shown, the molecular property prediction device 500 based on chemical bond representation learning includes a processing module 510 , a conversion module 520 and a prediction module 530 .
[0096] The processing module 510 is used to perform standardization processing on the simplified molecular linear input of the target molecule into the standard data to obtain the standardized data of the target molecule.
[0097] The conversion module 520 is used to convert the standardized data into chemical bond representation data based on a predetermined bond classification standard, where the predetermined bond classification standard is used to classify the characteristics of chemical bonds.
[0098] The prediction module 530 is used to predict the chemical properties of the target molecule based on the chemical bond characterization data to obtain a prediction result.
[0099] According to embodiments of the present invention, any multiple modules among the processing module 510, conversion module 520, and prediction module 530 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of the processing module 510, conversion module 520, and prediction module 530 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of software, hardware, and firmware, or any suitable combination of these. Alternatively, at least one of the processing module 510, conversion module 520, and prediction module 530 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0100] It should be noted that the molecular property prediction device part based on chemical bond representation learning in the embodiment of the present invention corresponds to the molecular property prediction method part based on chemical bond representation learning in the embodiment of the present invention. The description of the molecular property prediction device part based on chemical bond representation learning specifically refers to the molecular property prediction method part based on chemical bond representation learning, which will not be repeated here.
[0101] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.
[0102] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which are intended to fall within the scope of the present invention.< / mask> < / glo> < / pad> < / pad> < / glo> < / unk>
Claims
1. A molecular property prediction method based on chemical bond representation learning, characterized in that: The method comprises: Inputting the simplified molecular linearity of the target molecule into the standard data for standardization processing to obtain the standardized data of the target molecule; converting the standardized data into chemical bond characterization data based on a predetermined bond classification standard, wherein the predetermined bond classification standard is used to classify characteristics of chemical bonds; Based on the chemical bond characterization data, the chemical properties of the target molecule are predicted to obtain a prediction result.
2. The method according to claim 1, characterized in that The characteristics of the chemical bond include the start / end atoms of the chemical bond, bond type, conjugation, stereoisomerism, and ring size; The converting the standardized data into chemical bond characterization data based on a predetermined bond classification standard comprises: Performing a completion operation on the missing hydrogen atoms in the standardized data to obtain standardized data with hydrogen atoms; Based on the predetermined bond classification standard, classifying the characteristics of the plurality of chemical bonds in sequence according to the order of the plurality of chemical bonds present in the standardized data having hydrogen atoms to obtain a category of the characteristics of the plurality of chemical bonds; Based on a predetermined coding form, sequentially determining coding data that matches the category of the characteristics of each of the plurality of chemical bonds; The chemical bond characterization data is determined based on the coded data matched to the categories of the characteristics of each of the plurality of chemical bonds.
3. The method according to claim 2, characterized in that The determining of the chemical bond characterization data based on the coded data matched with the categories of the characteristics of each of the plurality of chemical bonds comprises: Encoding the plurality of chemical bonds in sequence according to the encoding data matching the categories of the characteristics of the respective chemical bonds to obtain encoding data of the respective chemical bonds; Determining, in sequence, target character strings that match the encoding data of each of the plurality of chemical bonds based on a predetermined dictionary, wherein the predetermined dictionary represents a mapping relationship between the encoding data of the chemical bonds and the character strings; Target character strings matching the respective encoding data of the plurality of chemical bonds are sequentially combined to obtain the chemical bond representation data.
4. The method according to claim 2, characterized in that When the characteristics of the chemical bonds are different, the predetermined coding forms are different.
5. The method according to claim 1, wherein The chemical bond representation data at least includes a starting concatenated string and a target string matching respective encoding data of a plurality of chemical bonds, wherein the starting concatenated string is used to connect the target string; The step of predicting the chemical properties of the target molecule based on the chemical bond characterization data to obtain a prediction result includes: The chemical bond characterization data is input into a prediction model, and the prediction result at the starting concatenated string is output, wherein the prediction model is pre-trained by a pre-trained model and a fine-tuning model, and the pre-trained parameters of the pre-trained model are used for fine-tuning the fine-tuning model.
6. The method according to claim 5, characterized in that The pre-training method includes: Obtaining a pre-training data set and a chemical property data set, wherein the pre-training data set includes a plurality of unlabeled sample simplified molecular linear input normative data, and the chemical property data set includes a plurality of sample simplified molecular linear input normative data with molecular chemical property labels; Standardizing the sample data in the pre-training dataset and the chemical property dataset respectively to obtain a standardized pre-training dataset and a standardized chemical property dataset; Based on the predetermined bond classification standard, converting the sample data in the standardized pre-training dataset and the standardized chemical property dataset into sample chemical bond representation data, respectively, to obtain a converted pre-training dataset and a converted chemical property dataset; Randomly selecting a predetermined proportion of the sample chemical bond representation data from the converted pre-training data set, and pre-training the pre-training model to obtain the pre-training parameters, wherein a first proportion of the sample chemical bond representation data in the predetermined proportion is masked, a second proportion of the sample chemical bond representation data is replaced, and a third proportion of the sample chemical bond representation data remains unchanged, and the sum of the first proportion, the second proportion, and the third proportion is 100%; Determining the pre-training parameters as initial parameters of the fine-tuning model; Based on the converted chemical property data set, the initial parameters of the fine-tuning model are fine-tuned to obtain the prediction model.
7. The method according to claim 6, characterized in that The sample chemical bond characterization data at least includes a starting splicing node symbol and sample coding data matching the categories of the sample features of the plurality of sample chemical bonds of the sample molecule, wherein the starting splicing node symbol is used to connect the sample coding data; The step of fine-tuning the initial parameters of the fine-tuning model based on the converted chemical property data set to obtain the prediction model comprises: Determining a training set and a validation set from the converted chemical property dataset according to a preset ratio; Based on the validation set, multiple sets of preset hyperparameters are screened to obtain target hyperparameters; Based on the training set and the target hyperparameters, the initial parameters of the fine-tuning model are fine-tuned to obtain the prediction model.
8. The method according to claim 7, characterized in that After fine-tuning the initial parameters of the fine-tuning model, the method further includes: performing a performance evaluation on the fine-tuned model based on a test set to obtain an evaluation result, wherein the test set is determined from the converted chemical property dataset; When it is determined that the evaluation result meets the preset conditions, the fine-tuned fine-tuning model is determined as the prediction model.
9. The method according to any one of claims 6 to 8, characterized in that The method further comprises: Normalization is performed after removing sample data of predetermined chemical substances from the pre-training dataset and the chemical property dataset, wherein the predetermined chemical substances include at least one of the following: mixtures, inorganic substances, and molecules whose heavy atom number does not meet a threshold.
10. A molecular property prediction device based on chemical bond representation learning, characterized in that: The device comprises: a processing module, configured to perform standardization processing on the simplified molecular linear input of the target molecule into the standard data to obtain the standardized data of the target molecule; a conversion module for converting the standardized data into chemical bond characterization data based on a predetermined bond classification standard, wherein the predetermined bond classification standard is used to classify characteristics of chemical bonds; The prediction module is used to predict the chemical properties of the target molecule based on the chemical bond characterization data to obtain a prediction result.
Citation Information
Cited By
Polymer property prediction method and device based on substructure and knowledge enhancement
CN121331270A