A method and apparatus for predicting molecular properties by integrating three-dimensional structure and prior features.
By constructing a molecular property prediction model that integrates three-dimensional structure and prior features, and utilizing big data and advanced network structures to predict the quantum chemical properties of photoelectric molecules, the problems of high computational complexity and long time in existing technologies are solved, and efficient prediction results are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies suffer from high computational complexity, long computation time, and low computational efficiency when analyzing key quantum chemical properties of photoelectric molecules, making it difficult to make efficient predictions.
A molecular property prediction model integrating three-dimensional structure and prior features is constructed. A training dataset for the model is built through big data collection. The Uni-Mol model, molecular feature mapping network, prior feature mapping network, feature fusion layer and multi-task prediction network are used for prediction. The quantum chemical properties of photoelectric molecules are predicted by combining three-dimensional molecular structure and prior property vector.
It reduces prediction complexity, shortens prediction time, and improves prediction efficiency, enabling efficient prediction of various quantum chemical properties of photoelectric molecules.
Smart Images

Figure CN120808952B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for predicting molecular properties by integrating three-dimensional structure and prior features. Background Technology
[0002] Photoelectric molecules are molecules capable of absorbing, converting, or emitting light energy, such as specific molecules in organic photovoltaic materials. Analyzing a series of key quantum chemical properties of photoelectric molecules is a necessary means to understand and optimize their performance. The key quantum chemical properties mentioned here typically include: the transition energy S0 to the first excited singlet state S1 (S0S1), the transition energy S0 to the first excited triplet state T1 (S0T1), the electron reorganization energy (ERE), the vertical ionization potential (VIP), the vertical electron affinity (VEA), the ground state potential energy surface reorganization energy (λ0), and the excited state potential energy surface reorganization energy (λ). + Adiabatic Electron Affinity (AEA), Hole Reorganization Energy (λ) hole This includes the highest occupied molecular orbital-lowest unoccupied molecular orbital gap (HOMO-LUMO gap), etc. Typically, the analysis of these key properties is based on a series of quantum chemical calculations (such as density functional theory calculations). Practical experience shows that these conventional analytical methods often suffer from high computational complexity, long computation time, and low computational efficiency. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by providing a method, apparatus, electronic device, and computer-readable storage medium for predicting molecular properties that integrates three-dimensional structure and prior features. After identifying the set of predicted properties and the set of prior properties, this invention configures a corresponding molecular property prediction model for these two sets. This molecular property prediction model is used to predict multiple types of quantum chemical properties specified in the set of predicted properties based on the three-dimensional molecular structure as input to the model and a prior characteristic vector corresponding to the set of prior properties, and outputs the corresponding property prediction vectors. A model training dataset is constructed through big data collection for model training. After model training, the molecular property prediction model is used to predict multiple types of quantum chemical properties of any photoelectric molecule. This invention can reduce prediction complexity, shorten prediction time, and improve prediction efficiency.
[0004] To achieve the above objectives, a first aspect of the present invention provides a method for predicting molecular properties by integrating three-dimensional structure and prior features, the method comprising:
[0005] A set of predicted properties is obtained by identifying the types of key quantum chemical properties of photoelectric molecules; a set of a priori properties is formed by selecting several other properties that can affect photoelectric conversion efficiency; and a corresponding molecular property prediction model is constructed based on the predicted property set and the priori property set. The predicted properties of the predicted property set include transition energy SOS1, transition energy SOT1, electron recombination energy, vertical ionization energy, ground potential energy surface recombination energy, excited potential energy surface recombination energy, adiabatic electron affinity, hole recombination energy, and the highest occupied orbital-lowest empty orbital gap. The a priori properties of the a priori property set include the total number of aromatic rings, the total number of non-aromatic rings, the proportion of heteroatoms, orbital overlap integral, and molecular dipole moment. The molecular property prediction model is used to predict the key quantum chemical properties of photoelectric molecules based on the three-dimensional molecular structure M and the prior characteristic vector X input to the model and output the corresponding property prediction vector Y. The prior characteristic vector X corresponds to the prior property set, and the property prediction vector Y corresponds to the predicted property set.
[0006] The model training dataset is constructed by collecting big data and denoted as the first dataset; and the molecular property prediction model is trained based on the first dataset;
[0007] After model training is completed, the system receives the first molecular information input by the user; and prepares the model input data based on the first molecular information and the prior property set to obtain the corresponding three-dimensional molecular structure M and the prior property vector X; and inputs the current three-dimensional molecular structure M and the prior property vector X into the molecular property prediction model to predict the corresponding property prediction vector Y and feeds it back to the current user; the first molecular information is a one-dimensional SMILES sequence, two-dimensional molecular topology, or three-dimensional molecular conformation of a photoelectrochemical molecule.
[0008] Preferably, the three-dimensional molecular structure M consists of multiple atomic features m i Composition, 1≤index i≤N A N A The total number of atoms in the current molecular structure; the atomic feature m i Includes atomic element type and atomic three-dimensional coordinates;
[0009] The prior feature vector X consists of multiple feature data x j Composition, 1 ≤ index j ≤ P, where P is the total number of prior properties in the prior property set; the characteristic data x j Each property corresponds one-to-one with the prior properties of the aforementioned set of prior properties;
[0010] The property prediction vector Y is composed of multiple prediction data y k Composition, 1 ≤ index k ≤ N, where N is the total number of predicted properties in the predicted property set; the predicted data y k Each of the predicted properties corresponds one-to-one with the predicted property set.
[0011] The first dataset includes multiple first data records; each first data record includes a first training structure, a first training vector, and a first label vector; the data structures of the first training structure, the first training vector, and the first label vector are consistent with the corresponding three-dimensional molecular structure M, the prior property vector X, and the property prediction vector Y; the first label vector consists of N label data. composition.
[0012] Preferably, the first model input terminal of the molecular property prediction model is used to receive the three-dimensional molecular structure M, the second model input terminal is used to receive the prior characteristic vector X, and the model output terminal is used to output the corresponding property prediction vector Y;
[0013] The molecular property prediction model includes a Uni-Mol model, a molecular feature mapping network, a prior feature mapping network, a feature fusion layer, a fusion feature encoder, and a multi-task prediction network; the multi-task prediction network consists of N parallel k-th prediction heads and a prediction output layer; the Uni-Mol model has been pre-trained.
[0014] The input of the Uni-Mol model is connected to the input of the first model, and its output is connected to the input of the molecular feature mapping network. The output of the molecular feature mapping network is connected to the first input of the feature fusion layer. The input of the prior feature mapping network is connected to the input of the second model, and its output is connected to the second input of the feature fusion layer. The output of the feature fusion layer is connected to the input of the fusion feature encoder. The output of the fusion feature encoder is connected to the input of each of the k-th prediction heads of the multi-task prediction network. The output of each of the k-th prediction heads of the multi-task prediction network is connected to one input of the prediction output layer. The output of the prediction output layer of the multi-task prediction network is connected to the model output.
[0015] The Uni-Mol model is used to perform atomic-level feature encoding based on the three-dimensional molecular structure M to obtain the corresponding atomic feature tensor H. M Send to the molecular feature mapping network; the atomic feature tensor H M The shape is N A ×C A C A The atomic feature dimension is preset; the atomic feature tensor H M By N A A length of C A atomic eigenvectors composition;
[0016] The molecular feature mapping network consists of a feature mapping layer and a feature pooling layer. The feature mapping layer is formed by sequentially connecting two or more first linear activation layers. Each first linear activation layer is composed of a corresponding first linear layer and a first nonlinear activation function, where the first nonlinear activation function is a ReLU activation function or a Sigmoid activation function. The molecular feature mapping network is used to map the atomic feature tensor H using the feature mapping layer. M Each of the atomic feature vectors Do a job from C A The feature mapping from a 3D feature space to a C1D feature space yields a mapped feature vector of length C1. And from the obtained N A The mapped feature vectors Form a shape of N A ×C1 mapping feature tensor H proj The feature pooling layer then applies attention pooling to the mapped feature tensor H. proj N for each feature dimension AThe feature data is pooled to obtain corresponding feature pooled data; and the obtained C1 feature pooled data form a corresponding mapping feature vector H1, which is sent to the feature fusion layer; the mapping feature vector H1 is a feature vector of length C1, where C1 is a preset first feature dimension;
[0017] The prior feature mapping network is composed of one or two sequentially connected second linear activation layers. Each second linear activation layer is composed of a corresponding second linear layer and a second nonlinear activation function, wherein the second nonlinear activation function is a ReLU activation function or a Sigmoid activation function. The prior feature mapping network is used to perform feature mapping processing on the prior feature vector X to obtain the corresponding mapped feature vector H2, which is then sent to the feature fusion layer. The mapped feature vector H2 is a feature vector of length C1.
[0018] The feature fusion layer is used to form a corresponding fused feature tensor H from the mapped feature vectors H1 and H2. R Send to the fusion feature encoder; the fusion feature tensor H R The shape is 2×C1;
[0019] The fusion feature encoder is implemented based on the encoder model of the Transformer architecture; the fusion feature encoder is used to process the fusion feature tensor H. R The corresponding encoded feature tensor E is obtained by performing feature encoding processing. R ; and the encoded feature tensor E R The encoded feature vector E1 corresponding to the mapped feature vector H1 is sent to each of the k-th prediction heads of the multi-task prediction network; the encoded feature tensor E R The shape is 2×C1, and it is composed of the encoded feature vector E1 and the encoded feature vector E2; each of the encoded feature vectors E1 and E2 is a feature vector of length C1, and the encoded feature vectors E1 and E2 correspond one-to-one with the mapped feature vectors H1 and H2.
[0020] Each of the k-th prediction heads in the multi-task prediction network is implemented based on a fully connected network; each of the k-th prediction heads is used to perform corresponding property prediction processing based on the encoded feature vector E1 to obtain the corresponding prediction data y. k Send to the prediction output layer;
[0021] The prediction output layer of the multi-task prediction network will obtain N prediction data y. k The corresponding property prediction vector Y is constructed and output.
[0022] Preferably, the dataset used to construct the model training dataset through big data collection is denoted as the corresponding first dataset, and specifically includes:
[0023] A molecular sequence library is obtained by collecting one-dimensional molecular sequences of photoelectric molecules through multiple preset data channels; the multiple data channels include publicly available photoelectric molecule information databases and publicly available technical documents; the molecular sequence library includes multiple first molecular sequences; each first molecular sequence is a SMILES sequence of a photoelectric molecule;
[0024] Based on a preset cheminformatics tool, a corresponding three-dimensional molecular conformation is constructed according to each of the first molecular sequences to obtain the corresponding first molecular conformation; the cheminformatics tool includes RDKit software;
[0025] Based on the aforementioned cheminformatics tools, the total number of aromatic rings, the total number of non-aromatic rings, and the proportion of heteroatoms in each of the first molecular conformations are identified to obtain the corresponding total number of first aromatic rings, the total number of first non-aromatic rings, and the proportion of first heteroatoms. Furthermore, based on a preset quantum chemical calculation tool, the orbital overlap integral, molecular dipole moment, transition energy SOS1, transition energy SOT1, electron recombination energy, vertical ionization energy, ground potential energy surface recombination energy, excited potential energy surface recombination energy, adiabatic electron affinity, hole recombination energy, and highest occupied orbital-lowest empty orbital gap of the first molecular conformation are calculated to obtain the corresponding first orbital overlap integral, first molecular dipole moment, first transition energy SOS1, second transition energy SOT1, first electron recombination energy, first vertical ionization energy, first ground potential energy surface recombination energy, first excited potential energy surface recombination energy, first adiabatic electron affinity, first hole recombination energy, and first highest occupied orbital-lowest empty orbital energy gap. According to the orbital-lowest empty orbital gap; and composed of the total number of the first aromatic rings, the total number of the first non-aromatic rings, the proportion of the first heteroatoms, the first orbital overlap integral, and the first molecular dipole moment corresponding to each of the first molecular conformations, a corresponding a priori characteristic vector X is formed; and composed of the first transition energy S0S1, the second transition energy S0T1, the first electron recombination energy, the first vertical ionization energy, the first ground state energy surface recombination energy, the first excited state energy surface recombination energy, the first adiabatic electron affinity energy, the first hole recombination energy, and the first highest occupied orbital-lowest empty orbital gap corresponding to each of the first molecular conformations, a corresponding property prediction vector Y is formed; the quantum chemical calculation tools include Multiwfn software, Gaussian software, ORCA software, GAMESS software, and NWChem software;
[0026] Each of the first molecular conformations is taken as the corresponding current molecular conformation; the total number of atoms in the current molecular conformation is identified, and the identification result is the corresponding total number of atoms N. AThe element type and three-dimensional coordinates of each atom in the current molecular conformation are identified, and the identification results are used as the corresponding element type and three-dimensional coordinates of the atom. A corresponding atomic feature m is then formed by the element type and three-dimensional coordinates of each atom. i ; and consisting of all the atomic features m corresponding to the current molecular conformation. i To form a corresponding three-dimensional molecular structure M;
[0027] The three-dimensional molecular structure M, the prior characteristic vector X, and the property prediction vector Y corresponding to each of the first molecular conformations are used as the corresponding first training structure, the first training vector, and the first label vector to form a corresponding first data record; and all the obtained first data records are deduplicated; and all the remaining first data records after deduplication are used to form the corresponding first dataset.
[0028] Preferably, training the molecular property prediction model based on the first dataset specifically includes:
[0029] Step 51: Set an overall loss value range, N classification loss value ranges, and an evaluation value range; and form the corresponding current parameter set by the model parameters of the molecular feature mapping network, the prior feature mapping network, the fusion feature encoder, and the multi-task prediction network of the molecular property prediction model;
[0030] Step 52: Based on a preset first segmentation ratio, randomly divide the first dataset into two subsets, denoted as the first training set and the first evaluation set; and count the total number of records in the first training set to obtain the corresponding total number N. TR The total number N is obtained by statistically analyzing the total number of records in the first training set. EV ;
[0031] Both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first partitioning ratio, N TR :N EV ≈First division ratio;
[0032] Step 53: Input the first training structure and the first training vector of each first data record in the first training set as the current three-dimensional molecular structure M and the prior characteristic vector X into the molecular property prediction model for prediction, and use the property prediction vector Y obtained in this prediction as the corresponding first prediction vector; and form a corresponding first prediction-label pair by each first prediction vector and the first label vector of the corresponding first data record.
[0033] Wherein, each of the first prediction vectors is denoted as the corresponding first prediction vector Y. g Each of the first label vectors is denoted as the corresponding first label vector. Each of the first prediction vectors Y g The various predicted data y in k Let y be the corresponding predicted data. g,k Each of the first label vectors The various tag data in Record as the corresponding tag data
[0034] Step 54, obtain N TR Each of the first prediction-label pairs is input into a preset first model loss function L. M The calculation yields N classification loss values and one overall loss value.
[0035] Wherein, the first model loss function L M Specifically:
[0036]
[0037] L k () represents the loss function corresponding to the k-th type of prediction property, and each of the aforementioned loss functions L k Implemented based on L1 loss function, Smooth L1 loss function, or L2 loss function; all of the aforementioned loss functions L k The choice of loss function is not mandatory;
[0038] The classification loss value and the loss function L k One-to-one correspondence, which is the corresponding loss function L. k The output loss value; the overall loss value is the loss function L of the first model. M The output loss value; the classification loss value corresponds one-to-one with the classification loss value range; the overall loss value corresponds to the overall loss value range;
[0039] Step 55: Identify whether the overall loss value meets the range of the overall loss value; if it does, proceed to step 56; if not, based on the first model optimizer, move towards making the first model loss function L... M The direction that reaches the minimum value is used to perform one round of parameter modulation on the current parameter set, and the process returns to step 53 when the current round of parameter modulation ends.
[0040] The first model optimizer includes the Adam optimizer and the SGD optimizer.
[0041] Step 56: Identify whether all the classification loss values satisfy their respective classification loss value ranges; if yes, proceed to step 57; if no, take each classification loss value that does not satisfy its corresponding classification loss value range as a corresponding first loss value, and take the model parameters of the k-th predictor corresponding to each first loss value as the corresponding first predictor parameters, and take the loss function L corresponding to each first loss value as the first predictor parameters. k As the corresponding first loss function, a corresponding second model optimizer is assigned to each first loss value, and the corresponding first prediction head parameters are modulated in one round based on each second model optimizer in the direction of minimizing the corresponding first loss function. After the parameter modulation of all the first prediction head parameters in this time is completed, the process returns to step 53.
[0042] The second model optimizer includes the Adam optimizer and the SGD optimizer;
[0043] Step 57: Perform a traversal of all the first data records in the first evaluation set; during this traversal, take the currently traversed first data record as the corresponding current evaluation record; and input the first training structure and first training vector of the current evaluation record as the current three-dimensional molecular structure M and the prior property vector X into the molecular property prediction model for prediction, and take the property prediction vector Y obtained in this prediction as the corresponding second prediction vector; and form a corresponding second prediction-label pair by the second prediction vector and the first label vector of the current evaluation record; and at the end of this traversal, take the obtained N EV The second prediction-label is fed into a preset first model evaluation function to obtain the corresponding first evaluation value;
[0044] The first model evaluation function is implemented based on the RMSE function;
[0045] Step 58: Identify whether the first evaluation value meets the evaluation value range; if not, return to step 52 to continue training; if it does, confirm that the model training has ended.
[0046] Preferably, the step of preparing the corresponding three-dimensional molecular structure M and the prior property vector X by preparing model input data based on the first molecule information and the prior property set specifically includes:
[0047] The first molecular information is identified; if the first molecular information is a one-dimensional SMILES sequence or a two-dimensional molecular topology, then the corresponding three-dimensional molecular conformation is constructed based on the current first molecular information using the cheminformatics tool to obtain the corresponding current molecular conformation; if the first molecular information is a three-dimensional molecular conformation, then the current first molecular information is used as the corresponding current molecular conformation.
[0048] Based on the cheminformatics tools, the total number of aromatic rings, the total number of non-aromatic rings, and the proportion of heteroatoms in the current molecular conformation are identified to obtain the corresponding total number of second aromatic rings, the total number of second non-aromatic rings, and the proportion of second heteroatoms; and based on the quantum chemical calculation tools, the orbital overlap integral and the molecular dipole moment of the current molecular conformation are calculated to obtain the corresponding second orbital overlap integral and the second molecular dipole moment.
[0049] The prior characteristic vector X is composed of the total number of the second aromatic rings, the total number of the second non-aromatic rings, the proportion of the second heteroatoms, the second orbital overlap integral, and the second molecular dipole moment obtained in this study.
[0050] The total number of atoms in the current molecular conformation is identified, and the identification result is the corresponding total number of atoms N. A The element type and three-dimensional coordinates of each atom in the current molecular conformation are identified, and the identification results are used as the corresponding element type and three-dimensional coordinates of the atom. A corresponding atomic feature m is then formed by the element type and three-dimensional coordinates of each atom. i ; and consisting of all the atomic features m corresponding to the current molecular conformation. i To form a corresponding three-dimensional molecular structure M;
[0051] The obtained three-dimensional molecular structure M and the prior characteristic vector X are then used as the processing results of the model input data preparation.
[0052] A second aspect of the present invention provides an apparatus for implementing the molecular property prediction method that integrates three-dimensional structure and prior features as described in the first aspect above. The apparatus includes: a model building module, a model training module, and a model prediction module.
[0053] The model building module is used to identify the set of key quantum chemical properties of photoelectric molecules to obtain a corresponding set of predicted properties; and to select multiple other properties that can affect photoelectric conversion efficiency to form a corresponding set of prior properties; and to construct a corresponding molecular property prediction model based on the set of predicted properties and the set of prior properties; the predicted properties of the set of predicted properties include transition energy SOS1, transition energy SOT1, electron recombination energy, vertical ionization energy, ground potential energy surface recombination energy, excited potential energy surface recombination energy, adiabatic electron affinity, hole recombination energy, and the highest occupied orbital-lowest empty orbital gap; the prior properties of the set of prior properties include the total number of aromatic rings, the total number of non-aromatic rings, the proportion of heteroatoms, orbital overlap integral, and molecular dipole moment; the molecular property prediction model is used to predict the key quantum chemical properties of photoelectric molecules based on the three-dimensional molecular structure M and the prior characteristic vector X input to the model and output the corresponding property prediction vector Y; the prior characteristic vector X corresponds to the set of prior properties, and the property prediction vector Y corresponds to the set of predicted properties.
[0054] The model training module is used to construct a model training dataset through big data collection, denoted as the corresponding first dataset; and to train the molecular property prediction model based on the first dataset;
[0055] The model prediction module is used to receive the first molecular information input by the user after the model training is completed; and to prepare the model input data according to the first molecular information and the prior property set to obtain the corresponding three-dimensional molecular structure M and the prior property vector X; and to input the current three-dimensional molecular structure M and the prior property vector X into the molecular property prediction model to predict the corresponding property prediction vector Y and feed it back to the current user; the first molecular information is a one-dimensional SMILES sequence, two-dimensional molecular topology or three-dimensional molecular conformation of a photoelectrochemical molecule.
[0056] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0057] The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above;
[0058] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0059] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.
[0060] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for predicting molecular properties by integrating three-dimensional structure and prior features. As described above, after confirming the set of predicted properties and the set of prior properties, this invention configures a corresponding molecular property prediction model for these two sets. This molecular property prediction model is used to predict multiple quantum chemical properties specified in the set of predicted properties based on the three-dimensional molecular structure as input to the model and a prior characteristic vector corresponding to the set of prior properties, and outputs the corresponding property prediction vector. A model training dataset is constructed through big data collection for model training. After model training, the molecular property prediction model is used to predict multiple quantum chemical properties of any photoelectric molecule. This invention not only reduces prediction complexity but also shortens prediction time and improves prediction efficiency. Attached Figure Description
[0061] Figure 1 This is a schematic diagram of a molecular property prediction method that integrates three-dimensional structure and prior features, provided in Embodiment 1 of the present invention.
[0062] Figure 2 This is a block diagram of the molecular property prediction model provided in Embodiment 1 of the present invention;
[0063] Figure 3 This is a module structure diagram of a molecular property prediction device that integrates three-dimensional structure and prior features, provided in Embodiment 2 of the present invention.
[0064] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0066] Embodiment 1 of this invention provides a method for predicting molecular properties by integrating three-dimensional structure and prior features; such as Figure 1 The schematic diagram of a molecular property prediction method integrating three-dimensional structure and prior features provided in Embodiment 1 of the present invention is shown. The method mainly includes the following steps:
[0067] Step 1: Identify the set of key quantum chemical properties of photoelectric molecules to obtain the corresponding set of predicted properties; select multiple other properties that can affect photoelectric conversion efficiency to form the corresponding set of prior properties; and construct a corresponding molecular property prediction model based on the set of predicted properties and the set of prior properties.
[0068] Here, the predicted properties of the predicted property set in this embodiment of the invention include transition energy SOS1, transition energy SOT1, electron recombination energy, vertical ionization energy, ground potential energy surface recombination energy, excited potential energy surface recombination energy, adiabatic electron affinity energy, hole recombination energy, and the highest occupied orbital-lowest unoccupied orbital gap. It should be noted that the predicted properties of this predicted property set can also be added or removed based on application requirements.
[0069] The prior properties of the prior property set in this embodiment of the invention include the total number of aromatic rings, the total number of non-aromatic rings, the heteroatom ratio, the orbital overlap integral, and the molecular dipole moment; wherein:
[0070] 1) The total number of aromatic rings refers to the total number of ring structures in a molecule that conform to Hückel's rule; the presence of aromatic rings affects the degree of conjugation and stability of the molecule, and thus affects the highest occupied orbital-lowest empty orbital energy gap, excited state properties, and molecular conductivity and other property parameters; the conjugation effect of aromatic rings helps to improve the photoelectric conversion efficiency of the molecule;
[0071] 2) The total number of non-aromatic rings refers to the total number of ring structures in a molecule that do not conform to Hückel's rule; non-aromatic rings may affect the local electronic structure and reactivity of a molecule;
[0072] 3) The heteroatom ratio refers to the ratio of the number of heteroatoms (such as nitrogen, oxygen, sulfur, etc.) in a molecule to the total number of atoms. Changes in the heteroatom ratio can affect the polarity, ionization energy and electron affinity of a molecule, and changes in these properties can further affect the photoelectric conversion efficiency of the molecule.
[0073] 4) The orbital overlap integral is an integral in quantum chemistry that reflects the degree of overlap between two orbital wave functions. It describes the overlap between atomic or molecular orbitals. The magnitude of the orbital overlap integral directly affects charge transfer and exciton behavior within molecules. A larger orbital overlap integral facilitates electron transport within molecules, thereby improving the photoelectric conversion efficiency of molecules.
[0074] 5) Molecular dipole moment is a physical quantity that describes the non-uniformity of charge distribution in a molecule. It is equal to the product of the distance between the centroids of positive and negative charges in the molecule and the amount of charge. The magnitude and direction of the molecular dipole moment affect the intermolecular interactions, solubility, reaction rate and stereoselectivity. Changes in these properties will further affect the photoelectric conversion efficiency of the molecule.
[0075] It should be noted that the prior properties in the prior property set of the embodiments of the present invention can be added or removed based on application requirements.
[0076] The molecular property prediction model of this invention is used to predict the key quantum chemical properties of photoelectric molecules based on the three-dimensional molecular structure M and the prior property vector X input to the model, and output the corresponding property prediction vector Y.
[0077] Here, the three-dimensional molecular structure M consists of multiple atomic features m i Composition, 1≤index i≤N A N A The total number of atoms in the current molecular structure; atomic characteristic m i This includes the type of atomic element and the three-dimensional coordinates of the atom.
[0078] The prior feature vector X consists of multiple feature data x j Composition, 1≤indexj≤P, where P is the total number of prior properties in the prior property set; the prior characteristic vector X corresponds to the prior property set, and the characteristic data x j Each property corresponds one-to-one with the a priori properties of the a priori property set.
[0079] The property prediction vector Y consists of multiple prediction data y k Composition, 1 ≤ index k ≤ N, where N is the total number of predicted properties in the predicted property set; the property prediction vector Y corresponds to the predicted property set, and the predicted data y k Each property corresponds one-to-one with the predicted property set.
[0080] like Figure 2 As shown in the module structure diagram of the molecular property prediction model provided in Embodiment 1 of the present invention, the first model input terminal of the molecular property prediction model is used to receive the three-dimensional molecular structure M, the second model input terminal is used to receive the prior characteristic vector X, and the model output terminal is used to output the corresponding property prediction vector Y.
[0081] like Figure 2 As shown, the model components of the molecular property prediction model include: Uni-Mol model, molecular feature mapping network, prior feature mapping network, feature fusion layer, fusion feature encoder and multi-task prediction network; the multi-task prediction network consists of N parallel k-th prediction heads and a prediction output layer.
[0082] like Figure 2As shown, the connection relationships of the model components in the molecular property prediction model are as follows: the input of the Uni-Mol model is connected to the input of the first model, and its output is connected to the input of the molecular feature mapping network; the output of the molecular feature mapping network is connected to the first input of the feature fusion layer; the input of the prior feature mapping network is connected to the input of the second model, and its output is connected to the second input of the feature fusion layer; the output of the feature fusion layer is connected to the input of the fusion feature encoder; the output of the fusion feature encoder is connected to the input of each k-th prediction head of the multi-task prediction network; the output of each k-th prediction head of the multi-task prediction network is connected to one input of the prediction output layer; and the output of the prediction output layer of the multi-task prediction network is connected to the model output.
[0083] The functional descriptions of the model components of the molecular property prediction model are shown below.
[0084] 1) Uni-Mol model:
[0085] The Uni-Mol model of this invention is used to obtain the corresponding atomic feature tensor H by atomic-level feature encoding based on the three-dimensional molecular structure M. M Send to the molecular feature mapping network. Among them, the atomic feature tensor H... M The shape is N A ×C A C A The atomic feature dimension is a preset dimension, given by the Uni-Mol model and generally not modified; the atomic feature tensor H M By N A A length of C A atomic eigenvectors composition.
[0086] It should be noted that the Uni-Mol model in this embodiment of the invention is a pre-trained model capable of atomic-level high-dimensional feature encoding of three-dimensional molecular structures, and its pre-training has been completed. Therefore, parameter modulation is unnecessary during the subsequent training of the prediction model. Detailed model structure and pre-training methods of the Uni-Mol model can be found in the publicly available technical document "Uni-Mol: A Universal 3D Dolecular Representation Learning Framework," and will not be elaborated further here. The reason this embodiment uses the Uni-Mol model as the feature encoder for the three-dimensional molecular structure M is that this model can refine the precision of the encoded features to the atomic level. This provides richer original feature information for the molecular-level features generated by the subsequent molecular feature mapping network, thereby helping to further improve the prediction accuracy of the final predicted properties.
[0087] 2) Molecular feature mapping network:
[0088] The molecular feature mapping network of this invention consists of a feature mapping layer and a feature pooling layer; wherein, the feature mapping layer is formed by sequentially connecting two or more first linear activation layers, and each first linear activation layer is formed by connecting a corresponding first linear layer and a first nonlinear activation function, wherein the first nonlinear activation function is a ReLU activation function or a Sigmoid activation function.
[0089] Molecular feature mapping networks are used to map atomic feature tensors H using feature mapping layers. M each atom eigenvector Do a job from C A The feature mapping from a 3D feature space to a C1D feature space yields a mapped feature vector of length C1. And from the obtained N A 1 mapping feature vector Form a shape of N A ×C1 mapping feature tensor H proj The feature pooling layer then maps the feature tensor H using attention pooling. proj N for each feature dimension A Each feature data point is pooled to obtain corresponding feature pooled data; and the resulting C1 feature pooled data points form a corresponding mapped feature vector H1, which is then sent to the feature fusion layer. Here, the mapped feature vector H1 is a feature vector of length C1, where C1 is a preset first feature dimension, which is a pre-defined positive integer, for example, C1 = 1024.
[0090] It should be noted that the attention pooling method first processes the mapping feature tensor H. proj N in each feature channel A The attention weights of each feature data are calculated, and then based on the obtained N... A Each attention weight is paired with N A The feature data is weighted and summed, and the result is used as the feature pooling data for the current feature channel. It should also be noted that the mapped feature vector H1 output by the molecular feature mapping network is actually the molecular-level feature of the three-dimensional molecular structure M.
[0091] 3) Prior Feature Mapping Network:
[0092] The prior feature mapping network of this invention consists of one or two second linear activation layers connected sequentially. Each second linear activation layer is composed of a corresponding second linear layer and a second nonlinear activation function, where the second nonlinear activation function is a ReLU activation function or a Sigmoid activation function.
[0093] The prior feature mapping network is used to perform feature mapping on the prior feature vector X to obtain the corresponding mapped feature vector H2, which is then sent to the feature fusion layer. The mapped feature vector H2 is a feature vector of length C1.
[0094] It should be noted that the mapped feature vector H2 output by the prior feature mapping network is actually the feature encoding vector of the prior feature vector X; the function of the prior feature mapping network is to map the encoded features, i.e. the mapped feature vector H2, to the same feature space as the mapped feature vector H1 while encoding the prior feature vector X.
[0095] 4) Feature fusion layer:
[0096] The feature fusion layer in this embodiment of the invention is used to form a corresponding fused feature tensor H from the mapped feature vectors H1 and H2. R Send to the fusion feature encoder. Wherein, the fusion feature tensor H... R The shape is 2×C1.
[0097] 5) Fusion Feature Encoder:
[0098] The fusion feature encoder of this invention is implemented based on the encoder model of the Transformer architecture.
[0099] The fusion feature encoder is used to process the fusion feature tensor H R The corresponding encoded feature tensor E is obtained by performing feature encoding processing. R ; and encode the feature tensor E R The encoded feature vector E1 corresponding to the mapped feature vector H1 is sent to each of the k-th prediction heads of the multi-task prediction network. Here, the encoded feature tensor E... R The shape is 2×C1, consisting of encoding feature vector E1 and encoding feature vector E2; each of the encoding feature vectors E1 and E2 is a feature vector of length C1, and the encoding feature vectors E1 and E2 correspond one-to-one with the mapping feature vectors H1 and H2.
[0100] It should be noted that, based on publicly available information about the Transformer architecture, the encoder model within this architecture actually uses a series of multi-head attention operations to correlate each token feature vector in the input features with all other token feature vectors. In this embodiment of the invention, the fused feature tensor H... RThe mapped feature vectors H1 and H2 can be considered as two token feature vectors. After fusing a series of multi-head attention operations of the feature encoder, the mapped feature vector H2 can be associated with the mapped feature vector H1 to obtain a corresponding encoded feature vector E1, and the mapped feature vector H1 can be associated with the mapped feature vector H2 to obtain a corresponding encoded feature vector E2. Because the mapped feature vector H2 carries prior feature information, the encoded feature vector E1 is actually a molecular feature information that incorporates prior feature information. Based on this molecular feature with prior information, property prediction helps to improve prediction accuracy. Therefore, in this embodiment of the invention, the encoded feature vector E1 is passed to each k-th prediction head of the multi-task prediction network.
[0101] 6) Multi-task prediction network:
[0102] In this embodiment of the invention, each k-th prediction head of the multi-task prediction network is implemented based on a fully connected network. Each k-th prediction head is used to perform corresponding property prediction processing based on the encoded feature vector E1 to obtain the corresponding prediction data y. k Send to the prediction output layer. The prediction output layer of the multi-task prediction network will receive N prediction data y. k The corresponding property prediction vector Y is constructed and output.
[0103] Here, each k-th prediction head is a regression computation task head implemented based on a fully connected network. The fully connected network of each prediction head is an internal network of the current prediction head and is not shared with each other.
[0104] Step 2: Collect big data to build a model training dataset, denoted as the first dataset; and train the molecular property prediction model based on the first dataset.
[0105] Specifically, this includes: Step 21, constructing a model training dataset through big data collection, denoted as the corresponding first dataset;
[0106] The first dataset includes multiple first data records; each first data record includes a first training structure, a first training vector, and a first label vector; the data structures of the first training structure, the first training vector, and the first label vector are consistent with the corresponding three-dimensional molecular structure M, prior property vector X, and property prediction vector Y; the first label vector consists of N label data. composition;
[0107] Specifically, this includes: Step 211, collecting big data on the one-dimensional molecular sequences of photoelectric molecules through multiple preset data channels to obtain the corresponding molecular sequence library;
[0108] The data sources include publicly available photoelectrochemical molecular information databases and publicly available technical documents; the molecular sequence library includes multiple first molecular sequences; each first molecular sequence is a SMILES sequence of a photoelectrochemical molecule;
[0109] Step 212: Based on the preset cheminformatics tools, construct the corresponding three-dimensional molecular conformation according to each first molecular sequence to obtain the corresponding first molecular conformation;
[0110] Among them, cheminformatics tools include RDKit software;
[0111] Step 213 involves identifying the total number of aromatic rings, the total number of non-aromatic rings, and the proportion of heteroatoms for each first molecule conformation using cheminformatics tools to obtain the corresponding total number of first aromatic rings, the total number of first non-aromatic rings, and the proportion of first heteroatoms; and calculating the orbital overlap integral, molecular dipole moment, transition energy SOS1, transition energy SOT1, electron recombination energy, vertical ionization energy, ground potential energy surface recombination energy, excited potential energy surface recombination energy, adiabatic electron affinity energy, hole recombination energy, and highest occupied orbital-lowest empty orbital gap for the first molecule conformation using pre-set quantum chemical calculation tools to obtain the corresponding first orbital overlap integral, first molecule dipole moment, first transition energy SOS1, second transition energy SOT1, first electron recombination energy, and first vertical ionization energy. The first ground state energy surface recombination energy, the first excited state energy surface recombination energy, the first adiabatic electron affinity energy, the first hole recombination energy, and the first highest occupied orbital-lowest empty orbital energy gap are all considered. A corresponding a priori characteristic vector X is formed by the total number of first aromatic rings, the total number of first non-aromatic rings, the proportion of first heteroatoms, the first orbital overlap integral, and the first molecule dipole moment corresponding to each first molecule conformation. A corresponding property prediction vector Y is formed by the first transition energy S0S1, the second transition energy S0T1, the first electron recombination energy, the first vertical ionization energy, the first ground state energy surface recombination energy, the first excited state energy surface recombination energy, the first adiabatic electron affinity energy, the first hole recombination energy, and the first highest occupied orbital-lowest empty orbital energy gap corresponding to each first molecule conformation.
[0112] Among them, quantum chemical calculation tools include Multiwfn software, Gaussian software, ORCA software, GAMESS software, and NWChem software;
[0113] Step 214, and take each first molecular conformation as the corresponding current molecular conformation; and identify the total number of atoms in the current molecular conformation and take the identification result as the corresponding total number of atoms N. A It identifies the element type and three-dimensional coordinates of each atom in the current molecular conformation and uses the identification results as the corresponding atomic element type and three-dimensional coordinates; and combines the atomic element type and three-dimensional coordinates of each atom to form a corresponding atomic feature m. i; and composed of all atomic features m corresponding to the current molecular conformation. i To form a corresponding three-dimensional molecular structure M;
[0114] Step 215: Take the three-dimensional molecular structure M, prior characteristic vector X, and property prediction vector Y corresponding to each first molecular conformation as the corresponding first training structure, first training vector, and first label vector to form a corresponding first data record; and perform record deduplication on all the obtained first data records; and form the corresponding first dataset from all the remaining first data records after deduplication.
[0115] Step 22, and train the molecular property prediction model based on the first dataset;
[0116] Specifically, this includes: Step 221, setting an overall loss value range, N classification loss value ranges, and an evaluation value range; and the current parameter set is composed of the model parameters of the molecular feature mapping network, prior feature mapping network, fusion feature encoder, and multi-task prediction network of the molecular property prediction model;
[0117] Step 222: Based on a preset first segmentation ratio, the first dataset is randomly divided into two subsets, denoted as the first training set and the first evaluation set; and the total number of records in the first training set is counted to obtain the corresponding total number N. TR The total number N is obtained by statistically analyzing the total number of records in the first training set. EV ;
[0118] Wherein, the first segmentation ratio is a pre-set ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio, N TR :N EV ≈First division ratio;
[0119] Step 223: The first training structure and first training vector of each first data record in the first training set are used as the current three-dimensional molecular structure M and prior property vector X, and input into the molecular property prediction model for prediction. The property prediction vector Y obtained in this prediction is used as the corresponding first prediction vector. Each first prediction vector and the first label vector of the corresponding first data record form a corresponding first prediction-label pair.
[0120] Here, each of the first prediction vectors is denoted as the corresponding first prediction vector Y. g Each first label vector is denoted as the corresponding first label vector. Each first prediction vector Y g The various prediction data y in k Let y be the corresponding predicted data.g,k Each first label vector The various tag data in Record as the corresponding tag data
[0121] Step 224, obtain N TR Each first prediction-label pair is input into the preset first model loss function L. M The calculation yields N classification loss values and one overall loss value.
[0122] Here, the first model loss function L in this embodiment of the invention M Specifically:
[0123]
[0124] Among them, L k () represents the loss function corresponding to the k-th prediction property. Each loss function L... k Implemented based on L1 loss function, Smooth L1 loss function, or L2 loss function; all loss functions L k The choice of loss function is not mandatory;
[0125] Classification loss value and loss function L k One-to-one correspondence, for the corresponding loss function L k The output loss value; the overall loss value is the first model loss function L. M The output loss value; the classification loss value corresponds one-to-one with the classification loss value range; the overall loss value corresponds to the overall loss value range;
[0126] Step 225: Identify whether the overall loss value meets the overall loss value range; if it does, proceed to step 226; if not, based on the first model optimizer, move towards making the first model loss function L... M The direction that reaches the minimum value performs one round of parameter modulation on the current parameter set, and returns to step 223 when the current round of parameter modulation ends;
[0127] The first model optimizer includes the Adam optimizer and the SGD optimizer;
[0128] Step 226: Identify whether all classification loss values satisfy their respective classification loss value ranges; if yes, proceed to step 227; if not, take each classification loss value that does not satisfy its corresponding classification loss value range as a corresponding first loss value, and take the model parameters of the k-th predictor corresponding to each first loss value as the corresponding first predictor parameters, and take the loss function L corresponding to each first loss value as the first predictor parameters. kAs the corresponding first loss function, a corresponding second model optimizer is assigned to each first loss value, and the corresponding first prediction head parameters are modulated in one round based on each second model optimizer in the direction of minimizing the corresponding first loss function. After the parameter modulation of all first prediction head parameters is completed, the process returns to step 223.
[0129] The second model optimizer includes the Adam optimizer and the SGD optimizer.
[0130] Step 227: Perform a round of traversal on all first data records in the first evaluation set; during this round of traversal, take the currently traversed first data record as the corresponding current evaluation record; and use the first training structure and first training vector of the current evaluation record as the current three-dimensional molecular structure M and prior property vector X as inputs into the molecular property prediction model for prediction, and use the property prediction vector Y obtained in this prediction as the corresponding second prediction vector; and form a corresponding second prediction-label pair with the second prediction vector and the first label vector of the current evaluation record; and at the end of this round of traversal, obtain N EV Each second prediction-label is input into a preset first model evaluation function to obtain the corresponding first evaluation value;
[0131] The first model evaluation function is implemented based on the RMSE function;
[0132] Step 228: Identify whether the first evaluation value meets the evaluation value range; if not, return to step 222 to continue training; if it does, confirm that the model training is complete.
[0133] Step 3: After the model training is completed, the first molecular information input by the user is received; and the model input data is prepared according to the first molecular information and the prior property set to obtain the corresponding three-dimensional molecular structure M and prior property vector X; and the current three-dimensional molecular structure M and prior property vector X are input into the molecular property prediction model to make predictions and obtain the corresponding property prediction vector Y, which is then fed back to the current user.
[0134] Specifically, this includes: Step 31, after the model training is completed, receiving the first molecule information input by the user;
[0135] Among them, the first molecule information is a one-dimensional SMILES sequence, two-dimensional molecular topology, or three-dimensional molecular conformation of a photoelectric molecule;
[0136] Step 32, and prepare the model input data based on the first molecule information and the prior property set to obtain the corresponding three-dimensional molecular structure M and prior property vector X;
[0137] Specifically, it includes: step 321, identifying the first molecule information; if the first molecule information is a one-dimensional SMILES sequence or a two-dimensional molecular topology, then constructing the corresponding three-dimensional molecular conformation based on the current first molecule information using cheminformatics tools to obtain the corresponding current molecular conformation; if the first molecule information is a three-dimensional molecular conformation, then using the current first molecule information as the corresponding current molecular conformation.
[0138] Step 322: Based on cheminformatics tools, the total number of aromatic rings, the total number of non-aromatic rings, and the proportion of heteroatoms in the current molecular conformation are identified to obtain the corresponding total number of second aromatic rings, the total number of second non-aromatic rings, and the proportion of second heteroatoms; and based on quantum chemical calculation tools, the orbital overlap integral and the molecular dipole moment of the current molecular conformation are calculated to obtain the corresponding second orbital overlap integral and the second molecular dipole moment.
[0139] Step 323, and a corresponding prior characteristic vector X is formed by the total number of second aromatic rings, the total number of second non-aromatic rings, the proportion of second heteroatoms, the second orbital overlap integral, and the second molecular dipole moment obtained in this step;
[0140] Step 324 involves identifying the total number of atoms in the current molecular conformation and setting the identification result as the corresponding total number of atoms N. A It identifies the element type and three-dimensional coordinates of each atom in the current molecular conformation and uses the identification results as the corresponding atomic element type and three-dimensional coordinates; and combines the atomic element type and three-dimensional coordinates of each atom to form a corresponding atomic feature m. i ; and composed of all atomic features m corresponding to the current molecular conformation. i To form a corresponding three-dimensional molecular structure M;
[0141] Step 325, and output the three-dimensional molecular structure M and prior characteristic vector X obtained in this step as the processing result of preparing the input data for this model;
[0142] Step 33: Input the current three-dimensional molecular structure M and the prior property vector X into the molecular property prediction model to make predictions and obtain the corresponding property prediction vector Y, which is then fed back to the current user.
[0143] Figure 3 This is a module structure diagram of a molecular property prediction device integrating three-dimensional structure and prior features provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 3 As shown, the device includes: a model building module 201, a model training module 202, and a model prediction module 203.
[0144] The model building module 201 is used to identify the set of key quantum chemical properties of photoelectric molecules to obtain the corresponding set of predicted properties; and to select multiple other properties that can affect photoelectric conversion efficiency to form the corresponding set of prior properties; and to construct a corresponding molecular property prediction model based on the set of predicted properties and the set of prior properties; the predicted properties in the set of predicted properties include the transition energy S. 0- S1, transition energy S0T1, electron recombination energy, vertical ionization energy, ground potential energy surface recombination energy, excited potential energy surface recombination energy, adiabatic electron affinity energy, hole recombination energy, highest occupied orbital-lowest empty orbital energy gap; the a priori properties of the a priori property set include the total number of aromatic rings, the total number of non-aromatic rings, the proportion of heteroatoms, orbital overlap integral, and molecular dipole moment; the molecular property prediction model is used to predict the key quantum chemical properties of photoelectric molecules based on the three-dimensional molecular structure M and the a priori property vector X input to the model and output the corresponding property prediction vector Y; the a priori property vector X corresponds to the a priori property set, and the property prediction vector Y corresponds to the predicted property set.
[0145] The model training module 202 is used to construct a model training dataset through big data collection, denoted as the corresponding first dataset; and to train the molecular property prediction model based on the first dataset.
[0146] The model prediction module 203 is used to receive the first molecular information input by the user after the model training is completed; and to prepare the model input data according to the first molecular information and the prior property set to obtain the corresponding three-dimensional molecular structure M and prior property vector X; and to input the current three-dimensional molecular structure M and prior property vector X into the molecular property prediction model to predict the corresponding property prediction vector Y and feed it back to the current user; the first molecular information is a one-dimensional SMILES sequence, two-dimensional molecular topology or three-dimensional molecular conformation of a photoelectric molecule.
[0147] The molecular property prediction device that integrates three-dimensional structure and prior features provided in this embodiment of the invention can execute the method steps in the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.
[0148] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing elements; they can be fully implemented in hardware; or some modules can be implemented by processing elements calling software, while others are implemented in hardware. For example, the model building module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0149] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-on-a-Chip (SOC).
[0150] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0151] Figure 4 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 4 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.
[0152] exist Figure 4The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0153] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0154] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.
[0155] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for predicting molecular properties by integrating three-dimensional structure and prior features. As described above, after confirming the set of predicted properties and the set of prior properties, this invention configures a corresponding molecular property prediction model for these two sets. This molecular property prediction model is used to predict multiple quantum chemical properties specified in the set of predicted properties based on the three-dimensional molecular structure as input to the model and a prior characteristic vector corresponding to the set of prior properties, and outputs the corresponding property prediction vector. A model training dataset is constructed through big data collection for model training. After model training, the molecular property prediction model is used to predict multiple quantum chemical properties of any photoelectric molecule. This invention not only reduces prediction complexity but also shortens prediction time and improves prediction efficiency.
[0156] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0157] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method of predicting molecular properties by fusing three-dimensional structures and prior features, characterized by, The method comprises: Confirming a corresponding set of predicted properties from a set of key quantum chemical properties of a photoelectric molecule; and selecting a set of other properties capable of affecting photoelectric conversion efficiency to form a corresponding set of prior properties; and constructing a corresponding molecular property prediction model based on the set of predicted properties and the set of prior properties; the predicted properties of the set of predicted properties include transition energy S0S1, transition energy S0T1, electronic reorganization energy, vertical ionization energy, ground state potential energy surface reorganization energy, excited state potential energy surface reorganization energy, adiabatic electron affinity, hole reorganization energy, highest occupied orbital-lowest unoccupied orbital energy gap; the prior properties of the set of prior properties include total number of aromatic rings, total number of non-aromatic rings, heteroatom ratio, orbital overlap integral, molecular dipole moment; the molecular property prediction model is used to predict the key quantum chemical properties of the photoelectric molecule according to the three-dimensional molecular structure M and the prior characteristic vector X input by the model, and output a corresponding property prediction vector Y; the prior characteristic vector X corresponds to the set of prior properties, and the property prediction vector Y corresponds to the set of predicted properties; A model training data set is constructed by big data collection, denoted as a corresponding first data set; and the molecular property prediction model is trained based on the first data set; After the model training is completed, the first molecular information input by a user is received; and model input data preparation is performed according to the first molecular information and the set of prior properties to obtain the three-dimensional molecular structure M and the prior characteristic vector X; and the three-dimensional molecular structure M and the prior characteristic vector X are input into the molecular property prediction model for prediction to obtain the property prediction vector Y corresponding to the current user feedback; the first molecular information is a one-dimensional SMILES sequence, a two-dimensional molecular topology or a three-dimensional molecular conformation of a photoelectric molecule; Wherein, the three-dimensional molecular structure M is composed of a plurality of atomic characteristics m i , 1≤index i≤N A , N A is the total number of atoms of the current molecular structure; the atomic characteristics m i include atomic element type, atomic three-dimensional coordinates; The model input data preparation according to the first molecular information and the set of prior properties to obtain the three-dimensional molecular structure M and the prior characteristic vector X specifically comprises: The first molecular information is identified; if the first molecular information is a one-dimensional SMILES sequence or a two-dimensional molecular topology, a corresponding three-dimensional molecular conformation is constructed based on a chemical informatics tool according to the current first molecular information to obtain a corresponding current molecular conformation; if the first molecular information is a three-dimensional molecular conformation, the current first molecular information is taken as the corresponding current molecular conformation; And the total number of aromatic rings, the total number of non-aromatic rings and the heteroatom ratio of the current molecular conformation are identified based on the chemical informatics tool to obtain a corresponding second total number of aromatic rings, a second total number of non-aromatic rings and a second heteroatom ratio; and the orbital overlap integral and the molecular dipole moment of the current molecular conformation are calculated based on a quantum chemical calculation tool to obtain a second orbital overlap integral and a second molecular dipole moment; And the second total number of aromatic rings, the second total number of non-aromatic rings, the second heteroatom ratio, the second orbital overlap integral and the second molecular dipole moment obtained this time form a corresponding prior characteristic vector X. and the total number of atoms of the current molecular conformation is identified and the identification result is the corresponding total number of atoms N A ; and the element type and three-dimensional coordinates of each atom of the current molecular conformation are identified, and the identification result is the corresponding atomic element type and atomic three-dimensional coordinates; and a corresponding atomic characteristic m i is composed of the atomic element type and the atomic three-dimensional coordinates corresponding to each atom; and a corresponding three-dimensional molecular structure M i is composed of all the atomic characteristics m corresponding to the current molecular conformation. The three-dimensional molecular structure M and the prior characteristic vector X obtained this time are output as the processing results of the preparation of the model input data this time.
2. The molecular property prediction method integrating three-dimensional structure and prior features according to claim 1, characterized in that, The prior characteristic vector X is composed of a plurality of characteristic data x j , 1≤index j≤P, P is the total number of prior properties of the prior property set; the characteristic data x j corresponds to a prior property of the prior property set one by one. The property prediction vector Y is composed of a plurality of prediction data y k , 1≤index k≤N, N being the total number of prediction properties of the set of prediction properties; the prediction data y k correspond one-to-one to a prediction property of the set of prediction properties. The first data set comprises a plurality of first data records; the first data records comprise a first training structure, a first training vector, and a first label vector; the data structures of the first training structure, the first training vector, and the first label vector are consistent with the corresponding three-dimensional molecular structure M, the prior property vector X, and the property prediction vector Y; the first label vector is composed of N label data .
3. The molecular property prediction method integrating three-dimensional structure and prior features according to claim 2, characterized in that, The first model input end of the molecular property prediction model is used to receive the three-dimensional molecular structure M, the second model input end is used to receive the prior characteristic vector X, and the model output end is used to output the corresponding property prediction vector Y; The molecular property prediction model comprises a Uni-Mol model, a molecular feature mapping network, a prior feature mapping network, a feature fusion layer, a fusion feature encoder and a multi-task prediction network; the multi-task prediction network is composed of N parallel kth prediction heads and a prediction output layer; the Uni-Mol model has been pre-trained; The input end of the Uni-Mol model is connected with the first model input end, and the output end is connected with the input end of the molecular feature mapping network; the output end of the molecular feature mapping network is connected with the first input end of the feature fusion layer; the input end of the prior feature mapping network is connected with the second model input end, and the output end is connected with the second input end of the feature fusion layer; the output end of the feature fusion layer is connected with the input end of the fusion feature encoder; the output end of the fusion feature encoder is connected with the input end of each kth prediction head of the multi-task prediction network; the output end of each kth prediction head of the multi-task prediction network is connected with one input end of the prediction output layer; the output end of the prediction output layer of the multi-task prediction network is connected with the model output end; The Uni-Mol model is used for atomic-level feature coding according to the three-dimensional molecular structure M to obtain a corresponding atomic feature tensor H M sending to the molecular feature mapping network; the atomic feature tensor H M The shape of N A ×C A , C A is a preset atomic feature dimension; the atomic feature tensor H M consists of N A atomic feature vectors of length C A The molecular feature mapping network is composed of a feature mapping layer and a feature pooling layer, the feature mapping layer is sequentially connected by two or more first linear activation layers, each first linear activation layer is connected by a corresponding first linear layer and a first nonlinear activation function, the first nonlinear activation function is a ReLU activation function or a Sigmoid activation function; the molecular feature mapping network is used to perform feature mapping processing on the atomic feature tensor H M of each atomic feature vector from a C A dimensional feature space to a C1 dimensional feature space to obtain a mapping feature vector with a vector length of C1 ; and N A mapping feature vectors are obtained to form a mapping feature tensor H A with a shape of N proj ×C1; and the feature pooling layer is used to perform pooling processing on N proj feature data of each feature dimension of the mapping feature tensor H A in an attention pooling manner to obtain corresponding feature pooling data; and C1 feature pooling data are obtained to form a corresponding mapping feature vector H1 with a length of C1, and C1 is a preset first feature dimension; The prior feature mapping network is sequentially connected by one or two layers of second linear activation layers; each second linear activation layer is connected by a corresponding second linear layer and a second nonlinear activation function; the second nonlinear activation function is a ReLU activation function or a Sigmoid activation function; the prior feature mapping network is used to perform feature mapping processing on the prior characteristic vector X to obtain a corresponding mapping feature vector H2, which is sent to the feature fusion layer; the mapping feature vector H2 is a feature vector with a length of C1; The feature fusion layer is used to form a corresponding fusion feature tensor H by the mapping feature vectors H1, H2 R The fusion feature encoder is used to send the fusion feature tensor H R The shape is 2xC1; The fusion feature encoder is implemented based on an encoder model of a Transformer architecture; the fusion feature encoder is used to encode the fusion feature tensor H R to obtain a corresponding encoded feature tensor E R ; and the encoded feature vector E1 corresponding to the mapping feature vector H1 in the encoded feature tensor E R is sent to each of the kth prediction heads of the multi-task prediction network; the shape of the encoded feature tensor E R is 2×C1, which is composed of the encoded feature vector E1 and the encoded feature vector E2; the encoded feature vector E1 and the encoded feature vector E2 are each a feature vector with a length of C1, and the encoded feature vector E1 and the encoded feature vector E2 correspond one-to-one to the mapping feature vector H1 and the mapping feature vector H2. Each of the kth prediction heads of the multi-task prediction network is implemented based on a fully connected network; each of the kth prediction heads is configured to perform corresponding property prediction processing on the encoded feature vector E1 to obtain corresponding prediction data y k send to the prediction output layer; The prediction output layer of the multi-task prediction network outputs N prediction data y k composing the corresponding property prediction vector Y and outputting. 4.The method of claim 2, wherein, The model training data set constructed by big data collection is denoted as a corresponding first data set, which specifically comprises: A one-dimensional molecular sequence of an optoelectronic molecule is collected by a plurality of preset data channels to obtain a corresponding molecular sequence library; the plurality of data channels include a public optoelectronic molecule information library and public technical literature; the molecular sequence library comprises a plurality of first molecular sequences; each first molecular sequence is a SMILES sequence of an optoelectronic molecule; A corresponding first molecular conformation is constructed based on a preset chemical informatics tool according to each first molecular sequence; the chemical informatics tool includes RDKit software; and the proportion of heteroatoms of each of the first molecular conformations are identified based on the chemoinformatics tool to obtain corresponding first total number of aromatic rings, first total number of non-aromatic rings and first proportion of heteroatoms; and the orbital overlap integral, molecular dipole moment, transition energy S0S1, transition energy S0T1, electronic reorganization energy, vertical ionization energy, ground state potential energy surface reorganization energy, excited state potential energy surface reorganization energy, adiabatic electron affinity, hole reorganization energy, highest occupied molecular orbital-lowest unoccupied orbital energy gap of the first molecular conformations are calculated based on a preset quantum chemistry calculation tool to obtain corresponding first orbital overlap integral, first molecular dipole moment, first transition energy S0S1, second transition energy S0T1, first electronic reorganization energy, first vertical ionization energy, first ground state potential energy surface reorganization energy, first excited state potential energy surface reorganization energy, first adiabatic electron affinity, first hole reorganization energy, first highest occupied molecular orbital-lowest unoccupied orbital energy gap; and a corresponding prior characteristic vector X is composed of the first total number of aromatic rings, the first total number of non-aromatic rings, the first proportion of heteroatoms, the first orbital overlap integral and the first molecular dipole moment corresponding to each of the first molecular conformations; and a corresponding property prediction vector Y is composed of the first transition energy S0S1, the second transition energy S0T1, the first electronic reorganization energy, the first vertical ionization energy, the first ground state potential energy surface reorganization energy, the first excited state potential energy surface reorganization energy, the first adiabatic electron affinity, the first hole reorganization energy and the first highest occupied molecular orbital-lowest unoccupied orbital energy gap corresponding to each of the first molecular conformations; the quantum chemistry calculation tool includes Multiwfn software, Gaussian software, ORCA software, GAMESS software and NWChem software; and each of the first molecular conformations is taken as a corresponding current molecular conformation; and the total number of atoms of the current molecular conformation is identified and the identification result is taken as a corresponding total number of atoms N A ; and the element type and three-dimensional coordinates of each atom of the current molecular conformation are identified and the identification result is taken as a corresponding atomic element type and atomic three-dimensional coordinates; and a corresponding atomic feature m i is composed of the atomic element type and the atomic three-dimensional coordinates corresponding to each atom i ; and a corresponding three-dimensional molecular structure M is composed of all the atomic features m i corresponding to the current molecular conformation. the three-dimensional molecular structure M, the prior characteristic vector X and the property prediction vector Y corresponding to each of the first molecular conformations are taken as corresponding first training structure, first training vector and first label vector to form a corresponding first data record; and all the obtained first data records are subjected to record deduplication processing; and all the first data records remaining after deduplication form a corresponding first data set. 5.The method of claim 3, wherein, The training of the molecular property prediction model based on the first data set specifically includes: Step 51, setting an overall loss value range, N classification loss value ranges and an evaluation value range; and the model parameters of the molecular feature mapping network, the prior feature mapping network, the fusion feature encoder and the multi-task prediction network of the molecular property prediction model form a corresponding current parameter set; Step 52, based on the preset first split ratio, randomly split the first data set into two sub-data sets, denoted as the corresponding first training set and the first evaluation set; and count the total number of records of the first training set to obtain the corresponding total number N TR ; and count the total number of records of the first training set to obtain the corresponding total number N EV ; The first training set and the first evaluation set are both composed of a plurality of the first data records; a proportion of the total number of records of the first training set to the first evaluation set meets the first split proportion, N TR : N EV ≈ the first split proportion; Step 53, input the first training structure and the first training vector of each first data record of the first training set as the current three-dimensional molecular structure M and the prior characteristic vector X into the molecular property prediction model for prediction, and take the property prediction vector Y obtained by this prediction as a corresponding first prediction vector; and a corresponding first prediction-label pair is composed of each first prediction vector and the first label vector of the corresponding first data record; wherein each of the first prediction vectors is denoted as a corresponding first prediction vector Y g , each of the first label vectors is denoted as a corresponding first label vector ; each of the prediction data y g in each of the first prediction vectors Y k is denoted as a corresponding prediction data y g,k ; each of the label data in each of the first label vectors is denoted as a corresponding label data ; Step 54, the obtained N TR first prediction-label pairs are taken into a preset first model loss function L M corresponding N classification loss values and an overall loss value are calculated. The first model loss function L M Specifically: ; L k () is the loss function corresponding to the kth predicted property, and each of the loss functions L k is realized based on an L1 loss function, a SmoothL1 loss function or an L2 loss function; all the loss functions L k The selection of the loss function is not forced to be consistent. The classification loss value corresponds to the loss function L k One-to-one correspondence, corresponding to the loss function L k The output loss value; the overall loss value is the first model loss function L M The output loss value; the classification loss value corresponds to the classification loss value range; the overall loss value corresponds to the overall loss value range; Step 55, identify whether the overall loss value meets the overall loss value range; if yes, go to step 56; if not, based on the first model optimizer, adjust the current parameter set towards the direction of making the first model loss function L M reach the minimum value, and return to step 53 at the end of this round of parameter modulation. Wherein, the first model optimizer includes an Adam optimizer and an SGD optimizer; Step 56, identify whether all the classification loss values each satisfy the corresponding classification loss value range; if yes, go to step 57; if no, take each classification loss value that does not satisfy the corresponding classification loss value range as a corresponding first loss value, take the model parameters of the kth prediction head corresponding to each first loss value as a corresponding first prediction head parameter, and take the loss function L k as a corresponding first loss function, and assign a corresponding second model optimizer to each first loss value, and perform a round of parameter modulation on the corresponding first prediction head parameter based on each second model optimizer towards the direction of making the corresponding first loss function reach a minimum value, and return to step 53 after the parameter modulation of all first prediction head parameters this time is completed; Wherein, the second model optimizer includes an Adam optimizer and an SGD optimizer; Step 57, a round of traversal is performed on all the first data records of the first evaluation set; and during the round of traversal, the first data record currently being traversed is taken as a corresponding current evaluation record; and the first training structure and the first training vector of the current evaluation record are taken as the three-dimensional molecular structure M and the prior characteristic vector X to input the molecular property prediction model for prediction, and the property prediction vector Y obtained by the prediction is taken as a corresponding second prediction vector; and a corresponding second prediction-label pair is formed by the second prediction vector and the first label vector of the current evaluation record; and when the round of traversal ends, the N EV second prediction-label pairs obtained are taken into a preset first model evaluation function to obtain a corresponding first evaluation value; Wherein, the first model evaluation function is realized based on an RMSE function; Step 58, identify whether the first evaluation value meets the evaluation value range; if not, return to step 52 for continuous training; if yes, confirm that the model training is completed.
6. An apparatus for performing the method of predicting molecular properties from fused three-dimensional structures and prior features according to any one of claims 1-5, characterized in that, The device includes a model construction module, a model training module, and a model prediction module; The model construction module is used to confirm a set of predicted properties corresponding to a set of types of key quantum chemical properties of optoelectronic molecules, and select a set of prior properties corresponding to a set of other properties that can affect the photoelectric conversion efficiency; and construct a corresponding molecular property prediction model based on the set of predicted properties and the set of prior properties; the predicted properties of the set of predicted properties include transition energy S0S1, transition energy S0T1, electronic reorganization energy, vertical ionization energy, ground state potential energy surface reorganization energy, excited state potential energy surface reorganization energy, adiabatic electron affinity, hole reorganization energy, highest occupied orbital-lowest unoccupied orbital energy gap; the prior properties of the set of prior properties include total number of aromatic rings, total number of non-aromatic rings, proportion of heteroatoms, orbital overlap integral, molecular dipole moment; the molecular property prediction model is used to predict the key quantum chemical properties of optoelectronic molecules according to the three-dimensional molecular structure M and the prior characteristic vector X input by the model, and output a corresponding property prediction vector Y; the prior characteristic vector X corresponds to the set of prior properties, and the property prediction vector Y corresponds to the set of predicted properties; The model training module is used to construct a model training data set by big data acquisition, denoted as a corresponding first data set; and train the molecular property prediction model based on the first data set; The model prediction module is used to receive first molecular information input by a user after the model training is completed; and prepare model input data according to the first molecular information and the set of prior properties to obtain the three-dimensional molecular structure M and the prior characteristic vector X; and input the current three-dimensional molecular structure M and the prior characteristic vector X into the molecular property prediction model for prediction to obtain a corresponding property prediction vector Y, which is fed back to the current user; the first molecular information is a one-dimensional SMILES sequence, a two-dimensional molecular topology, or a three-dimensional molecular conformation of an optoelectronic molecule.
7. An electronic device, comprising: It includes: a memory, a processor, and a transceiver; The processor is coupled with the memory to read and execute instructions in the memory to implement the method of any one of claims 1-5. The transceiver is coupled with the processor to be controlled by the processor to perform message transceiving.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, when the computer instructions are executed by a computer, cause the computer to perform the method of any one of claims 1-5.
Citation Information
Patent Citations
Training method and device for dihedral angle energy prediction model
CN118538328A
Method and device for predicting optical energy characteristics of organic photovoltaic material molecules
CN120183563A
Cited By
A method, system, and equipment for predicting the performance and verifying the synthesis of organic cathode materials for lithium batteries.
CN122575593A