A processing method and apparatus for predicting the photoelectric properties of solvent molecules.
By constructing a multi-task prediction model and combining the structure and prior properties of solute and solvent, the problems of high complexity and low accuracy in the analysis of photoelectric properties of solvent molecules in existing technologies are solved, and efficient and accurate prediction of photoelectric properties of solvent molecules is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies have high computational complexity and long computation time when analyzing the photoelectric properties of solvent molecules. Furthermore, conventional models fail to effectively consider the influence of solute structure, solute properties, and solute-solvent interactions on solvent properties, resulting in insufficient prediction accuracy.
A multi-task prediction model is constructed, which combines the solute molecule structure, solvent molecule structure and five types of prior properties, and performs feature fusion through a deep learning model to predict the photoelectric properties of solvent molecules, including emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak and absorption half-width.
It reduces the complexity and time of analysis, improves the accuracy and efficiency of prediction, and enhances the flexibility and generalization of the model.
Smart Images

Figure CN120808953B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a processing method and apparatus for predicting the photoelectric properties of solvent molecules. Background Technology
[0002] Optoelectronic material solutions are complex systems composed of solvent and solute molecules, whose photoelectric properties are jointly determined by intermolecular interactions. Solute molecules are the direct executors of photoelectric conversion, while solvent molecules influence the photoelectric properties of solute molecules through dissolution, conformational modulation, and spectral modulation. Given a fixed solute molecule count, analyzing the photoelectric properties of different solvent molecules is essential for solvent optimization engineering in optoelectronic materials. Commonly used photoelectric properties of solvent molecules include: the peak point of the solvent molecule's emission spectrum (emission peak), the residence time of the solvent molecule in the excited state (molecular lifetime), the photoelectric conversion efficiency of the solvent molecule, the peak point of the solvent molecule's absorption spectrum (absorption peak), and the full width at half maximum (FWHM) of the solvent molecule's absorption spectrum.
[0003] Currently, most techniques for analyzing the aforementioned photoelectric properties of solvent molecules are based on quantum chemical calculations (such as density functional theory). The biggest problem with this conventional approach is its high computational complexity, long computation time, and low efficiency, making it difficult to meet the requirements of high-throughput computing. With the application of artificial intelligence models in the field of optoelectronic materials, constructing predictive models for target properties can reduce analytical complexity, shorten analysis time, and improve analytical efficiency. However, conventional models only use the solvent molecule's own structure as a reference for prediction, without considering the influence of solute structure, solute properties, and solute-solvent interaction characteristics on solvent properties. This means that the prediction accuracy of conventional models still needs further improvement. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a method, apparatus, electronic device, and computer-readable storage medium for predicting the photoelectric properties of solvent molecules. This invention uses five types of solute / solvent properties that can be quickly obtained through simple calculations as prior properties (total number of aromatic rings in the solute, interatomic interaction energy of the solute, solvent dipole moment, solvent polarity index, and solute-solvent orbital overlap integral calculated using the GFN2-xTB method). It uses five types of photoelectric properties of solvent molecules that require lengthy calculations (emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum) as prediction targets. A multi-task prediction model is constructed that can predict the five prediction targets based on the solute molecule structure, solvent molecule structure, and the five types of prior properties. Two optional structures are provided for the multi-task prediction model. A model training dataset is constructed to train the model. After training, the multi-task prediction model assists users in handling high-throughput solvent selection tasks. This invention not only reduces analytical complexity, shortens analysis time, and improves analytical efficiency, but also improves prediction accuracy and the flexibility and generalization of the prediction model.
[0005] To achieve the above objectives, a first aspect of the present invention provides a method for predicting the photoelectric properties of solvent molecules, the method comprising:
[0006] A deep learning model is constructed that fuses features of solute molecule structure, solvent molecule structure, and prior properties of the solute-solvent system, and predicts multiple photoelectric properties of solvent molecules based on the fused features. This model is denoted as the corresponding multi-task prediction model. The multi-task prediction model is used to predict multiple photoelectric properties of solvent molecules based on the solute molecule structure M input to the model. 1 Solvent molecular structure M 2 The prior property vector X is used to perform multi-task prediction of the photoelectric properties of solvent molecules and output the corresponding solvent property prediction vector Y; both the solute molecule structure and the solvent molecule structure are three-dimensional molecular structures; the prior properties of the solute-solvent system include the total number of aromatic rings of the solute molecule, the interatomic interaction energy of the solute molecule, the dipole moment of the solvent molecule, the polarity index of the solvent molecule, and the orbital overlap integral of the solute molecule and the solvent molecule; the various photoelectric properties include emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum; the solute molecule structure M... 1 Characterized by multiple atoms Composition, 1≤index i≤N 1 N 1 The total number of atoms in the solute molecule; the atomic characteristics Includes atomic element types and atomic three-dimensional coordinates; the solvent molecule structure M 2 Characterized by multiple atoms Composition, 1≤indexj≤N 2N 2 The total number of atoms in the solvent molecules; the atomic characteristics The a priori characteristic vector X includes the type of atomic element and the three-dimensional coordinates of the atom; the a priori characteristic vector X includes the total number of aromatic rings in the solute x1, the interatomic interaction energy of the solute x2, the solvent dipole moment x3, the solvent polarity index x4, and the solute-solvent orbital overlap integral x5; the solvent characteristic prediction vector Y includes the emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5.
[0007] The model training dataset is constructed and denoted as the corresponding first dataset; and the multi-task prediction model is trained based on the first dataset;
[0008] After model training, the multi-task prediction model is used to perform solvent optimization on the batch optimization task input by the user and the processing results are fed back to the current user. The batch optimization task includes a first solute molecule sequence, a first solvent molecule sequence set, and five types of index threshold ranges. The first solvent molecule sequence set includes multiple first solvent molecule sequences. Each first solvent molecule sequence and each first solvent molecule sequence is a SMILES sequence. The five types of index threshold ranges include emission peak threshold range, molecular lifetime threshold range, photoelectric conversion efficiency threshold range, absorption peak threshold range, and absorption half-width threshold range.
[0009] Preferably, the first dataset includes multiple first data records; each first data record includes a first solute molecule structure, a first solvent molecule structure, a first prior property vector, and a first solvent property label vector; the data structures of the first solute molecule structure, the first solvent molecule structure, the first prior property vector, and the first solvent property label vector are related to the corresponding solute molecule structure M. 1 The solvent molecule structure M 2 The prior characteristic vector X and the solvent characteristic prediction vector Y are consistent.
[0010] Preferably, the first model input of the multi-task prediction model is used to receive the solute molecular structure M. 1 The second model input terminal is used to receive the solvent molecule structure M. 2 The third model input terminal is used to receive the prior characteristic vector X, and the model output terminal is used to output the corresponding solvent characteristic prediction vector Y;
[0011] The multi-task prediction model has two optional structure options: a first optional structure and a second optional structure. The first optional structure includes the following components: a first structural feature encoder, a solute feature extraction module, a second structural feature encoder, a solvent feature extraction module, a molecular feature fusion module, a fusion feature encoder, a prior feature extraction module, a prior feature fusion module, a first regression prediction model, a second regression prediction model, a third regression prediction model, a fourth regression prediction model, a fifth regression prediction model, and a prediction output layer. The second optional structure includes the following components: a first structural feature encoder, a solute feature extraction module, a second structural feature encoder, a solvent feature extraction module, a molecular feature fusion module, a fusion feature encoder, a prior feature extraction module, a prior feature fusion module, and a fusion feature decoder.
[0012] The shared components of the first type of optional structure and the second type of optional structure are: the first structural feature encoder, the solute feature extraction module, the second structural feature encoder, the solvent feature extraction module, the molecular feature fusion module, the fusion feature encoder, the prior feature extraction module, and the prior feature fusion module;
[0013] The personalized components of the optional structure are: the first regression prediction model, the second regression prediction model, the third regression prediction model, the fourth regression prediction model, the fifth regression prediction model, and the prediction output layer;
[0014] The personalized component of the two optional structures is the fusion feature decoder.
[0015] Furthermore, the component connection relationships of the optional structure are as follows: the input end of the first structural feature encoder is connected to the input end of the first model, and its output end is connected to the input end of the solute feature extraction module; the output end of the solute feature extraction module is connected to the first input end of the molecular feature fusion module; the input end of the second structural feature encoder is connected to the input end of the second model, and its output end is connected to the input end of the solvent feature extraction module; the output end of the solvent feature extraction module is connected to the second input end of the molecular feature fusion module; the output end of the molecular feature fusion module is connected to the input end of the fusion feature encoder; the output end of the fusion feature encoder is connected to the first input end of the prior feature fusion module; the input end of the prior feature extraction module is connected to the input end of the third model, and its output end is connected to the second input end of the prior feature fusion module; the output end of the prior feature fusion module is connected to the input ends of the first, second, third, fourth, and fifth regression prediction models, respectively; the input end of the prediction output layer is connected to the output ends of the first, second, third, fourth, and fifth regression prediction models, respectively, and its output end is connected to the model output end.
[0016] The component connection relationships of the two optional structures are as follows: the input end of the first structural feature encoder is connected to the input end of the first model, and the output end is connected to the input end of the solute feature extraction module; the output end of the solute feature extraction module is connected to the first input end of the molecular feature fusion module; the input end of the second structural feature encoder is connected to the input end of the second model, and the output end is connected to the input end of the solvent feature extraction module; the output end of the solvent feature extraction module is connected to the second input end of the molecular feature fusion module; the output end of the molecular feature fusion module is connected to the input end of the fusion feature encoder; the output end of the fusion feature encoder is connected to the first input end of the prior feature fusion module; the input end of the prior feature extraction module is connected to the input end of the third model, and the output end is connected to the second input end of the prior feature fusion module; the output end of the prior feature fusion module is connected to the input end of the fusion feature decoder; and the output end of the fusion feature decoder is connected to the model output end.
[0017] Furthermore, the component functions of the shared components of the first type of optional structure and the second type of optional structure are as follows:
[0018] The first structural feature encoder is implemented based on the Uni-Mol model; the first structural feature encoder is used to analyze the solute molecule structure M. 1 Atomic-level high-dimensional feature encoding is performed to obtain the corresponding feature tensor E1, which is then sent to the solute feature extraction module; the feature tensor E1 has an N shape. 1 ×CA C A The preset atomic feature dimension;
[0019] The solute feature extraction module is used to perform pooling processing on each feature channel of the feature tensor E1 according to a preset solute feature pooling rule to obtain a vector of length C. A The pooled feature vector P1 is obtained; and a molecular feature space mapping is performed on the pooled feature vector P1 based on the built-in MLP model to obtain a vector of length C. B The molecular feature vector H1 is sent to the molecular feature fusion module; the solute feature pooling rules include max pooling, average pooling, and attention pooling; C B These are preset molecular feature dimensions;
[0020] The second structural feature encoder is implemented based on the Uni-Mol model; the second structural feature encoder is used to analyze the solvent molecule structure M. 2 Atom-level high-dimensional feature encoding processing is performed to obtain the corresponding feature tensor E2, which is then sent to the solvent feature extraction module; the shape of the feature tensor E2 is N. 2 ×C A ;
[0021] The solvent feature extraction module is used to perform pooling processing on each feature channel of the feature tensor E2 according to a preset solvent feature pooling rule to obtain a vector of length C. A The pooling feature vector P2 is obtained; and a molecular feature space mapping is performed on the pooling feature vector P2 based on the built-in MLP model to obtain a vector of length C. B The molecular feature vector H2 is sent to the molecular feature fusion module; the solvent feature pooling rules include max pooling, average pooling, and attention pooling;
[0022] The molecular feature fusion module is used to send a corresponding feature vector sequence H3, which is formed by sequentially sorting the molecular feature vectors H1 and H2, to the fusion feature encoder.
[0023] The fusion feature encoder is implemented based on the Transformer architecture encoder model; the fusion feature encoder is used to perform feature encoding processing on the feature vector sequence H3 through a multi-head self-attention encoding mechanism to obtain the corresponding encoded feature sequence H4, which is then sent to the prior feature fusion module; the encoded feature sequence H4 consists of two vectors of length C. B Feature encoding vector The feature encoding vector is formed by sequential sorting; Each of the molecular feature vectors H1 and H2 corresponds one-to-one;
[0024] The prior feature extraction module is used to perform prior feature space mapping on the prior feature vector X using a built-in MLP model to obtain a vector of length C. C The prior feature vector H5 is sent to the prior feature fusion module; C C The preset prior feature dimensions;
[0025] The prior feature fusion module is used to extract the corresponding feature encoding vector from the encoded feature sequence H4. And the feature encoding vector The vector is concatenated with the prior feature vector H5 to obtain a vector of length C. D The concatenated feature vector H6 is obtained; and in the first optional structure, the concatenated feature vector H6 is sent to the first regression prediction model, the second regression prediction model, the third regression prediction model, the fourth regression prediction model, and the fifth regression prediction model; and in the second optional structure, the concatenated feature vector H6 is sent to the fused feature decoder; C D C is the preset splicing feature dimension. D =C B +C C .
[0026] Furthermore, the component functions of the aforementioned optional structure personalized component are as follows:
[0027] The first regression prediction model is implemented based on a feedforward neural network; the first regression prediction model is used to predict the peak wavelength of the emission spectrum of solvent molecules based on the spliced feature vector H6 to obtain the corresponding first wavelength and send it to the prediction output layer;
[0028] The second regression prediction model is implemented based on a feedforward neural network; the second regression prediction model is used to predict the residence time of solvent molecules in the excited state based on the spliced feature vector H6 to obtain the corresponding first duration and send it to the prediction output layer;
[0029] The third regression prediction model is implemented based on a feedforward neural network; the third regression prediction model is used to predict the photoelectric conversion efficiency of solvent molecules based on the spliced feature vector H6 to obtain the corresponding first conversion efficiency and send it to the prediction output layer;
[0030] The fourth regression prediction model is implemented based on a feedforward neural network; the fourth regression prediction model is used to predict the peak wavelength of the solvent molecule absorption spectrum based on the spliced feature vector H6 to obtain the corresponding second wavelength and send it to the prediction output layer.
[0031] The fifth regression prediction model is implemented based on a feedforward neural network; the fifth regression prediction model is used to predict the half-width of the solvent molecule absorption spectrum based on the spliced feature vector H6 to obtain the corresponding first half-width and send it to the prediction output layer.
[0032] The prediction output layer takes the first wavelength, the first duration, the first conversion efficiency, the second wavelength, and the first half-width as the corresponding emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5 to form the corresponding solvent characteristic prediction vector Y and outputs it.
[0033] Furthermore, the component functions of the personalized components with the two optional structures are as follows:
[0034] The fusion feature decoder is implemented based on the decoder model of the Transformer architecture; the fusion feature decoder is used to convert the concatenated feature vector H6 into a feature encoding vector. The encoding feature sequence is composed of the prior feature vector H5; a corresponding decoding output sequence is initialized based on the data format of the solvent property prediction vector Y; and the output sequence is decoded according to the sequence decoding logic of the Transformer architecture decoder to obtain the corresponding solvent property prediction vector Y and output it.
[0035] Preferably, the model training dataset is denoted as the corresponding first dataset, and specifically includes:
[0036] The corresponding molecular pair datasets are obtained by collecting large amounts of solute-solvent molecular pairs in the field of optoelectronic materials through multiple preset data channels; the multiple data channels include publicly available electrolyte information databases, publicly available optoelectronic material information databases, and publicly available technical documents; the molecular pair datasets include multiple sets of the solute-solvent molecular pairs; the solute-solvent molecular pairs include solute molecule sequences and solvent molecule sequences; the solute molecule sequence and the solvent molecule sequence are each a one-dimensional SMILES sequence;
[0037] Each set of solute-solvent molecule pairs is taken as the corresponding current molecule pair;
[0038] Based on preset cheminformatics tools, a three-dimensional molecular conformation is created according to the solute molecule sequence and the solvent molecule sequence of the current molecular pair to obtain the corresponding current solute conformation and current solvent conformation; the cheminformatics tools include Open Babel software and RDKit software;
[0039] Based on preset molecular dynamics simulation tools, the stable conformations of the current solute and solvent are optimized to obtain corresponding optimized conformations of the current solute and solvent; and based on the molecular dynamics simulation tools, the three-dimensional solute-solvent system conformation composed of the optimized conformations of the current solute and solvent is simulated to obtain the corresponding current system conformation; the molecular dynamics simulation tools include Gaussian software and GROMACS software;
[0040] The element type and three-dimensional coordinates of each atom in the solute molecule in the current system conformation are extracted and used as the corresponding atomic element type and atomic three-dimensional coordinates to form the corresponding atomic features. And by all the atomic features obtained The corresponding solute molecular structure M 1 The element type and three-dimensional coordinates of each atom in the solvent molecule of the current system conformation are extracted and used as the corresponding atomic element type and atomic three-dimensional coordinates to form the corresponding atomic features. And by all the atomic features obtained The corresponding solvent molecular structure M 2 ; and the current solute molecular structure M 1 and the solvent molecular structure M 2 The first solute molecule structure and the first solvent molecule structure corresponding to the current molecule pair;
[0041] The system identifies the total number of aromatic rings of solute molecules in the current system conformation using the cheminformatics tool, and uses the identification result as the corresponding total number of aromatic rings of the solute x1; it calculates the interatomic interaction energy of solute molecules in the current system conformation using a preset quantum chemical calculation tool, and uses the calculation result as the corresponding interatomic interaction energy of the solute x2; it calculates the molecular dipole moment of solvent molecules in the current system conformation using the quantum chemical calculation tool, and uses the calculation result as the corresponding solvent dipole moment x3; it calculates the polarity index of solvent molecules in the current system conformation using the quantum chemical calculation tool, and uses the calculation result as the corresponding solvent polarity index x4; and it calculates the total number of aromatic rings of solute molecules in the current system conformation using the quantum chemical calculation tool. The orbital overlap integral of solute and solvent molecules in the system conformation is calculated, and the result is used as the corresponding solute-solvent orbital overlap integral x5; the total number of aromatic rings of the solute x1, the interatomic interaction energy of the solute x2, the dipole moment of the solvent x3, the polarity index of the solvent x4, and the solute-solvent orbital overlap integral x5 are used to form the corresponding a priori characteristic vector X; and the a priori characteristic vector X obtained in this case is used as the first a priori characteristic vector corresponding to the current molecule pair; the quantum chemical calculation tools include Multiwfn software, Gaussian software, ORCA software, GAMESS software, NWChem software, SHARC software, GPAW software, and Newton-X software;
[0042] The quantum chemical calculation tool is used to calculate the peak wavelength of the emission spectrum of solute molecules in the current system conformation and the result is taken as the corresponding emission peak y1. The quantum chemical calculation tool is also used to calculate the residence time of solvent molecules in the excited state in the current system conformation and the result is taken as the corresponding molecular lifetime y2. Furthermore, the quantum chemical calculation tool is used to calculate the photoelectric conversion efficiency of solvent molecules in the current system conformation and the result is taken as the corresponding photoelectric conversion efficiency y3. The quantum chemical calculation tool is also used to calculate the peak wavelength of the absorption spectrum of solvent molecules in the current system conformation and the result is taken as the corresponding absorption peak y4. Finally, the quantum chemical calculation tool is used to calculate the half-width at half-maximum (WHM) of the absorption spectrum of solvent molecules in the current system conformation and the result is taken as the corresponding absorption WHM y5. The obtained emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption WHM y5 form the corresponding solvent characteristic prediction vector Y. This solvent characteristic prediction vector Y is used as the first solvent characteristic label vector corresponding to the current molecule pair.
[0043] The first data record is composed of the first solute molecule structure, the first solvent molecule structure, the first prior property vector, and the first solvent property tag vector corresponding to the current molecule pair.
[0044] The first dataset is composed of all the first data records obtained.
[0045] Preferably, training the multi-task prediction model based on the first dataset specifically includes:
[0046] Step 91: Based on a preset first segmentation ratio, randomly divide the first dataset into two subsets, denoted as the first training set and the first evaluation set; and count the total number of records in the first training set to obtain the corresponding total number N. TR The total number N is obtained by statistically analyzing the total number of records in the first training set. EV ;
[0047] Both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first partitioning ratio, N TR :N EV ≈First division ratio;
[0048] Step 92: Take each of the first data records in the first training set as the corresponding current training record; and add noise to the first solute molecule structure, the first solvent molecule structure, and the first prior characteristic vector of the current training record to obtain the corresponding noisy solute molecule structure, noisy solvent molecule structure, and noisy prior characteristic vector; input the noisy solute molecule structure, the noisy solvent molecule structure, and the noisy prior characteristic vector into the multi-task prediction model for prediction, and take the solvent characteristic prediction vector Y obtained in this prediction as the corresponding first prediction vector; and form a corresponding first prediction-label pair by the first prediction vector and the first solvent characteristic label vector of the current training record.
[0049] The noise-adding method for the noisy solute molecule structure involves adding noise to some or all of the three-dimensional coordinates of the atoms in the first solute molecule structure, and the overall added noise should satisfy a Gaussian distribution rule with zero average noise. The noise-adding method for the noisy solvent molecule structure involves adding noise to all or some of the three-dimensional coordinates of the atoms in the first solvent molecule structure, and the overall added noise should satisfy a Gaussian distribution rule with zero average noise. The noise-adding method for the noisy prior characteristic vector involves adding noise to all vector data of the first prior characteristic vector, and the overall added noise should satisfy a Gaussian distribution rule with zero average noise.
[0050] Each of the first prediction vectors is denoted as the corresponding prediction vector Y. g Each of the first solvent characteristic label vectors is denoted as the corresponding label vector. 1≤index g≤N TR Each of the predicted vectors Y g The emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width at half maximum (WHM) y5 are denoted as the corresponding predicted values y. g,1 y g,2 y g,3 y g,4 y g,5 Each of the aforementioned label vectors The emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width at half maximum (WHM) y5 are denoted as the corresponding predicted values.
[0051] Step 93, obtain N TR Each of the first prediction-label pairs is input into a preset overall loss function L. total The calculation yields five corresponding classification loss values and an overall loss value;
[0052] Wherein, the overall loss function L total Specifically:
[0053] L total =w1·L1+w2·L2+w3·L3+w4·L4+w5·L5,
[0054]
[0055] w1, w2, w3, w4, and w5 are five preset weighting coefficients;
[0056] The five classification loss values include a first loss value, a second loss value, a third loss value, a fourth loss value, and a fifth loss value; the first loss value, the second loss value, the third loss value, the fourth loss value, and the fifth loss value are the loss values output by the corresponding first loss function L1, the second loss function L2, the third loss function L3, the fourth loss function L4, and the fifth loss function L5, respectively;
[0057] The overall loss value is the overall loss function L. total The output loss value;
[0058] Step 94: Identify whether the overall loss value meets the preset overall loss value range; if it does, proceed to step 95; if not, based on the preset first model optimizer, move towards making the overall loss function L... total The multi-task prediction model is subjected to one round of parameter modulation in the direction that reaches the minimum value, and the process returns to step 92 when the parameter modulation ends.
[0059] The first model optimizer includes the Adam optimizer and the SGD optimizer.
[0060] Step 95: Identify whether the first loss value, the second loss value, the third loss value, the fourth loss value, and the fifth loss value all satisfy their respective first loss value range, second loss value range, third loss value range, fourth loss value range, and fifth loss value range; if all satisfy, proceed to step 96; if at least one loss value does not satisfy, take each of the first loss value, second loss value, third loss value, fourth loss value, or fifth loss value that does not satisfy its corresponding loss value range as a corresponding loss value to be optimized, and assign the first loss function L1, second loss function L2, third loss function L3, and fourth loss function L4 to each loss value to be optimized. 4 or the fifth loss function L5 is used as the corresponding target loss function, and the first regression prediction model, the second regression prediction model, the third regression prediction model, the fourth regression prediction model or the fifth regression prediction model corresponding to each of the loss values to be optimized are used as the corresponding target regression models, and the model parameters of each of the target regression models are used as the corresponding target model parameters. A corresponding second model optimizer is assigned to each of the target regression models, and the corresponding target model parameters are modulated in one round based on each of the second model optimizers in the direction of minimizing the corresponding target loss function. After the parameter modulation of all the target model parameters in this round is completed, the process returns to step 92.
[0061] Each of the second model optimizers includes the Adam optimizer and the SGD optimizer;
[0062] Step 96: Perform a traversal of all the first data records in the first evaluation set; and during this traversal, take the currently traversed first data record as the corresponding current evaluation record; and take the first solute molecule structure, the first solvent molecule structure, and the first prior property vector of the current evaluation record as the current solute molecule structure M. 1 The solvent molecule structure M 2The prior characteristic vector X is input into the multi-task prediction model for prediction, and the solvent characteristic prediction vector Y obtained in this prediction is used as the corresponding second prediction vector; a corresponding second prediction-label pair is formed by the second prediction vector and the first solvent characteristic label vector of the current evaluation record; and at the end of this round of traversal, the obtained N EV The second prediction-label is fed into a preset first model evaluation function to obtain the corresponding first evaluation value;
[0063] The first model evaluation function is implemented based on the RMSE function;
[0064] Step 97: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 91 to continue training; if it does, confirm that the model training is complete.
[0065] Preferably, the step of using the multi-task prediction model to perform solvent optimization processing on the batch optimization task input by the user and feeding back the processing results to the current user specifically includes:
[0066] Extract the corresponding first solute molecule sequence, first solvent molecule sequence set, and the threshold range of the five categories of indicators from the batch optimization task;
[0067] Based on a pre-defined cheminformatics tool, a three-dimensional molecular conformation is created according to the first solute molecule sequence to obtain the corresponding first solute conformation; and the element type and three-dimensional coordinates of each atom in the first solute conformation are extracted as the corresponding atomic element type and atomic three-dimensional coordinates to form a corresponding atomic feature. And by all the atomic features obtained To form a corresponding solute molecular structure M 1 Based on the cheminformatics tool, the total number of aromatic rings in the first solute conformation is identified and the identification result is taken as a corresponding total number of aromatic rings in the solute x1; and based on the preset quantum chemical calculation tool, the interatomic interaction energy corresponding to the first solute conformation is calculated and the calculation result is taken as a corresponding interatomic interaction energy in the solute x2.
[0068] Based on the aforementioned cheminformatics tool, a three-dimensional molecular conformation is created according to each of the first solvent molecule sequences in the first solvent molecule sequence set to obtain the corresponding first solvent conformation; each of the first solvent conformations is used as the corresponding current solvent conformation; and the element type and three-dimensional coordinates of each atom in the current solvent conformation are extracted as the corresponding atom element type and atom three-dimensional coordinates to form a corresponding atom feature. And by all the atomic features obtained To form a corresponding solvent molecule structure M 2 Based on the quantum chemical calculation tool, the molecular dipole moment corresponding to the current solvent conformation is calculated, and the calculation result is taken as a corresponding solvent dipole moment x3; based on the quantum chemical calculation tool, the polarity index corresponding to the current solvent molecular structure is calculated, and the calculation result is taken as a corresponding solvent polarity index x4; based on the quantum chemical calculation tool, the solute-solvent orbital overlap integral corresponding to the first solute conformation and the current solvent conformation is calculated according to the GFN2-xTB method, and the calculation result is taken as a corresponding solute-solvent orbital overlap integral x5;
[0069] A priori characteristic vector X is formed by combining the total number of aromatic rings of the solute corresponding to each of the first solute molecule sequences (x1), the interatomic interaction energy of the solute (x2), the solvent dipole moment (x3), the solvent polarity index (x4), and the solute-solvent orbital overlap integral (x5); and the solute molecule structure M corresponding to each of the first solvent molecule sequences is further combined. 1 Solvent molecular structure M 2 The prior characteristic vector X is input into the multi-task prediction model to obtain a corresponding solvent characteristic prediction vector Y; and the first solvent molecule sequence corresponding to each solvent characteristic prediction vector Y that satisfies the threshold range of the five categories of indicators is recorded as the corresponding preferred molecule sequence.
[0070] Each preferred molecule sequence and its corresponding solvent property prediction vector Y are combined to form a preferred record; and all the preferred records are combined to form a corresponding solvent molecule preferred report, which is then fed back to the current user.
[0071] A second aspect of the present invention provides an apparatus for implementing the processing method for predicting the photoelectric properties of solvent molecules as described in the first aspect above, the apparatus comprising: a model building module, a model training module, and a model application module;
[0072] The model building module is used to construct a deep learning model that fuses features of solute molecule structure, solvent molecule structure, and prior properties of the solute-solvent system, and predicts multiple photoelectric properties of solvent molecules based on the fused features; denoted as the corresponding multi-task prediction model; wherein, the multi-task prediction model is used to predict multiple photoelectric properties of solvent molecules based on the solute molecule structure M input to the model. 1 Solvent molecular structure M 2The prior property vector X is used to perform multi-task prediction of the photoelectric properties of solvent molecules and output the corresponding solvent property prediction vector Y; both the solute molecule structure and the solvent molecule structure are three-dimensional molecular structures; the prior properties of the solute-solvent system include the total number of aromatic rings of the solute molecule, the interatomic interaction energy of the solute molecule, the dipole moment of the solvent molecule, the polarity index of the solvent molecule, and the orbital overlap integral of the solute molecule and the solvent molecule; the various photoelectric properties include emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum; the solute molecule structure M... 1 Characterized by multiple atoms Composition, 1≤index i≤N 1 N 1 The total number of atoms in the solute molecule; the atomic characteristics Includes atomic element types and atomic three-dimensional coordinates; the solvent molecule structure M 2 Characterized by multiple atoms Composition, 1≤indexj≤N 2 N 2 The total number of atoms in the solvent molecules; the atomic characteristics The a priori characteristic vector X includes the type of atomic element and the three-dimensional coordinates of the atom; the a priori characteristic vector X includes the total number of aromatic rings in the solute x1, the interatomic interaction energy of the solute x2, the solvent dipole moment x3, the solvent polarity index x4, and the solute-solvent orbital overlap integral x5; the solvent characteristic prediction vector Y includes the emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5.
[0073] The model training module is used to construct a model training dataset, denoted as the corresponding first dataset; and to train the multi-task prediction model based on the first dataset;
[0074] The model application module is used to perform solvent optimization processing on the batch optimization task input by the user after the model training is completed, and to feed back the processing results to the current user. The batch optimization task includes a first solute molecule sequence, a first solvent molecule sequence set, and five types of index threshold ranges. The first solvent molecule sequence set includes multiple first solvent molecule sequences. Each first solvent molecule sequence and each first solvent molecule sequence is a SMILES sequence. The five types of index threshold ranges include emission peak threshold range, molecular lifetime threshold range, photoelectric conversion efficiency threshold range, absorption peak threshold range, and absorption half-width threshold range.
[0075] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0076] The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above;
[0077] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0078] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.
[0079] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for predicting the photoelectric properties of solvent molecules. As described above, this invention uses five types of solute / solvent properties that can be quickly obtained through simple calculations as prior properties (total number of aromatic rings in the solute, interatomic interaction energy of the solute, solvent dipole moment, solvent polarity index, and solute-solvent orbital overlap integral calculated using the GFN2-xTB method). It also uses five types of photoelectric properties of solvent molecules that require lengthy calculations (emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum) as prediction targets. A multi-task prediction model is constructed that can predict the five prediction targets based on the solute molecule structure, solvent molecule structure, and the five types of prior properties. Two optional structures are provided for the multi-task prediction model. A model training dataset is constructed to train the model. After model training, the multi-task prediction model assists users in handling high-throughput solvent selection tasks. This invention not only reduces analytical complexity, shortens analysis time, and improves analytical efficiency, but also improves prediction accuracy, flexibility, and generalization of the prediction model. Attached Figure Description
[0080] Figure 1 This is a schematic diagram of a processing method for predicting the photoelectric properties of solvent molecules provided in Embodiment 1 of the present invention;
[0081] Figure 2 This is a module structure diagram of an optional structure of the multi-task prediction model provided in Embodiment 1 of the present invention;
[0082] Figure 3 This is a module structure diagram of the two optional structures of the multi-task prediction model provided in Embodiment 1 of the present invention;
[0083] Figure 4 This is a module structure diagram of a processing device for predicting the photoelectric properties of solvent molecules provided in Embodiment 2 of the present invention;
[0084] Figure 5This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0085] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0086] Embodiment 1 of the present invention provides a method for predicting the photoelectric properties of solvent molecules; such as Figure 1 The schematic diagram shows a method for predicting the photoelectric properties of solvent molecules provided in Embodiment 1 of the present invention. The method mainly includes the following steps:
[0087] Step 1: Construct a deep learning model that fuses features of solute molecule structure, solvent molecule structure, and prior properties of solute-solvent system, and predicts multiple photoelectric properties of solvent molecules based on the fused features. This model is denoted as the corresponding multi-task prediction model.
[0088] Here, the solute and solvent molecular structures in this embodiment of the invention are both three-dimensional molecular structures. The a priori properties of the solute-solvent system in this embodiment include the total number of aromatic rings in the solute molecule, the interatomic interaction energy of the solute molecule, the dipole moment of the solvent molecule, the polarity index of the solvent molecule, and the orbital overlap integral between the solute and solvent molecules. The various photoelectric properties in this embodiment include emission peaks, molecular lifetime, photoelectric conversion efficiency, absorption peaks, and absorption half-width at half-maximum (HWHM).
[0089] It should be noted that the five types of a priori properties of the solute-solvent system can be quickly obtained through computational tools: 1) Total number of aromatic rings in the solute molecule: This property parameter can be obtained by inputting the solute molecule structure into a preset cheminformatics tool to count the total number of aromatic rings. Under normal circumstances, this statistical process takes only seconds; 2) Interatomic interaction energy of the solute molecule: This property parameter can be obtained by inputting the solute molecule structure into a preset quantum chemical calculation tool. Under normal circumstances, this calculation process takes only minutes or even faster; 3) Dipole moment + polarity index of the solvent molecule: These two property parameters can be obtained simultaneously by inputting the solvent molecule structure into a quantum chemical calculation tool. Under normal circumstances, this calculation process takes only minutes or even faster; 4) Orbital overlap integral between solute and solvent molecules: This property parameter can be obtained by inputting the molecular structures of the solute and solvent into a quantum chemical calculation tool and calculating the orbital overlap integral between the solute and solvent molecules based on the GFN2-xTB method. Under normal circumstances, this calculation process takes only seconds.
[0090] It should also be noted that the various photoelectric properties (emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum) in the embodiments of this invention are the five types of prediction targets of this invention. Obtaining these five types of prediction targets using quantum chemical calculations is time-consuming: 1) The calculation process for emission peak, absorption peak, and absorption half-width at half-maximum is typically on the order of hours, and even for some complex molecular structures, the calculation process can reach the order of days; 2) The calculation process for molecular lifetime and photoelectric conversion efficiency is typically on the order of days or weeks. If high-throughput screening of solvent molecules based on these five types of prediction targets is to be performed using quantum chemical calculations, it is difficult to complete in a short time without sufficient computing power.
[0091] The multi-task prediction model of this invention is used to predict the solute molecule structure M based on the input model. 1 Solvent molecular structure M 2 The photoelectric properties of solvent molecules are predicted using the prior property vector X, and the corresponding solvent property prediction vector Y is output.
[0092] Among them: 1) Solute molecular structure M 1 Characterized by multiple atoms Composition, 1≤index i≤N 1 N 1 The total number of atoms in a solute molecule; atomic characteristics. This includes atomic element types and atomic three-dimensional coordinates. 2) Solvent molecular structure M 2 Characterized by multiple atoms Composition, 1≤indexj≤N 2 N 2 The total number of atoms in a solvent molecule; atomic characteristics 3) The prior characteristic vector X includes the total number of aromatic rings in the solute (x1), the interatomic interaction energy of the solute (x2), the solvent dipole moment (x3), the solvent polarity index (x4), and the solute-solvent orbital overlap integral (x5). 4) The solvent characteristic prediction vector Y includes the emission peak (y1), molecular lifetime (y2), photoelectric conversion efficiency (y3), absorption peak (y4), and absorption half-width (WHM) (y5).
[0093] The first model input of the multi-task prediction model in this embodiment of the invention is used to receive the solute molecular structure M. 1 The second model input is used to receive the solvent molecule structure M. 2 The third model input is used to receive the prior characteristic vector X, and the model output is used to output the corresponding solvent characteristic prediction vector Y.
[0094] It should be noted that the embodiments of the present invention provide two types of optional model structures for the multi-task prediction model, namely, a first type of optional structure and a second type of optional structure; the first type of optional structure is as follows: Figure 2 This is a module structure diagram of one type of optional structure of the multi-task prediction model provided in Embodiment 1 of the present invention. The second type of optional structure is as follows: Figure 3 The module structure diagram of the two optional structures of the multi-task prediction model provided in Embodiment 1 of the present invention is shown.
[0095] like Figure 2 As shown, a model component with an optional structure according to an embodiment of the present invention includes: a first structural feature encoder, a solute feature extraction module, a second structural feature encoder, a solvent feature extraction module, a molecular feature fusion module, a fusion feature encoder, a priori feature extraction module, a priori feature fusion module, a first regression prediction model, a second regression prediction model, a third regression prediction model, a fourth regression prediction model, a fifth regression prediction model, and a prediction output layer.
[0096] The component connection relationships of this optional structure are as follows: the input end of the first structural feature encoder is connected to the input end of the first model, and the output end is connected to the input end of the solute feature extraction module; the output end of the solute feature extraction module is connected to the first input end of the molecular feature fusion module; the input end of the second structural feature encoder is connected to the input end of the second model, and the output end is connected to the input end of the solvent feature extraction module; the output end of the solvent feature extraction module is connected to the second input end of the molecular feature fusion module; the output end of the molecular feature fusion module is connected to the input end of the fusion feature encoder; the output end of the fusion feature encoder is connected to the first input end of the prior feature fusion module; the input end of the prior feature extraction module is connected to the input end of the third model, and the output end is connected to the second input end of the prior feature fusion module; the output end of the prior feature fusion module is connected to the input ends of the first, second, third, fourth, and fifth regression prediction models, respectively; the input end of the prediction output layer is connected to the output ends of the first, second, third, fourth, and fifth regression prediction models, respectively, and the output end is connected to the model output end.
[0097] like Figure 3 As shown, the model components of the second type of optional structure in this embodiment include: a first structural feature encoder, a solute feature extraction module, a second structural feature encoder, a solvent feature extraction module, a molecular feature fusion module, a fusion feature encoder, a priori feature extraction module, a priori feature fusion module, and a fusion feature decoder.
[0098] The component connection relationships of the two optional structures are as follows: the input end of the first structural feature encoder is connected to the input end of the first model, and the output end is connected to the input end of the solute feature extraction module; the output end of the solute feature extraction module is connected to the first input end of the molecular feature fusion module; the input end of the second structural feature encoder is connected to the input end of the second model, and the output end is connected to the input end of the solvent feature extraction module; the output end of the solvent feature extraction module is connected to the second input end of the molecular feature fusion module; the output end of the molecular feature fusion module is connected to the input end of the fusion feature encoder; the output end of the fusion feature encoder is connected to the first input end of the prior feature fusion module; the input end of the prior feature extraction module is connected to the input end of the third model, and the output end is connected to the second input end of the prior feature fusion module; the output end of the prior feature fusion module is connected to the input end of the fusion feature decoder; and the output end of the fusion feature decoder is connected to the model output end.
[0099] like Figure 2 , Figure 3 As shown in the embodiments of the present invention, the first and second type optional structures have eight shared components, namely: a first structural feature encoder, a solute feature extraction module, a second structural feature encoder, a solvent feature extraction module, a molecular feature fusion module, a fusion feature encoder, a priori feature extraction module, and a priori feature fusion module. These eight shared components have the same function in both types of optional structures, as detailed below.
[0100] 1) First structural feature encoder:
[0101] The first structural feature encoder is implemented based on the Uni-Mol model. The first structural feature encoder is used to analyze the solute molecular structure M... 1 Atomic-level high-dimensional feature encoding is performed to obtain the corresponding feature tensor E1, which is then sent to the solute feature extraction module. The feature tensor E1 has a shape of N. 1 ×C A C A This refers to the preset atomic feature dimensions.
[0102] Here, the Uni-Mol model of the first structural feature encoder has been pre-trained. Detailed model structure and pre-training methods of the Uni-Mol model can be found in the publicly available technical document "Uni-Mol: A Universal 3D Dolecular Representation Learning Framework," and will not be elaborated further here. The reason why this embodiment uses the Uni-Mol model as the solute molecule structure M... 1 The feature encoder is chosen because the model can refine the precision of the encoded features to the atomic level, which helps to further improve the prediction accuracy of the final prediction properties.
[0103] 2) Solute Feature Extraction Module:
[0104] The solute feature extraction module is used to perform pooling processing on each feature channel of the feature tensor E1 according to the preset solute feature pooling rules to obtain a vector of length C. A The pooling feature vector P1 is obtained; and a molecular feature space mapping is performed on the pooling feature vector P1 based on the built-in MLP model to obtain a vector of length C. B The molecular feature vector H1 is sent to the molecular feature fusion module.
[0105] Here, the solute feature pooling rules in this embodiment of the invention include max pooling, average pooling, and attention pooling; the max pooling rule is to pool N features in each feature channel. 1 The maximum value within each feature data point is used as the pooling value for the current feature channel. The average pooling rule is to use the maximum value within each feature channel as the pooling value. 1 The average value within each feature data point is used as the pooling value for the current feature channel. The attention pooling rule is to first pool the N features in each feature channel. 1 The attention weights of each feature data are calculated, and then based on the obtained N... 1 Each attention weight is paired with N 1 The feature data is weighted and summed, and the result is used as the pooling value for the current feature channel; in specific implementation, one of these three types of rules can be selected as the adaptation rule based on specific implementation requirements; C B The preset molecular feature dimensions.
[0106] 3) Second structural feature encoder:
[0107] The second structural feature encoder is implemented based on the Uni-Mol model. The second structural feature encoder is used to analyze the solvent molecule structure M. 2 Atomic-level high-dimensional feature encoding is performed to obtain the corresponding feature tensor E2, which is then sent to the solvent feature extraction module. The feature tensor E2 has a shape of N. 2 ×C A .
[0108] Here, the Uni-Mol model of the second structural feature encoder has also been pre-trained. This embodiment of the invention uses the Uni-Mol model as the solvent molecule structure M. 2 The feature encoder is useful because the model can refine the precision of encoded features to the atomic level, thereby helping to further improve the prediction accuracy of the final prediction properties.
[0109] 4) Solvent Feature Extraction Module:
[0110] The solvent feature extraction module is used to perform pooling processing on each feature channel of the feature tensor E2 according to a preset solvent feature pooling rule to obtain a vector of length C. A The pooling feature vector P2 is obtained; and a molecular feature space mapping is performed on the pooling feature vector P2 based on the built-in MLP model to obtain a vector of length C. B The molecular feature vector H2 is sent to the molecular feature fusion module.
[0111] Here, the solvent feature pooling rules in this embodiment of the invention include max pooling, average pooling, and attention pooling. In specific implementation, one of these three types of rules can be selected as the adaptation rule based on specific implementation requirements.
[0112] 5) Molecular Feature Fusion Module:
[0113] The molecular feature fusion module is used to send a corresponding feature vector sequence H3, which is formed by sequentially sorting molecular feature vectors H1 and H2, to the fusion feature encoder.
[0114] 6) Fusion Feature Encoder:
[0115] The fusion feature encoder is implemented based on the Transformer architecture encoder model. It uses a multi-head self-attention encoding mechanism to encode the feature vector sequence H3, obtaining the corresponding encoded feature sequence H4, which is then sent to the prior feature fusion module. The encoded feature sequence H4 consists of two vectors of length C. B Feature encoding vector Arranged in sequence; feature encoding vector Each corresponds one-to-one with the molecular feature vectors H1 and H2.
[0116] It should be noted that, according to publicly available information about the Transformer architecture, the encoder model in this architecture actually associates each token feature vector in the input features with all token feature vectors through a series of multi-head attention operations. In this embodiment of the invention, the molecular feature vectors H1 and H2 in the feature vector sequence H3 can be regarded as two token feature vectors. After fusing a series of multi-head attention operations of the feature encoder, the molecular feature vector H2 can be associated with the molecular feature vector H1 to obtain a corresponding feature encoding vector. By associating the molecular feature vector H1 with the molecular feature vector H2, a corresponding feature encoding vector is obtained. Among them, the feature encoding vector In essence, it is a solvent molecule feature that incorporates solute characteristics. Using this molecular feature with solute information to predict properties helps improve prediction accuracy.
[0117] 7) Prior Feature Extraction Module:
[0118] The prior feature extraction module uses a built-in MLP model to perform a prior feature space mapping on the prior feature vector X to obtain a vector of length C. C The prior feature vector H5 is sent to the prior feature fusion module. Wherein, C C The preset prior feature dimension.
[0119] 8) Prior Feature Fusion Module:
[0120] The prior feature fusion module is used to extract the corresponding feature encoding vector from the encoded feature sequence H4. And the feature encoding vector Concatenate the prior feature vector H5 with the vector to obtain a vector of length C. D The concatenated feature vector H6 is obtained; and in one optional structure, the concatenated feature vector H6 is sent to the first regression prediction model, the second regression prediction model, the third regression prediction model, the fourth regression prediction model, and the fifth regression prediction model; and in another optional structure, the concatenated feature vector H6 is sent to the fused feature decoder; C D C is the preset splicing feature dimension. D =C B +C C .
[0121] Here, as can be seen from the preceding text, the feature encoding vector A solvent molecule feature that incorporates solute characteristic information is represented by a feature encoding vector. The concatenated feature vector H6, formed by concatenating the prior feature vector H5, can be regarded as a solvent molecule feature that integrates solute and prior information.
[0122] like Figure 2 As shown, the optional structure of the personalized components in this embodiment of the invention includes: a first regression prediction model, a second regression prediction model, a third regression prediction model, a fourth regression prediction model, a fifth regression prediction model, and a prediction output layer. The component functions of the six optional structure personalized components are as follows.
[0123] 1) First regression prediction model:
[0124] The first regression prediction model is implemented based on a feedforward neural network. The first regression prediction model is used to predict the peak wavelength of the emission spectrum of solvent molecules based on the spliced feature vector H6 to obtain the corresponding first wavelength, which is then sent to the prediction output layer.
[0125] 2) Second regression prediction model:
[0126] The second regression prediction model is implemented based on a feedforward neural network. The second regression prediction model is used to predict the residence time of solvent molecules in the excited state based on the spliced feature vector H6 to obtain the corresponding first duration, which is then sent to the prediction output layer.
[0127] 3) Third regression prediction model:
[0128] The third regression prediction model is implemented based on a feedforward neural network. The third regression prediction model is used to predict the photoelectric conversion efficiency of solvent molecules based on the spliced feature vector H6, and then send the corresponding first conversion efficiency to the prediction output layer.
[0129] 4) Fourth regression prediction model:
[0130] The fourth regression prediction model is based on a feedforward neural network. The fourth regression prediction model is used to predict the peak wavelength of the solvent molecule absorption spectrum based on the spliced feature vector H6 to obtain the corresponding second wavelength, which is then sent to the prediction output layer.
[0131] 5) Fifth regression prediction model:
[0132] The fifth regression prediction model is implemented based on a feedforward neural network. The fifth regression prediction model is used to predict the half-width of the solvent molecule absorption spectrum based on the spliced feature vector H6, and then send the corresponding first half-width to the prediction output layer.
[0133] 6) Predicting the output layer:
[0134] The predicted output layer uses the first wavelength, first duration, first conversion efficiency, second wavelength, and first half-width as the corresponding emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5 to form the corresponding solvent characteristic prediction vector Y and outputs it.
[0135] like Figure 3 As shown, the personalized component of the optional structure in Embodiment 2 of the present invention is a fusion feature decoder. The component functions of this decoder are as follows.
[0136] The fusion feature decoder in this embodiment of the invention is implemented based on the decoder model of the Transformer architecture. This fusion feature decoder is used to convert the concatenated feature vector H6 into a feature-encoded vector. The encoding feature sequence is composed of the prior feature vector H5; a corresponding decoding output sequence is initialized based on the data format of the solvent property prediction vector Y; and the output sequence is decoded according to the sequence decoding logic of the Transformer architecture decoder to obtain the corresponding solvent property prediction vector Y and output it.
[0137] It should be noted that, based on the publicly available decoder model structure of the Transformer architecture, the fusion feature decoder consists of multiple decoder blocks connected sequentially to a linear prediction network. Each decoder block comprises a multi-head self-attention unit and a cross-attention unit. The input to the fusion feature decoder is the encoded feature sequence (denoted as encoded feature sequence A) and the initialized decoded output sequence, and the output is the solvent property prediction vector Y. From the publicly available decoding logic of the Transformer architecture, the decoding logic of the fusion feature decoder can be simply described as follows: first, embedding feature encoding is performed on the initialized decoded output sequence to obtain an initialized encoded feature sequence Z. 0 and encode the feature sequence Z 0 The input is fed into the first decoding module for computation; the multi-head self-attention unit of the first decoding module is used to process the encoded feature sequence Z. 0 Perform multi-head self-attention computation and use the result as the corresponding self-attention encoding feature sequence S. 1 Send to the cross-attention unit of this module; the cross-attention unit of the first decoding module is used to process the encoded feature sequence A and the self-attention encoded feature sequence S. 1 Perform cross-attention operation and use the result as the new encoded feature sequence Z. 1 Send to the second decoding module; the multi-head self-attention unit of the second decoding module is used to encode the feature sequence Z. 1 Perform multi-head self-attention computation and use the result as the corresponding self-attention encoding feature sequence S. 2 The signal is sent to the cross-attention unit of this module. The cross-attention unit of the second decoding module is used to process the encoded feature sequence A and the self-attention encoded feature sequence S. 2 Perform cross-attention operation and use the result as the new encoded feature sequence Z. 2 The signal is sent to the third decoding module; and so on, until the cross-attention unit of the last decoding module takes the cross-attention operation result as the final encoded feature sequence Z. end Send to the linear prediction network; the linear prediction network is then used to determine the encoded feature sequence Z. end Perform five types of solvent property prediction and output the corresponding solvent property prediction vector Y.
[0138] It should also be noted that although the embodiments of the present invention provide two optional structures for the multi-task prediction model, before processing prediction tasks based on the multi-task prediction model, users need to select one of these two optional structures as the base model based on user habits and / or specific task implementation requirements. After confirming the base model, training is performed first through step 2, and then applied through step 3 after training is completed.
[0139] Step 2: Construct the model training dataset, denoted as the first dataset; and train the multi-task prediction model based on the first dataset.
[0140] Specifically, this includes: Step 21, constructing the model training dataset, denoted as the corresponding first dataset;
[0141] The first dataset includes multiple first data records; each first data record includes a first solute molecule structure, a first solvent molecule structure, a first prior property vector, and a first solvent property label vector; the data structures of the first solute molecule structure, the first solvent molecule structure, the first prior property vector, and the first solvent property label vector are related to the corresponding solute molecule structure M. 1 Solvent molecular structure M 2 The prior characteristic vector X and the solvent characteristic prediction vector Y must be kept consistent.
[0142] Specifically, it includes: Step 211, collecting large data on solute-solvent molecule pairs in the field of optoelectronic materials through multiple preset data channels to obtain the corresponding molecule pair dataset;
[0143] Here, multiple data sources are included, such as publicly available electrolyte databases, publicly available optoelectronic materials databases, and publicly available technical documents; the molecular pair dataset includes multiple sets of solute-solvent molecular pairs; the solute-solvent molecular pairs include solute molecule sequences and solvent molecule sequences; each solute molecule sequence and solvent molecule sequence is a one-dimensional SMILES sequence;
[0144] Step 212, and take each set of solute-solvent molecule pairs as the corresponding current molecule pairs;
[0145] Step 213, and based on the preset cheminformatics tools, create a three-dimensional molecular conformation according to the solute molecule sequence and solvent molecule sequence of the current molecular pair to obtain the corresponding current solute conformation and current solvent conformation.
[0146] Here, the cheminformatics tools in this embodiment of the invention include Open Babel software and RDKit software;
[0147] Step 214: Based on the preset molecular dynamics simulation tool, optimize the stable conformations of the current solute conformation and the current solvent conformation to obtain the corresponding optimized conformation of the current solute and the current solvent; and based on the molecular dynamics simulation tool, simulate the three-dimensional solute-solvent system conformation composed of the optimized conformation of the current solute and the optimized conformation of the current solvent to obtain the corresponding current system conformation.
[0148] Here, the molecular dynamics simulation tools in this embodiment of the invention include Gaussian software and GROMACS software;
[0149] Step 215: Extract the element type and three-dimensional coordinates of each atom in the solute molecule in the current system conformation as the corresponding atomic element type and atomic three-dimensional coordinates to form the corresponding atomic features. And from all the atomic features obtained The corresponding solute molecular structure M 1 Furthermore, the element type and three-dimensional coordinates of each atom in the solvent molecule of the current system conformation are extracted and used as the corresponding atomic element type and three-dimensional coordinates to form the corresponding atomic features. And from all the atomic features obtained The corresponding solvent molecule structure M 2 ; and the current solute molecular structure M 1 and solvent molecular structure M 2 As the first solute molecular structure and the first solvent molecular structure corresponding to the current molecular pair;
[0150] Step 216 involves identifying the total number of aromatic rings in the solute molecules of the current system conformation using cheminformatics tools and using the identification result as the corresponding total number of aromatic rings in the solute x1; calculating the interatomic interaction energy of the solute molecules of the current system conformation using a preset quantum chemical calculation tool and using the calculation result as the corresponding interatomic interaction energy of the solute x2; calculating the molecular dipole moment of the solvent molecules of the current system conformation using a quantum chemical calculation tool and using the calculation result as the corresponding solvent dipole moment x3; and calculating the molecular dipole moment of the solvent molecules of the current system conformation using a quantum chemical calculation tool. The polarity index is calculated and the result is used as the corresponding solvent polarity index x4; the orbital overlap integral of solute and solvent molecules in the current system conformation is calculated using quantum chemical calculation tools and the result is used as the corresponding solute-solvent orbital overlap integral x5; the total number of aromatic rings of solute x1, the interatomic interaction energy of solute x2, the solvent dipole moment x3, the solvent polarity index x4, and the solute-solvent orbital overlap integral x5 are used to form the corresponding a priori characteristic vector X; and the a priori characteristic vector X obtained in this process is used as the first a priori characteristic vector corresponding to the current molecule pair.
[0151] Here, the quantum chemical calculation tools in this embodiment of the invention include Multiwfn software, Gaussian software, ORCA software, GAMESS software, NWChem software, SHARC software, GPAW software, and Newton-X software;
[0152] Step 217 involves calculating the peak wavelength of the emission spectrum of solute molecules in the current system conformation using quantum chemical calculation tools, and using the calculation result as the corresponding emission peak y1; calculating the residence time of solvent molecules in the excited state in the current system conformation using quantum chemical calculation tools, and using the calculation result as the corresponding molecular lifetime y2; calculating the photoelectric conversion efficiency of solvent molecules in the current system conformation using quantum chemical calculation tools, and using the calculation result as the corresponding photoelectric conversion efficiency y3; calculating the peak wavelength of the absorption spectrum of solvent molecules in the current system conformation using quantum chemical calculation tools, and using the calculation result as the corresponding absorption peak y4; calculating the half-width at half-maximum (WHM) of the absorption spectrum of solvent molecules in the current system conformation using quantum chemical calculation tools, and using the calculation result as the corresponding absorption WHM y5; and using the obtained emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption WHM y5 to form the corresponding solvent characteristic prediction vector Y; and using the solvent characteristic prediction vector Y obtained in this step as the first solvent characteristic label vector corresponding to the current molecule pair.
[0153] Step 218, and a corresponding first data record is formed by the first solute molecule structure, the first solvent molecule structure, the first prior property vector, and the first solvent property tag vector corresponding to the current molecule pair;
[0154] Step 219, and the first dataset is composed of all the first data records obtained;
[0155] Step 22, and train the multi-task prediction model based on the first dataset;
[0156] Specifically, this includes: Step 221, randomly dividing the first dataset into two subsets based on a preset first segmentation ratio, denoted as the first training set and the first evaluation set; and calculating the total number of records in the first training set to obtain the corresponding total number N. TR The total number N is obtained by statistically analyzing the total number of records in the first training set. EV ;
[0157] Wherein, the first segmentation ratio is a pre-set ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio, N TR :N EV ≈First division ratio;
[0158] Step 222: Take each first data record of the first training set as the corresponding current training record; and add noise to the first solute molecular structure, first solvent molecular structure, and first prior characteristic vector of the current training record to obtain the corresponding noisy solute molecular structure, noisy solvent molecular structure, and noisy prior characteristic vector; input the noisy solute molecular structure, noisy solvent molecular structure, and noisy prior characteristic vector into the multi-task prediction model for prediction, and take the solvent characteristic prediction vector Y obtained in this prediction as the corresponding first prediction vector; and form a corresponding first prediction-label pair by the first prediction vector and the first solvent characteristic label vector of the current training record.
[0159] Here, the noise-adding method for the solute molecule structure is to add noise to some or all of the three-dimensional coordinates of the atoms in the first solute molecule structure, and the overall added noise should satisfy the Gaussian distribution rule with zero average noise; the noise-adding method for the solvent molecule structure is to add noise to all or some of the three-dimensional coordinates of the atoms in the first solvent molecule structure, and the overall added noise should satisfy the Gaussian distribution rule with zero average noise; the noise-adding method for the prior characteristic vector is to add noise to all vector data of the first prior characteristic vector, and the overall added noise should satisfy the Gaussian distribution rule with zero average noise; it should be noted that the reason for adding noise to the first solute molecule structure, the first solvent molecule structure, and the first prior characteristic vector is to improve the generalization ability of the model through model training;
[0160] Furthermore, each of the first prediction vectors in the embodiments of the present invention is denoted as the corresponding prediction vector Y. g Each first solvent characteristic label vector can be denoted as the corresponding label vector. 1≤index g≤N TR Each prediction vector Y g The emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width at half maximum (WHM) y5 are then denoted as the corresponding predicted values y. g,1 y g,2 y g,3 y g,4 y g,5 Each label vector The emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width at half maximum (HWHM) y5 are then recorded as the corresponding predicted values.
[0161] Step 223, obtain N TR Each first prediction-label pair is input into the preset overall loss function L. total The calculation yields five corresponding classification loss values and an overall loss value;
[0162] Here, the overall loss function L in this embodiment of the invention total Specifically:
[0163] L total =w1·L1+w2·L2+w3·L3+w4·L4+w5·L5,
[0164]
[0165] Where w1, w2, w3, w4, and w5 are five preset weight coefficients; L1, L2, L3, L4, and L5 are the corresponding first, second, third, fourth, and fifth loss functions; the five classification loss values include the first, second, third, fourth, and fifth loss values; the first, second, third, fourth, and fifth loss values are the loss values output by the corresponding first loss function L1, second loss function L2, third loss function L3, fourth loss function L4, and fifth loss function L5, respectively; the overall loss value is the overall loss function L... total The output loss value;
[0166] Step 224: Identify whether the overall loss value meets the preset overall loss value range; if it does, proceed to step 225; if not, based on the preset first model optimizer, move towards making the overall loss function L... total The direction that reaches the minimum value is used to perform one round of parameter modulation on the multi-task prediction model, and the process returns to step 222 when the parameter modulation ends.
[0167] The overall loss value range is a pre-set numerical range; the first model optimizer includes the Adam optimizer and the SGD optimizer.
[0168] Step 225: Identify whether the first, second, third, fourth, and fifth loss values all satisfy their respective first, second, third, fourth, and fifth loss value ranges. If all satisfy, proceed to step 226. If at least one loss value does not satisfy, treat each first, second, third, fourth, or fifth loss value that does not satisfy its corresponding loss value range as a corresponding loss value to be optimized, and assign the corresponding first loss function L1, second loss function L2, third loss function L3, and fourth loss function L4 to each loss value to be optimized. Alternatively, the fifth loss function L5 can be used as the corresponding target loss function, and the first regression prediction model, second regression prediction model, third regression prediction model, fourth regression prediction model or fifth regression prediction model corresponding to each loss value to be optimized can be used as the corresponding target regression model. The model parameters of each target regression model can be used as the corresponding target model parameters, and a corresponding second model optimizer can be assigned to each target regression model. Based on each second model optimizer, the corresponding target model parameters can be modulated in one round in the direction of minimizing the corresponding target loss function. After the parameter modulation of all target model parameters is completed, return to step 222.
[0169] Here, the first loss value range, the second loss value range, the third loss value range, the fourth loss value range, and the fifth loss value range are five pre-set numerical ranges; each second model optimizer includes the Adam optimizer and the SGD optimizer;
[0170] Step 226: Perform a traversal of all first data records in the first evaluation set; during this traversal, take the currently traversed first data record as the corresponding current evaluation record; and take the first solute molecule structure, first solvent molecule structure, and first prior characteristic vector of the current evaluation record as the current solute molecule structure M. 1 Solvent molecular structure M 2 The prior characteristic vector X is input into the multi-task prediction model for prediction, and the solvent characteristic prediction vector Y obtained from this prediction is used as the corresponding second prediction vector; a corresponding second prediction-label pair is formed by the second prediction vector and the first solvent characteristic label vector of the current evaluation record; and at the end of this round of traversal, the obtained N is... EV Each second prediction-label is input into a preset first model evaluation function to calculate the corresponding first evaluation value;
[0171] Here, the first model evaluation function is implemented based on the RMSE function;
[0172] Step 227: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 221 to continue training; if it meets the range, confirm that the model training is complete.
[0173] Here, the first evaluation value range is a pre-set numerical range.
[0174] Step 3: After the model training is completed, the multi-task prediction model is used to perform solvent optimization processing on the batch optimization task input by the user and the processing results are fed back to the current user.
[0175] The batch optimization task includes a first solute molecule sequence, a first solvent molecule sequence set, and five types of index threshold ranges; the first solvent molecule sequence set includes multiple first solvent molecule sequences; each first solvent molecule sequence and each first solvent molecule sequence is a SMILES sequence; the five types of index threshold ranges include emission peak threshold range, molecular lifetime threshold range, photoelectric conversion efficiency threshold range, absorption peak threshold range, and absorption half-width threshold range.
[0176] Specifically, this includes: Step 31, extracting the corresponding first solute molecule sequence, first solvent molecule sequence set and five types of index threshold ranges from the batch optimization task;
[0177] Step 32: Based on cheminformatics tools, a three-dimensional molecular conformation is created according to the first solute molecule sequence to obtain the corresponding first solute conformation; and the element type and three-dimensional coordinates of each atom in the first solute conformation are extracted as the corresponding atomic element type and atomic three-dimensional coordinates to form a corresponding atomic feature. And from all the atomic features obtained To form a corresponding solute molecular structure M 1 Based on cheminformatics tools, the total number of aromatic rings in the first solute conformation is identified and the identification result is used as a corresponding total number of aromatic rings in the solute x1; based on quantum chemical calculation tools, the interatomic interaction energy corresponding to the first solute conformation is calculated and the calculation result is used as a corresponding interatomic interaction energy in the solute x2.
[0178] Step 33: Based on cheminformatics tools, three-dimensional molecular conformations are created according to each first solvent molecule sequence in the first solvent molecule sequence set to obtain the corresponding first solvent conformation; each first solvent conformation is used as the corresponding current solvent conformation; and the element type and three-dimensional coordinates of each atom in the current solvent conformation are extracted as the corresponding atomic element type and atomic three-dimensional coordinates to form a corresponding atomic feature. And from all the atomic features obtained To form a corresponding solvent molecule structure M 2Based on quantum chemical calculation tools, the molecular dipole moment corresponding to the current solvent conformation is calculated, and the calculation result is used as a corresponding solvent dipole moment x3; based on quantum chemical calculation tools, the polarity index corresponding to the current solvent molecular structure is calculated, and the calculation result is used as a corresponding solvent polarity index x4; based on quantum chemical calculation tools, the solute-solvent orbital overlap integral corresponding to the first solute conformation and the current solvent conformation is calculated according to the GFN2-xTB method, and the calculation result is used as a corresponding solute-solvent orbital overlap integral x5;
[0179] Step 34, and compose a corresponding prior characteristic vector X from the total number of aromatic rings of the first solute molecule sequence and each first solvent molecule sequence, x1, the interatomic interaction energy of the solute x2, the solvent dipole moment x3, the solvent polarity index x4, and the solute-solvent orbital overlap integral x5; and compose a set of solute molecule structures M corresponding to each first solvent molecule sequence. 1 Solvent molecular structure M 2 The prior characteristic vector X is input into the multi-task prediction model to obtain a corresponding solvent characteristic prediction vector Y; and the first solvent molecule sequence corresponding to each solvent characteristic prediction vector Y that satisfies the threshold range of the five categories of indicators is recorded as the corresponding preferred molecule sequence.
[0180] Step 35, and a corresponding preferred record is formed by each preferred molecule sequence and its corresponding solvent property prediction vector Y; and a corresponding solvent molecule preferred report is formed by all the obtained preferred records and fed back to the current user.
[0181] Figure 4 This is a module structure diagram of a processing device for predicting the photoelectric properties of solvent molecules according to Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 4 As shown, the device includes: a model building module 201, a model training module 202, and a model application module 203.
[0182] Model building module 201 is used to construct a deep learning model that fuses features of solute molecule structure, solvent molecule structure, and prior properties of the solute-solvent system, and predicts multiple photoelectric properties of solvent molecules based on the fused features; denoted as the corresponding multi-task prediction model; wherein, the multi-task prediction model is used to predict multiple photoelectric properties of solvent molecules based on the solute molecule structure M input to the model. 1 Solvent molecular structure M 2The prior property vector X is used to perform multi-task prediction of the photoelectric properties of solvent molecules and output the corresponding solvent property prediction vector Y; both solute and solvent molecules have three-dimensional molecular structures; prior properties of the solute-solvent system include the total number of aromatic rings in the solute molecule, the interatomic interaction energy of the solute molecule, the dipole moment of the solvent molecule, the polarity index of the solvent molecule, and the orbital overlap integral of the solute and solvent molecules; multiple photoelectric properties include emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum; solute molecular structure M 1 Characterized by multiple atoms Composition, 1≤index i≤N 1 N 1 The total number of atoms in a solute molecule; atomic characteristics. Includes atomic element type, atomic three-dimensional coordinates; solvent molecule structure M 2 Characterized by multiple atoms Composition, 1≤indexj≤N 2 N 2 The total number of atoms in a solvent molecule; atomic characteristics It includes atomic element type and atomic three-dimensional coordinates; the prior characteristic vector X includes the total number of aromatic rings in the solute x1, the interaction energy between solute atoms x2, the solvent dipole moment x3, the solvent polarity index x4, and the solute-solvent orbital overlap integral x5; the solvent characteristic prediction vector Y includes the emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5.
[0183] The model training module 202 is used to construct the model training dataset, denoted as the corresponding first dataset; and to train the multi-task prediction model based on the first dataset.
[0184] The model application module 203 is used to perform solvent optimization processing on the batch optimization task input by the user after the model training is completed, and to feed back the processing results to the current user. The batch optimization task includes a first solute molecule sequence, a first solvent molecule sequence set, and five types of index threshold ranges. The first solvent molecule sequence set includes multiple first solvent molecule sequences. Each first solvent molecule sequence and each first solvent molecule sequence is a SMILES sequence. The five types of index threshold ranges include emission peak threshold range, molecular lifetime threshold range, photoelectric conversion efficiency threshold range, absorption peak threshold range, and absorption half-width threshold range.
[0185] The present invention provides a processing device for predicting the photoelectric properties of solvent molecules, which can execute the method steps in the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0186] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing elements; they can be fully implemented in hardware; or some modules can be implemented by processing elements calling software, while others are implemented in hardware. For example, the model building module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0187] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-on-a-Chip (SOC).
[0188] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0189] Figure 5 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 5 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.
[0190] exist Figure 5The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0191] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0192] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.
[0193] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for predicting the photoelectric properties of solvent molecules. As described above, this invention uses five types of solute / solvent properties that can be quickly obtained through simple calculations as prior properties (total number of aromatic rings in the solute, interatomic interaction energy of the solute, solvent dipole moment, solvent polarity index, and solute-solvent orbital overlap integral calculated using the GFN2-xTB method). It also uses five types of photoelectric properties of solvent molecules that require lengthy calculations (emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum) as prediction targets. A multi-task prediction model is constructed that can predict the five prediction targets based on the solute molecule structure, solvent molecule structure, and the five types of prior properties. Two optional structures are provided for the multi-task prediction model. A model training dataset is constructed to train the model. After model training, the multi-task prediction model assists users in handling high-throughput solvent selection tasks. This invention not only reduces analytical complexity, shortens analysis time, and improves analytical efficiency, but also improves prediction accuracy, flexibility, and generalization of the prediction model.
[0194] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0195] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for predicting the photoelectric properties of solvent molecules, characterized in that, The method includes: A deep learning model is constructed that fuses features of solute molecule structure, solvent molecule structure, and prior properties of the solute-solvent system, and predicts multiple photoelectric properties of solvent molecules based on the fused features. This model is denoted as the corresponding multi-task prediction model. The multi-task prediction model is used to predict multiple photoelectric properties of solvent molecules based on the solute molecule structure M input to the model. 1 Solvent molecular structure M 2 The prior property vector X is used to perform multi-task prediction of the photoelectric properties of solvent molecules and output the corresponding solvent property prediction vector Y; both the solute molecule structure and the solvent molecule structure are three-dimensional molecular structures; the prior properties of the solute-solvent system include the total number of aromatic rings of the solute molecule, the interatomic interaction energy of the solute molecule, the dipole moment of the solvent molecule, the polarity index of the solvent molecule, and the orbital overlap integral of the solute molecule and the solvent molecule; the various photoelectric properties include emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum; the solute molecule structure M... 1 Characterized by multiple atoms Composition, 1≤index i≤N 1 N 1 The total number of atoms in the solute molecule; the atomic characteristics Includes atomic element types and atomic three-dimensional coordinates; the solvent molecule structure M 2 Characterized by multiple atoms Composition, 1≤indexj≤N 2 N 2 The total number of atoms in the solvent molecules; the atomic characteristics The a priori characteristic vector X includes the type of atomic element and the three-dimensional coordinates of the atom; the a priori characteristic vector X includes the total number of aromatic rings in the solute x1, the interatomic interaction energy of the solute x2, the solvent dipole moment x3, the solvent polarity index x4, and the solute-solvent orbital overlap integral x5; the solvent characteristic prediction vector Y includes the emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5. The model training dataset is constructed and denoted as the corresponding first dataset; and the multi-task prediction model is trained based on the first dataset; After model training, the multi-task prediction model is used to perform solvent optimization on the batch optimization task input by the user and the processing results are fed back to the current user. The batch optimization task includes a first solute molecule sequence, a first solvent molecule sequence set, and five types of index threshold ranges. The first solvent molecule sequence set includes multiple first solvent molecule sequences. Each first solvent molecule sequence and each first solvent molecule sequence is a SMILES sequence. The five types of index threshold ranges include emission peak threshold range, molecular lifetime threshold range, photoelectric conversion efficiency threshold range, absorption peak threshold range, and absorption half-width threshold range. Specifically, the step of using the multi-task prediction model to perform solvent optimization processing on the batch optimization task input by the user and feeding back the processing results to the current user includes: Extract the corresponding first solute molecule sequence, first solvent molecule sequence set, and the threshold range of the five categories of indicators from the batch optimization task; Based on a pre-defined cheminformatics tool, a three-dimensional molecular conformation is created according to the first solute molecule sequence to obtain the corresponding first solute conformation; and the element type and three-dimensional coordinates of each atom in the first solute conformation are extracted as the corresponding atomic element type and atomic three-dimensional coordinates to form a corresponding atomic feature. And by all the atomic features obtained To form a corresponding solute molecular structure M 1 Based on the cheminformatics tool, the total number of aromatic rings in the first solute conformation is identified and the identification result is taken as a corresponding total number of aromatic rings in the solute x1; and based on the preset quantum chemical calculation tool, the interatomic interaction energy corresponding to the first solute conformation is calculated and the calculation result is taken as a corresponding interatomic interaction energy in the solute x2. Based on the aforementioned cheminformatics tool, a three-dimensional molecular conformation is created according to each of the first solvent molecule sequences in the first solvent molecule sequence set to obtain the corresponding first solvent conformation; each of the first solvent conformations is used as the corresponding current solvent conformation; and the element type and three-dimensional coordinates of each atom in the current solvent conformation are extracted as the corresponding atom element type and atom three-dimensional coordinates to form a corresponding atom feature. And by all the atomic features obtained To form a corresponding solvent molecule structure M 2 Based on the quantum chemical calculation tool, the molecular dipole moment corresponding to the current solvent conformation is calculated, and the calculation result is taken as a corresponding solvent dipole moment x3; based on the quantum chemical calculation tool, the polarity index corresponding to the current solvent molecular structure is calculated, and the calculation result is taken as a corresponding solvent polarity index x4; based on the quantum chemical calculation tool, the solute-solvent orbital overlap integral corresponding to the first solute conformation and the current solvent conformation is calculated according to the GFN2-xTB method, and the calculation result is taken as a corresponding solute-solvent orbital overlap integral x5; A priori characteristic vector X is formed by combining the total number of aromatic rings of the solute corresponding to each of the first solute molecule sequences (x1), the interatomic interaction energy of the solute (x2), the solvent dipole moment (x3), the solvent polarity index (x4), and the solute-solvent orbital overlap integral (x5); and the solute molecule structure M corresponding to each of the first solvent molecule sequences is further combined. 1 Solvent molecular structure M 2 The prior characteristic vector X is input into the multi-task prediction model to obtain a corresponding solvent characteristic prediction vector Y; and the first solvent molecule sequence corresponding to each solvent characteristic prediction vector Y that satisfies the threshold range of the five categories of indicators is recorded as the corresponding preferred molecule sequence. Each preferred molecule sequence and its corresponding solvent property prediction vector Y are combined to form a preferred record; and all the preferred records are combined to form a corresponding solvent molecule preferred report, which is then fed back to the current user.
2. The processing method for predicting the photoelectric properties of solvent molecules according to claim 1, characterized in that, The first dataset includes multiple first data records; each first data record includes a first solute molecule structure, a first solvent molecule structure, a first prior property vector, and a first solvent property label vector; the data structures of the first solute molecule structure, the first solvent molecule structure, the first prior property vector, and the first solvent property label vector are related to the corresponding solute molecule structure M. 1 The solvent molecule structure M 2 The prior characteristic vector X and the solvent characteristic prediction vector Y are consistent.
3. The processing method for predicting the photoelectric properties of solvent molecules according to claim 1, characterized in that, The first model input of the multi-task prediction model is used to receive the solute molecule structure M. 1 The second model input terminal is used to receive the solvent molecule structure M. 2 The third model input terminal is used to receive the prior characteristic vector X, and the model output terminal is used to output the corresponding solvent characteristic prediction vector Y; The multi-task prediction model has two optional structures: a first-class optional structure and a second-class optional structure. The first type of optional structure model component includes: a first structural feature encoder, a solute feature extraction module, a second structural feature encoder, a solvent feature extraction module, a molecular feature fusion module, a fusion feature encoder, a priori feature extraction module, a priori feature fusion module, a first regression prediction model, a second regression prediction model, a third regression prediction model, a fourth regression prediction model, a fifth regression prediction model, and a prediction output layer; the second type of optional structure model component includes: a first structural feature encoder, a solute feature extraction module, a second structural feature encoder, a solvent feature extraction module, a molecular feature fusion module, a fusion feature encoder, a priori feature extraction module, a priori feature fusion module, and a fusion feature decoder; The shared components of the first type of optional structure and the second type of optional structure are: the first structural feature encoder, the solute feature extraction module, the second structural feature encoder, the solvent feature extraction module, the molecular feature fusion module, the fusion feature encoder, the prior feature extraction module, and the prior feature fusion module; The personalized components of the optional structure are: the first regression prediction model, the second regression prediction model, the third regression prediction model, the fourth regression prediction model, the fifth regression prediction model, and the prediction output layer; The personalized component of the two optional structures is the fusion feature decoder.
4. The processing method for predicting the photoelectric properties of solvent molecules according to claim 3, characterized in that, The component connection relationships of the optional structure are as follows: the input end of the first structural feature encoder is connected to the input end of the first model, and the output end is connected to the input end of the solute feature extraction module; the output end of the solute feature extraction module is connected to the first input end of the molecular feature fusion module; the input end of the second structural feature encoder is connected to the input end of the second model, and the output end is connected to the input end of the solvent feature extraction module; the output end of the solvent feature extraction module is connected to the second input end of the molecular feature fusion module; the output end of the molecular feature fusion module is connected to the input end of the fusion feature encoder; the output end of the fusion feature encoder is connected to the first input end of the prior feature fusion module; the input end of the prior feature extraction module is connected to the input end of the third model, and the output end is connected to the second input end of the prior feature fusion module; the output end of the prior feature fusion module is connected to the input ends of the first, second, third, fourth, and fifth regression prediction models respectively; the input end of the prediction output layer is connected to the output ends of the first, second, third, fourth, and fifth regression prediction models respectively, and the output end is connected to the model output end. The component connection relationships of the two optional structures are as follows: the input end of the first structural feature encoder is connected to the input end of the first model, and the output end is connected to the input end of the solute feature extraction module; the output end of the solute feature extraction module is connected to the first input end of the molecular feature fusion module; the input end of the second structural feature encoder is connected to the input end of the second model, and the output end is connected to the input end of the solvent feature extraction module; the output end of the solvent feature extraction module is connected to the second input end of the molecular feature fusion module; the output end of the molecular feature fusion module is connected to the input end of the fusion feature encoder; the output end of the fusion feature encoder is connected to the first input end of the prior feature fusion module; the input end of the prior feature extraction module is connected to the input end of the third model, and the output end is connected to the second input end of the prior feature fusion module; the output end of the prior feature fusion module is connected to the input end of the fusion feature decoder; and the output end of the fusion feature decoder is connected to the model output end.
5. The processing method for predicting the photoelectric properties of solvent molecules according to claim 3, characterized in that, The shared component functions of the first type of optional structure and the second type of optional structure are as follows: The first structural feature encoder is implemented based on the Uni-Mol model; the first structural feature encoder is used to analyze the solute molecule structure M. 1 Atomic-level high-dimensional feature encoding is performed to obtain the corresponding feature tensor E1, which is then sent to the solute feature extraction module; the feature tensor E1 has an N shape. 1 ×C A C A The preset atomic feature dimension; The solute feature extraction module is used to perform pooling processing on each feature channel of the feature tensor E1 according to a preset solute feature pooling rule to obtain a vector of length C. A The pooled feature vector P1 is obtained; and a molecular feature space mapping is performed on the pooled feature vector P1 based on the built-in MLP model to obtain a vector of length C. B The molecular feature vector H1 is sent to the molecular feature fusion module; the solute feature pooling rules include max pooling, average pooling, and attention pooling; C B These are preset molecular feature dimensions; The second structural feature encoder is implemented based on the Uni-Mol model; the second structural feature encoder is used to analyze the solvent molecule structure M. 2 Atom-level high-dimensional feature encoding processing is performed to obtain the corresponding feature tensor E2, which is then sent to the solvent feature extraction module; the shape of the feature tensor E2 is N. 2 ×C A ; The solvent feature extraction module is used to perform pooling processing on each feature channel of the feature tensor E2 according to a preset solvent feature pooling rule to obtain a vector of length C. A The pooling feature vector P2 is obtained; and a molecular feature space mapping is performed on the pooling feature vector P2 based on the built-in MLP model to obtain a vector of length C. B The molecular feature vector H2 is sent to the molecular feature fusion module; the solvent feature pooling rules include max pooling, average pooling, and attention pooling; The molecular feature fusion module is used to send a corresponding feature vector sequence H3, which is formed by sequentially sorting the molecular feature vectors H1 and H2, to the fusion feature encoder. The fusion feature encoder is implemented based on the Transformer architecture encoder model; the fusion feature encoder is used to perform feature encoding processing on the feature vector sequence H3 through a multi-head self-attention encoding mechanism to obtain the corresponding encoded feature sequence H4, which is then sent to the prior feature fusion module; the encoded feature sequence H4 consists of two vectors of length C. B Feature encoding vector , The feature encoding vector is formed by sequential sorting; , Each of the molecular feature vectors H1 and H2 corresponds one-to-one; The prior feature extraction module is used to perform prior feature space mapping on the prior feature vector X using a built-in MLP model to obtain a vector of length C. C The prior feature vector H5 is sent to the prior feature fusion module; C C The preset prior feature dimensions; The prior feature fusion module is used to extract the corresponding feature encoding vector from the encoded feature sequence H4. ; and the feature encoding vector The vector is concatenated with the prior feature vector H5 to obtain a vector of length C. D The concatenated feature vector H6 is obtained; and in the first optional structure, the concatenated feature vector H6 is sent to the first regression prediction model, the second regression prediction model, the third regression prediction model, the fourth regression prediction model, and the fifth regression prediction model; and in the second optional structure, the concatenated feature vector H6 is sent to the fused feature decoder; C D C is the preset splicing feature dimension. D =C B +C C .
6. The processing method for predicting the photoelectric properties of solvent molecules according to claim 5, characterized in that, The component functions of the aforementioned optional structure personalized component are as follows: The first regression prediction model is implemented based on a feedforward neural network; the first regression prediction model is used to predict the peak wavelength of the emission spectrum of solvent molecules based on the spliced feature vector H6 to obtain the corresponding first wavelength and send it to the prediction output layer; The second regression prediction model is implemented based on a feedforward neural network; the second regression prediction model is used to predict the residence time of solvent molecules in the excited state based on the spliced feature vector H6 to obtain the corresponding first duration and send it to the prediction output layer; The third regression prediction model is implemented based on a feedforward neural network; the third regression prediction model is used to predict the photoelectric conversion efficiency of solvent molecules based on the spliced feature vector H6 to obtain the corresponding first conversion efficiency and send it to the prediction output layer; The fourth regression prediction model is implemented based on a feedforward neural network; the fourth regression prediction model is used to predict the peak wavelength of the solvent molecule absorption spectrum based on the spliced feature vector H6 to obtain the corresponding second wavelength and send it to the prediction output layer. The fifth regression prediction model is implemented based on a feedforward neural network; the fifth regression prediction model is used to predict the half-width of the solvent molecule absorption spectrum based on the spliced feature vector H6 to obtain the corresponding first half-width and send it to the prediction output layer. The prediction output layer takes the first wavelength, the first duration, the first conversion efficiency, the second wavelength, and the first half-width as the corresponding emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5 to form the corresponding solvent characteristic prediction vector Y and outputs it.
7. The processing method for predicting the photoelectric properties of solvent molecules according to claim 5, characterized in that, The component functions of the personalized components with the two optional structures are as follows: The fusion feature decoder is implemented based on the decoder model of the Transformer architecture; the fusion feature decoder is used to convert the concatenated feature vector H6 into a feature encoding vector. The encoding feature sequence is composed of the prior feature vector H5; a corresponding decoding output sequence is initialized based on the data format of the solvent property prediction vector Y; and the output sequence is decoded according to the sequence decoding logic of the Transformer architecture decoder to obtain the corresponding solvent property prediction vector Y and output it.
8. The processing method for predicting the photoelectric properties of solvent molecules according to claim 2, characterized in that, The training dataset for building the model is denoted as the first dataset, and specifically includes: The corresponding molecular pair datasets are obtained by collecting large amounts of solute-solvent molecular pairs in the field of optoelectronic materials through multiple preset data channels; the multiple data channels include publicly available electrolyte information databases, publicly available optoelectronic material information databases, and publicly available technical documents; the molecular pair datasets include multiple sets of the solute-solvent molecular pairs; the solute-solvent molecular pairs include solute molecule sequences and solvent molecule sequences; the solute molecule sequence and the solvent molecule sequence are each a one-dimensional SMILES sequence; Each set of solute-solvent molecule pairs is taken as the corresponding current molecule pair; Based on preset cheminformatics tools, a three-dimensional molecular conformation is created according to the solute molecule sequence and the solvent molecule sequence of the current molecular pair to obtain the corresponding current solute conformation and current solvent conformation; the cheminformatics tools include Open Babel software and RDKit software; Based on preset molecular dynamics simulation tools, the stable conformations of the current solute and solvent are optimized to obtain corresponding optimized conformations of the current solute and solvent; and based on the molecular dynamics simulation tools, the three-dimensional solute-solvent system conformation composed of the optimized conformations of the current solute and solvent is simulated to obtain the corresponding current system conformation; the molecular dynamics simulation tools include Gaussian software and GROMACS software; The element type and three-dimensional coordinates of each atom in the solute molecule in the current system conformation are extracted and used as the corresponding atomic element type and atomic three-dimensional coordinates to form the corresponding atomic features. And by all the atomic features obtained The corresponding solute molecular structure M 1 The element type and three-dimensional coordinates of each atom in the solvent molecule of the current system conformation are extracted and used as the corresponding atomic element type and atomic three-dimensional coordinates to form the corresponding atomic features. And by all the atomic features obtained The corresponding solvent molecular structure M 2 ; and the current solute molecular structure M 1 and the solvent molecular structure M 2 The first solute molecule structure and the first solvent molecule structure corresponding to the current molecule pair; The system identifies the total number of aromatic rings of solute molecules in the current system conformation using the cheminformatics tool, and uses the identification result as the corresponding total number of aromatic rings of the solute x1; it calculates the interatomic interaction energy of solute molecules in the current system conformation using a preset quantum chemical calculation tool, and uses the calculation result as the corresponding interatomic interaction energy of the solute x2; it calculates the molecular dipole moment of solvent molecules in the current system conformation using the quantum chemical calculation tool, and uses the calculation result as the corresponding solvent dipole moment x3; it calculates the polarity index of solvent molecules in the current system conformation using the quantum chemical calculation tool, and uses the calculation result as the corresponding solvent polarity index x4; and it calculates the total number of aromatic rings of solute molecules in the current system conformation using the quantum chemical calculation tool. The orbital overlap integral of solute and solvent molecules in the system conformation is calculated, and the result is used as the corresponding solute-solvent orbital overlap integral x5; the total number of aromatic rings of the solute x1, the interatomic interaction energy of the solute x2, the dipole moment of the solvent x3, the polarity index of the solvent x4, and the solute-solvent orbital overlap integral x5 are used to form the corresponding a priori characteristic vector X; and the a priori characteristic vector X obtained in this case is used as the first a priori characteristic vector corresponding to the current molecule pair; the quantum chemical calculation tools include Multiwfn software, Gaussian software, ORCA software, GAMESS software, NWChem software, SHARC software, GPAW software, and Newton-X software; The quantum chemical calculation tool is used to calculate the peak wavelength of the emission spectrum of solute molecules in the current system conformation and the result is taken as the corresponding emission peak y1. The quantum chemical calculation tool is also used to calculate the residence time of solvent molecules in the excited state in the current system conformation and the result is taken as the corresponding molecular lifetime y2. Furthermore, the quantum chemical calculation tool is used to calculate the photoelectric conversion efficiency of solvent molecules in the current system conformation and the result is taken as the corresponding photoelectric conversion efficiency y3. The quantum chemical calculation tool is also used to calculate the peak wavelength of the absorption spectrum of solvent molecules in the current system conformation and the result is taken as the corresponding absorption peak y4. Finally, the quantum chemical calculation tool is used to calculate the half-width at half-maximum (WHM) of the absorption spectrum of solvent molecules in the current system conformation and the result is taken as the corresponding absorption WHM y5. The obtained emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption WHM y5 form the corresponding solvent characteristic prediction vector Y. This solvent characteristic prediction vector Y is used as the first solvent characteristic label vector corresponding to the current molecule pair. The first data record is composed of the first solute molecule structure, the first solvent molecule structure, the first prior property vector, and the first solvent property tag vector corresponding to the current molecule pair. The first dataset is composed of all the first data records obtained.
9. The processing method for predicting the photoelectric properties of solvent molecules according to claim 2, characterized in that, The training of the multi-task prediction model based on the first dataset specifically includes: Step 91: Based on a preset first segmentation ratio, randomly divide the first dataset into two subsets, denoted as the first training set and the first evaluation set; and count the total number of records in the first training set to obtain the corresponding total number N. TR The total number N is obtained by statistically analyzing the total number of records in the first training set. EV ; Both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first partitioning ratio, N TR :N EV ≈First division ratio; Step 92: Take each of the first data records in the first training set as the corresponding current training record; and add noise to the first solute molecule structure, the first solvent molecule structure, and the first prior characteristic vector of the current training record to obtain the corresponding noisy solute molecule structure, noisy solvent molecule structure, and noisy prior characteristic vector; input the noisy solute molecule structure, the noisy solvent molecule structure, and the noisy prior characteristic vector into the multi-task prediction model for prediction, and take the solvent characteristic prediction vector Y obtained in this prediction as the corresponding first prediction vector; and form a corresponding first prediction-label pair by the first prediction vector and the first solvent characteristic label vector of the current training record. The noise-adding method for the noisy solute molecule structure involves adding noise to some or all of the three-dimensional coordinates of the atoms in the first solute molecule structure, and the overall added noise should satisfy a Gaussian distribution rule with zero average noise. The noise-adding method for the noisy solvent molecule structure involves adding noise to all or some of the three-dimensional coordinates of the atoms in the first solvent molecule structure, and the overall added noise should satisfy a Gaussian distribution rule with zero average noise. The noise-adding method for the noisy prior characteristic vector involves adding noise to all vector data of the first prior characteristic vector, and the overall added noise should satisfy a Gaussian distribution rule with zero average noise. Each of the first prediction vectors is denoted as the corresponding prediction vector Y. g Each of the first solvent characteristic label vectors is denoted as the corresponding label vector. 1 ≤ index g ≤ N TR Each of the predicted vectors Y g The emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width at half maximum (WHM) y5 are denoted as the corresponding predicted values y. g,1 y g,2 y g,3 y g,4 y g,5 Each of the aforementioned label vectors The emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width at half maximum (WHM) y5 are denoted as the corresponding predicted values. , , , , ; Step 93, obtain N TR Each of the first prediction-label pairs is input into a preset overall loss function L. total The calculation yields five corresponding classification loss values and an overall loss value; Wherein, the overall loss function L total Specifically: , , , , , ; w1, w2, w3, w4, and w5 are five preset weight coefficients; The five classification loss values include a first loss value, a second loss value, a third loss value, a fourth loss value, and a fifth loss value; the first loss value, the second loss value, the third loss value, the fourth loss value, and the fifth loss value are the loss values output by the corresponding first loss function L1, the second loss function L2, the third loss function L3, the fourth loss function L4, and the fifth loss function L5, respectively; The overall loss value is the overall loss function L. total The output loss value; Step 94: Identify whether the overall loss value meets the preset overall loss value range; if it does, proceed to step 95; if not, based on the preset first model optimizer, move towards making the overall loss function L... total The multi-task prediction model is subjected to one round of parameter modulation in the direction that reaches the minimum value, and the process returns to step 92 when the parameter modulation ends. The first model optimizer includes the Adam optimizer and the SGD optimizer; Step 95: Identify whether the first loss value, the second loss value, the third loss value, the fourth loss value, and the fifth loss value all satisfy their respective first loss value range, second loss value range, third loss value range, fourth loss value range, and fifth loss value range; if all satisfy, proceed to step 96; if at least one loss value does not satisfy, take each of the first loss value, second loss value, third loss value, fourth loss value, or fifth loss value that does not satisfy its corresponding loss value range as a corresponding loss value to be optimized, and assign the first loss function L1, the second loss function L2, the third loss function L3, and the fourth loss function L4 corresponding to each loss value to be optimized. Loss function L4 or the fifth loss function L5 is used as the corresponding target loss function. The first regression prediction model, the second regression prediction model, the third regression prediction model, the fourth regression prediction model or the fifth regression prediction model corresponding to each of the loss values to be optimized are used as the corresponding target regression models. The model parameters of each of the target regression models are used as the corresponding target model parameters. A corresponding second model optimizer is assigned to each of the target regression models. Based on each of the second model optimizers, the corresponding target model parameters are modulated in one round in the direction of minimizing the corresponding target loss function. After the parameter modulation of all the target model parameters in this round is completed, the process returns to step 92. Each of the second model optimizers includes the Adam optimizer and the SGD optimizer; Step 96: Perform a traversal of all the first data records in the first evaluation set; and during this traversal, take the currently traversed first data record as the corresponding current evaluation record; and take the first solute molecule structure, the first solvent molecule structure, and the first prior property vector of the current evaluation record as the current solute molecule structure M. 1 The solvent molecule structure M 2 The prior characteristic vector X is input into the multi-task prediction model for prediction, and the solvent characteristic prediction vector Y obtained in this prediction is used as the corresponding second prediction vector; a corresponding second prediction-label pair is formed by the second prediction vector and the first solvent characteristic label vector of the current evaluation record; and at the end of this round of traversal, the obtained N EV The second prediction-label is fed into a preset first model evaluation function to calculate the corresponding first evaluation value; The first model evaluation function is implemented based on the RMSE function; Step 97: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 91 to continue training; if it does, confirm that the model training is complete.
10. An apparatus for performing the processing method for predicting the photoelectric properties of solvent molecules according to any one of claims 1-9, characterized in that, The device includes: a model building module, a model training module, and a model application module; The model building module is used to construct a deep learning model that fuses features of solute molecule structure, solvent molecule structure, and prior properties of the solute-solvent system, and predicts multiple photoelectric properties of solvent molecules based on the fused features; denoted as the corresponding multi-task prediction model; wherein, the multi-task prediction model is used to predict multiple photoelectric properties of solvent molecules based on the solute molecule structure M input to the model. 1 Solvent molecular structure M 2 The prior property vector X is used to perform multi-task prediction of the photoelectric properties of solvent molecules and output the corresponding solvent property prediction vector Y; both the solute molecule structure and the solvent molecule structure are three-dimensional molecular structures; the prior properties of the solute-solvent system include the total number of aromatic rings of the solute molecule, the interatomic interaction energy of the solute molecule, the dipole moment of the solvent molecule, the polarity index of the solvent molecule, and the orbital overlap integral of the solute molecule and the solvent molecule; the various photoelectric properties include emission peak, molecular lifetime, photoelectric conversion efficiency, absorption peak, and absorption half-width at half-maximum; the solute molecule structure M... 1 Characterized by multiple atoms Composition, 1≤index i≤N 1 N 1 The total number of atoms in the solute molecule; the atomic characteristics Includes atomic element types and atomic three-dimensional coordinates; the solvent molecule structure M 2 Characterized by multiple atoms Composition, 1≤indexj≤N 2 N 2 The total number of atoms in the solvent molecules; the atomic characteristics The a priori characteristic vector X includes the type of atomic element and the three-dimensional coordinates of the atom; the a priori characteristic vector X includes the total number of aromatic rings in the solute x1, the interatomic interaction energy of the solute x2, the solvent dipole moment x3, the solvent polarity index x4, and the solute-solvent orbital overlap integral x5; the solvent characteristic prediction vector Y includes the emission peak y1, molecular lifetime y2, photoelectric conversion efficiency y3, absorption peak y4, and absorption half-width y5. The model training module is used to construct a model training dataset, denoted as the corresponding first dataset; and to train the multi-task prediction model based on the first dataset; The model application module is used to perform solvent optimization processing on the batch optimization task input by the user after the model training is completed, and to feed back the processing results to the current user. The batch optimization task includes a first solute molecule sequence, a first solvent molecule sequence set, and five types of index threshold ranges. The first solvent molecule sequence set includes multiple first solvent molecule sequences. Each first solvent molecule sequence and each first solvent molecule sequence is a SMILES sequence. The five types of index threshold ranges include emission peak threshold range, molecular lifetime threshold range, photoelectric conversion efficiency threshold range, absorption peak threshold range, and absorption half-width threshold range.
11. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-9; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1-9.
Citation Information
Patent Citations
Processing method and device of organic molecule maximum absorption peak prediction model
CN120183561A
Method and device for predicting optical energy characteristics of organic photovoltaic material molecules
CN120183563A