A processing method and device of a two-molecule transfer integral prediction model
By constructing a deep learning model and utilizing the bimolecular properties for feature fusion, the problems of long calculation time and low efficiency of bimolecular transfer integrals in existing technologies are solved, achieving fast and accurate bimolecular transfer integral prediction and optimizing the high-throughput screening process for optoelectronic materials.
Patent Information
- Application Number
- CN202510933552.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-07-08
AI Technical Summary
In existing technologies, the computation time and efficiency of bimolecular transfer integrals based on quantum chemical calculations are long, resulting in low high-throughput screening efficiency for thin film/crystal structure optoelectronic materials and hindering the iterative development of new materials.
A deep learning model is constructed, which uses bimolecular properties as prior characteristics. The bimolecular transfer integral is predicted through feature fusion. The model parameters are trained by combining a dataset of thin film/crystal structures to achieve fast and accurate prediction of bimolecular transfer integral.
Shorten analysis time, improve analysis efficiency, save computing resources, optimize screening efficiency, and provide support for the research and development of new materials.
Smart Images

Figure CN120783879B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a processing method and apparatus for a bimolecular transfer integral prediction model. Background Technology
[0002] The bimolecular transfer integral (BTI) is a parameter in quantum mechanics describing the strength of the electronic interaction between two molecules, and is a key physical quantity describing the electron transfer or transition process between the two molecules. In the design of optoelectronic materials with thin film or crystal structures, the BTI is one of the key indicators for characterizing the optoelectronic properties of the materials. Conventionally, the BTI is mostly calculated using quantum chemical calculation methods. The biggest drawback of this conventional method is its long computation time and low efficiency. For example, if quantum chemical calculation tools are used based on density functional theory (DFT) or time-dependent density functional theory (TD-DFT), the computation time is measured in hours or days. If high-throughput screening of thin film / crystal structure optoelectronic material molecules is performed using this conventional method, even with a large amount of computing resources, it is difficult to optimize the screening efficiency, thus hindering the acceleration of the research and development of new materials. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies by providing a method, apparatus, electronic device, and computer-readable storage medium for processing bimolecular transfer integral prediction models. This invention uses six types of bimolecular properties that can be quickly obtained through simple calculations as prior properties: the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral and the intermolecular interaction energy calculated based on semi-empirical methods (such as the GFN2-xTB method), and uses the bimolecular transfer integral as the prediction target. Based on this, a deep learning model (bimolecular transfer integral prediction model) is constructed that can predict the bimolecular transfer integral based on two three-dimensional molecular structures (first molecular structure M1, second molecular structure M2) and the six types of prior properties (bimolecular prior property vector X). Corresponding first / second datasets are constructed by collecting bimolecular information of thin-film / crystal structure optoelectronic materials. The bimolecular transfer integral prediction model is then trained on the first / second datasets to obtain two sets of model parameters (thin-film model parameter set, crystal model parameter set) for the two types of structures. After model training, the two sets of model parameters from the bimolecular transfer integral prediction model are used to predict the bimolecular transfer integral of any bimolecular system in the field of thin-film / crystal structure optoelectronic materials. The invention can shorten the analysis time, improve the analysis efficiency, and improve the prediction accuracy. Based on the invention, high-throughput screening of optoelectronic material molecules with thin film / crystal structure can effectively save computing resources and optimize screening efficiency, providing effective assistance for accelerating the research and development of new materials.
[0004] To achieve the above objectives, a first aspect of the present invention provides a method for processing a bimolecular transfer integral prediction model, the method comprising:
[0005] A deep learning model is constructed to fuse the structure and prior properties of bimolecules and predict bimolecule transfer integrals based on the fused features, denoted as the corresponding bimolecule transfer integral prediction model. This model fuses structural features of the first molecular structure M1 and the second molecular structure M2 input to the model, fuses features based on the bimolecule fusion features and the prior property vector X input to the model, and predicts bimolecule transfer integrals based on the structure-priority fusion features, outputting the corresponding transfer integral prediction data. Both the first molecular structure M1 and the second molecular structure M2 are composed of multiple atomic features, each consisting of an atomic element type and three-dimensional atomic coordinates. The bimolecule prior property vector X is composed of the prior properties of the first and second molecules corresponding to the first molecular structure M1 and the second molecular structure M2, specifically including the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecule orbital overlap integral calculated based on a semi-empirical method, and the intermolecular interaction energy. The semi-empirical method includes the GFN2-xTB method.
[0006] The first dataset was constructed by collecting bimolecular information of thin-film structure optoelectronic materials; and the second dataset was constructed by collecting bimolecular information of crystal structure optoelectronic materials.
[0007] The bimolecular transfer integral prediction model is trained based on the first dataset to obtain the corresponding thin-film model parameter set; and the bimolecular transfer integral prediction model is trained based on the second dataset to obtain the corresponding crystalline model parameter set.
[0008] The system receives the material structure type and bimolecular system input by the user; it then sets the model parameters of the bimolecular transfer integral prediction model using the parameter set of the thin film model or the parameter set of the crystal model corresponding to the material structure type to obtain the corresponding current prediction model; based on the bimolecular system, it performs bimolecular structure extraction and prior property calculation to obtain a set of corresponding first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X; and inputs the current first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X into the current prediction model to obtain the corresponding transfer integral prediction data, which is then fed back to the current user; the material structure type includes thin film structure and crystal structure; the material bimolecular system includes multiple system atoms; the atomic properties of each system atom include the atomic element type, the atomic three-dimensional coordinates, and the molecular identifier; the molecular identifier includes the first molecular identifier and the second molecular identifier.
[0009] Preferably, the first model input terminal of the bimolecular transfer integral prediction model is used to receive the first molecular structure M1, the second model input terminal is used to receive the second molecular structure M2, the third model input terminal is used to receive the bimolecular prior characteristic vector X, and the model output terminal is used to output the corresponding transfer integral prediction data.
[0010] The bimolecular transfer integral prediction model includes a first structural encoder, a second structural encoder, a first feature fusion module, a fused feature encoder, a feature mapping network, a prior feature encoder, a second feature fusion module, and a prediction output layer.
[0011] The input of the first structural encoder is connected to the input of the first model, and its output is connected to the first input of the first feature fusion module; the input of the second structural encoder is connected to the input of the second model, and its output is connected to the second input of the first feature fusion module; the output of the first feature fusion module is connected to the input of the fused feature encoder; the output of the fused feature encoder is connected to the input of the feature mapping network; the output of the feature mapping network is connected to the first input of the second feature fusion module; the output of the prior feature encoder is connected to the second input of the second feature fusion module; the output of the second feature fusion module is connected to the input of the prediction output layer; and the output of the prediction output layer is connected to the model output.
[0012] The first structure encoder is implemented based on a pre-trained Uni-Mol model; the first structure encoder is used to perform atomic-level high-dimensional feature encoding on the first molecular structure M1 using the pre-trained Uni-Mol model to obtain the corresponding feature tensor E1, which is then sent to the first feature fusion module; the feature tensor E1 has a shape of N1×C. A N1 is the total number of atoms in the first molecular structure M1, C A The atomic feature dimension is preset; the feature tensor E1 consists of N1 vectors of length C. A Composed of atomic feature vectors;
[0013] The second structure encoder is implemented based on a pre-trained Uni-Mol model; the second structure encoder is used to perform atomic-level high-dimensional feature encoding on the second molecular structure M2 using the pre-trained Uni-Mol model to obtain the corresponding feature tensor E2, which is then sent to the first feature fusion module; the feature tensor E2 has a shape of N2×C. A N2 is the total number of atoms in the second molecular structure M2; the feature tensor E2 is composed of N2 atomic feature vectors;
[0014] The first feature fusion module is used to merge the feature tensor E1 and the feature tensor E2 to obtain the corresponding merged feature tensor H1 and send it to the fusion feature encoder; the shape of the merged feature tensor H1 is (N1+N2)×C A The merged feature tensor H1 is composed of N1+N2 atomic feature vectors.
[0015] The fusion feature encoder is implemented based on the encoder model of the Transformer architecture; the fusion feature encoder is used to treat each atomic feature vector of the merged feature tensor H1 as a corresponding token embedding encoding vector t. i 1 ≤ index i ≤ (N1 + N2); and in ascending order of index i, the N1 + N2 token embedding encoding vectors t are processed. i Sort the data to obtain the corresponding first vector sequence {t} i}; and based on the position embedding encoding rules of the Transformer architecture, according to each of the token embedding encoding vectors t i Position embedding encoding is performed on index i to obtain a vector of length C. A Position embedding encoding vector p i ; and based on each of the token embedding encoding vectors t i and its corresponding position embedding encoding vector p i A new token embedding encoding vector is calculated. And in ascending order of index i, the N1+N2 token embedding encoding vectors are processed. Sort to obtain the corresponding second vector sequence and the second vector sequence The encoder model of the Transformer architecture is input for encoding processing to obtain the corresponding encoded feature vector sequence H2, which is then sent to the feature mapping network; the encoded feature vector sequence H2 consists of N1+N2 encoded feature vectors. The encoded feature vector is formed by sequential sorting; With the token embedding encoding vector One-to-one correspondence;
[0016] The feature mapping network is used to map each of the encoded feature vectors in the encoded feature vector sequence H2 using a built-in MLP model. Do a job from C A 3D feature space to C B Vector mapping in the 3D feature space yields a vector of length C. B Mapping feature vector And consisting of N1+N2 mapping feature vectors Form a shape of (N1+N2)×CB The mapping feature tensor H proj And based on attention pooling, the mapping feature tensor H is processed. proj The attention weights of N1+N2 feature data points in each feature dimension are calculated, and a weighted sum is calculated on the N1+N2 feature data points of the current feature dimension based on the N1+N2 attention weights of each feature dimension. The weighted sum calculation result corresponding to each feature dimension is used as the feature pooling data corresponding to the current feature dimension; and the obtained C... B Each of the aforementioned feature pooling data forms a corresponding feature vector H3, which is then sent to the second feature fusion module.
[0017] Among them, C B These are preset molecular feature dimensions;
[0018] The mapping feature vector The vector mapping process is as follows:
[0019]
[0020] σ1 and σ2 are two activation functions corresponding to the feature mapping network, where σ1 is either a ReLU or a Sigmoid activation function, and σ2 is either a ReLU or a Sigmoid activation function; W1 and W2 are two weight matrix parameters corresponding to the feature mapping network, and b1 and b2 are two offset vector parameters corresponding to the feature mapping network; the encoded feature vector The shape is C A ×1, the shape of the weight matrix parameter W1 is C B ×C A The shape of the offset vector parameter b1 is C. B ×1, the shape of the weight matrix parameter W2 is C B ×C B The shape of the offset vector parameter b2 is C. B ×1, the mapping feature vector The shape is C B ×1;
[0021] The prior feature encoder is used to normalize the prior feature data of the bimolecular prior feature vector X to obtain the corresponding normalized vector V; and based on the built-in MLP model, it performs a step from C to V on the normalized vector V. X 3D feature space to C B Vector mapping in the 3D feature space yields a vector of length C. B The feature vector H4 is sent to the second feature fusion module;
[0022] Wherein, the vector length of the bimolecular prior property vector X is C. X The length of the feature vector H4 is C. B C X The preset length of the bimolecular prior property vector;
[0023] The vector mapping process of the feature vector H4 is as follows:
[0024] H4 = σ4(W4σ3(W3V+b3)+b4);
[0025] σ3 and σ4 are two activation functions of the prior feature encoder, where σ3 is either a ReLU or Sigmoid activation function, and σ4 is either a ReLU or Sigmoid activation function; W3 and W4 are two weight matrix parameters corresponding to the prior feature encoder, and b3 and b4 are two offset vector parameters corresponding to the prior feature encoder; the normalized vector V has a shape of C. X ×1, the shape of the weight matrix parameter W3 is C B ×C X The shape of the offset vector parameter b3 is C. B ×1, the shape of the weight matrix parameter W4 is C B ×C B The shape of the offset vector parameter b4 is C. B ×1, the shape of the feature vector H4 is C B ×1;
[0026] The second feature fusion module is used to concatenate the feature vectors H3 and H4 to obtain the corresponding concatenated vector H5, which is then sent to the prediction output layer; the length of the concatenated vector H5 is 2C. B ;
[0027] The prediction output layer is used to perform bimolecular transfer integral prediction based on the splicing vector H5 using the built-in fully connected layer and output the corresponding transfer integral prediction data.
[0028] The prediction process for the transfer integral prediction data is as follows:
[0029] Y = W5H5 + b5;
[0030] Y represents the transfer integral prediction data, W5 and b5 are the weight matrix parameters and offset vector parameters corresponding to the prediction output layer, respectively; the shape of the concatenation vector H5 is 2C. B ×1, the shape of the weight matrix parameter W5 is 1×2C BThe offset vector parameter b5 is a scalar with a shape of 1×1, and the transfer integral prediction data Y is a scalar with a shape of 1×1.
[0031] Preferably, the first dataset includes multiple first data records; each first data record corresponds to a bimolecular structure of a class of thin-film optoelectronic materials; the first data record includes a first training structure, a second training structure, a first training feature vector, and first label data; the data structures of the first training structure and the second training structure are consistent with the corresponding first molecular structure M1 and second molecular structure M2; the data structure of the first training feature vector is consistent with the bimolecular prior feature vector X; the first label data is the bimolecular transfer integral of the current bimolecular structure.
[0032] The second dataset includes multiple second data records; each second data record corresponds to a bimolecular structure of a class of crystal structure optoelectronic materials; the second data record includes a third training structure, a fourth training structure, a second training feature vector, and second label data; the data structures of the third training structure and the fourth training structure are consistent with the corresponding first molecular structure M1 and second molecular structure M2; the data structure of the second training feature vector is consistent with the bimolecular prior feature vector X; the second label data is the bimolecular transfer integral of the current bimolecular structure.
[0033] Preferably, the step of constructing the first dataset by collecting bimolecular information of thin-film optoelectronic materials specifically includes:
[0034] Step 41: Collect big data on various molecular pair structures of thin-film optoelectronic materials through a preset first big data channel to obtain the corresponding first molecular pair dataset;
[0035] The first big data channel includes publicly available molecular information databases for optoelectronic materials, publicly available molecular information databases for thin-film structure optoelectronic materials, publicly available experimental information databases for thin-film structure optoelectronic materials, and publicly available technical literature on thin-film structure optoelectronic materials.
[0036] The first molecular pair dataset includes multiple first molecular pairs; the two molecules corresponding to the first molecular pair are denoted as molecule A1 and molecule A2; the first molecular pair includes molecular sequence S1 and molecular sequence S2; the molecular sequence S1 and the molecular sequence S2 are respectively the SMILES sequences of the corresponding molecules A1 and A2;
[0037] Step 42: Take each of the first molecule pairs as the corresponding current molecule pair;
[0038] Step 43: Based on a preset cheminformatics tool, a three-dimensional molecular conformation is created according to the molecular sequence S1 and the molecular sequence S2 of the current molecular pair to obtain the corresponding first molecular conformation and second molecular conformation.
[0039] The cheminformatics tools mentioned include Open Babel software and RDKit software;
[0040] Step 44: Based on a preset molecular dynamics simulation tool, optimize the stable conformations of the first molecular conformation and the second molecular conformation to obtain the corresponding first optimized conformation and the second optimized conformation; and based on the molecular dynamics simulation tool, simulate the thin film structure bimolecular system composed of the first optimized conformation and the second optimized conformation to obtain the corresponding first bimolecular system.
[0041] The molecular dynamics simulation tools include Gaussian software and GROMACS software;
[0042] Step 45: Extract the element type and three-dimensional coordinates of each atom of molecule A1 in the first bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding first training structure from all the atomic features corresponding to molecule A1.
[0043] Step 46: Extract the element type and three-dimensional coordinates of each atom of molecule A2 in the first bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding second training structure from all the atomic features corresponding to molecule A2.
[0044] Step 47: Based on the cheminformatics tool, identify the total number of aromatic rings of molecule A1 in the first bimolecular system to obtain a corresponding total number of aromatic rings of the first molecule; and based on a preset quantum chemical calculation tool, calculate the molecular polarity index of molecule A1 in the first bimolecular system to obtain a corresponding first molecular polarity index; and based on the quantum chemical calculation tool, calculate the molecular dipole moment of molecule A2 in the first bimolecular system to obtain a corresponding second molecular dipole moment; and based on the quantum chemical calculation tool, calculate the molecular polarity index of molecule A2 in the first bimolecular system to obtain a corresponding second molecular polarity index. Based on the quantum chemical calculation tool, the orbital overlap integral between molecule A1 and molecule A2 in the first bimolecular system is calculated using the semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; based on the quantum chemical calculation tool, the intermolecular interaction energy between molecule A1 and molecule A2 in the first bimolecular system is calculated using the semi-empirical method; and a corresponding first training characteristic vector is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current first bimolecular system;
[0045] The quantum chemical calculation tools include Multiwfn software, Gaussian software, ORCA software, GAMESS software, NWChem software, SHARC software, GPAW software, and Newton-X software.
[0046] Step 48: Based on the quantum chemical calculation tool, calculate the bimolecular transfer integral of molecule A1 and molecule A2 in the first bimolecular system according to the preset calculation method, and use the calculation result as a corresponding first tag data;
[0047] The preset calculation methods include the DFT method and the TD-DFT method;
[0048] Step 49: The first training structure, the second training structure, the first training feature vector, and the first label data corresponding to the current molecular pair are combined to form a corresponding first data record; and all the obtained first data records are combined to form the corresponding first dataset.
[0049] Preferably, the step of constructing the second dataset by collecting bimolecular information of crystal structure optoelectronic materials specifically includes:
[0050] Step 51: Collect big data on various molecular pair structures of crystal structure optoelectronic materials through a preset second big data channel to obtain the corresponding second molecular pair dataset;
[0051] The second big data channel includes publicly available molecular information databases for optoelectronic materials, publicly available molecular information databases for crystal structure optoelectronic materials, publicly available experimental information databases for crystal structure optoelectronic materials, and publicly available technical literature on crystal structure optoelectronic materials.
[0052] The second molecular pair dataset includes multiple second molecular pairs; the two molecules corresponding to the second molecular pairs are denoted as molecule A3 and molecule A4; the second molecular pair includes molecular sequence S3 and molecular sequence S4; the molecular sequence S3 and the molecular sequence S4 are the SM-ILES sequences of the corresponding molecules A3 and A4, respectively;
[0053] Step 52: Take each of the second molecular pairs as the corresponding current molecular pairs;
[0054] Step 53: Based on the cheminformatics tool, a three-dimensional molecular conformation is created according to the molecular sequence S3 and the molecular sequence S4 of the current molecular pair to obtain the corresponding third molecular conformation and fourth molecular conformation.
[0055] Step 54: Based on the molecular dynamics simulation tool, optimize the stable conformations of the third and fourth molecular conformations to obtain the corresponding third optimized conformation and fourth optimized conformation; and based on the molecular dynamics simulation tool, simulate the crystal structure bimolecular system composed of the third optimized conformation and the fourth optimized conformation to obtain the corresponding second bimolecular system.
[0056] Step 55: Extract the element type and three-dimensional coordinates of each atom of molecule A3 in the second bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding third training structure from all the atomic features corresponding to molecule A3.
[0057] Step 56: Extract the element type and three-dimensional coordinates of each atom of molecule A4 in the second bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding fourth training structure from all the atomic features corresponding to molecule A4.
[0058] Step 57: Based on the cheminformatics tool, identify the total number of aromatic rings of molecule A3 in the second bimolecular system to obtain a corresponding total number of aromatic rings of the first molecule; and based on the quantum chemical calculation tool, calculate the molecular polarity index of molecule A3 in the second bimolecular system to obtain a corresponding first molecular polarity index; and based on the quantum chemical calculation tool, calculate the molecular dipole moment of molecule A4 in the second bimolecular system to obtain a corresponding second molecular dipole moment; and based on the quantum chemical calculation tool, calculate the molecular polarity index of molecule A4 in the second bimolecular system to obtain a corresponding second molecular polarity index. Based on the quantum chemical calculation tool, the orbital overlap integral between molecule A3 and molecule A4 in the second bimolecular system is calculated using the semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; based on the quantum chemical calculation tool, the intermolecular interaction energy between molecule A3 and molecule A4 in the second bimolecular system is calculated using the semi-empirical method; and a corresponding second training characteristic vector is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current second bimolecular system;
[0059] Step 58: Based on the quantum chemical calculation tool, calculate the bimolecular transfer integral of molecule A3 and molecule A4 in the second bimolecular system according to the preset calculation method, and use the calculation result as a corresponding second tag data;
[0060] Step 59: The third training structure, the fourth training structure, the second training feature vector, and the second label data corresponding to the current molecular pair are combined to form a corresponding second data record; and all the obtained second data records are combined to form the corresponding second dataset.
[0061] Preferably, the step of training the bimolecular transfer integral prediction model based on the first dataset to obtain the corresponding thin-film model parameter set specifically includes:
[0062] Step 61: Initialize the model parameters of the bimolecular transfer integral prediction model based on the preset initialization model parameter set;
[0063] Step 62: Based on a preset first segmentation ratio, the first dataset is randomly divided into two sub-datasets, denoted as the first training set and the first evaluation set.
[0064] Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio;
[0065] Step 63: Take each of the first data records in the first training set as the corresponding current training record; and take the first training structure, the second training structure, and the first training feature vector of the current training record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding first prediction data; and take the first prediction data corresponding to the current training record and the first label data as a corresponding first prediction-label pair.
[0066] Step 64: Input all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value;
[0067] The first model loss function is implemented based on the L1 loss function, the Smooth L1 loss function, or the L2 loss function.
[0068] Step 65: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 66; if it does not, modulate the model parameters of the bimolecular transfer integral prediction model in one round based on the preset first model optimizer in the direction of minimizing the first model loss function, and return to step 63 to continue training when this round of modulation is combined.
[0069] The first model optimizer includes the Adam optimizer and the SGD optimizer.
[0070] Step 66: Take each of the first data records in the first evaluation set as the corresponding current evaluation record; and take the first training structure, the second training structure, and the first training feature vector of the current evaluation record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding second prediction data; and take the second prediction data corresponding to the current evaluation record and the first label data as a corresponding second prediction-label pair; and take all the obtained second prediction-label pairs into the preset first model evaluation function to calculate the corresponding first evaluation value;
[0071] The first model evaluation function is implemented based on the MAE function, MSE function, or RMSE function.
[0072] Step 67: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 62 to continue training; if yes, extract the current model parameters of the bimolecular transfer integral prediction model as the corresponding thin film model parameter set and save it.
[0073] Preferably, the step of training the bimolecular transfer integral prediction model based on the second dataset to obtain the corresponding crystal version model parameter set specifically includes:
[0074] Step 71: Initialize the model parameters of the bimolecular transfer integral prediction model based on the preset initialization model parameter set;
[0075] Step 72: Based on the preset second segmentation ratio, the second dataset is randomly divided into two subsets, which are denoted as the corresponding second training set and second evaluation set;
[0076] The second training set and the second evaluation set are both composed of multiple second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set satisfies the second segmentation ratio.
[0077] Step 73: Take each of the second data records in the second training set as the corresponding current training record; and take the third training structure, the fourth training structure, and the second training feature vector of the current training record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding third prediction data; and take the third prediction data corresponding to the current training record and the second label data as a corresponding third prediction-label pair.
[0078] Step 74: Input all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value;
[0079] The second model loss function is implemented based on the L1 loss function, the Smooth L1 loss function, or the L2 loss function.
[0080] Step 75: Identify whether the second loss value meets the preset second loss value range; if it does, proceed to step 76; if it does not, based on the preset second model optimizer, perform a round of modulation on the model parameters of the bimolecular transfer integral prediction model in the direction of minimizing the second model loss function, and return to step 73 to continue training when this round of modulation is combined.
[0081] The second model optimizer includes the Adam optimizer and the SGD optimizer;
[0082] Step 76: Take each of the second data records in the second evaluation set as the corresponding current evaluation record; and take the third training structure, the fourth training structure, and the second training feature vector of the current evaluation record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding fourth prediction data; and take the fourth prediction data corresponding to the current evaluation record and the second label data as a corresponding fourth prediction-label pair; and take all the obtained fourth prediction-label pairs into the preset second model evaluation function to calculate the corresponding second evaluation value.
[0083] The second model evaluation function is implemented based on the MAE function, MSE function, or RMSE function.
[0084] Step 77: Identify whether the second evaluation value meets the preset range of the second evaluation value; if not, return to step 72 to continue training; if so, extract the current model parameters of the bimolecular transfer integral prediction model as the corresponding crystal version model parameter set and save it.
[0085] Preferably, the step of setting the model parameters of the bimolecular transfer integral prediction model using the thin film model parameter set or the crystal model parameter set corresponding to the material structure type to obtain the corresponding current prediction model specifically includes:
[0086] The material structure type is identified; if the material structure type is a thin film structure, the model parameters of the bimolecular transfer integral prediction model are set based on the thin film model parameter set; if the material structure type is a crystal structure, the model parameters of the bimolecular transfer integral prediction model are set based on the crystal model parameter set; and the bimolecular transfer integral prediction model whose parameters have been set is used as the corresponding current prediction model.
[0087] Preferably, the step of extracting the bimolecular structure and calculating prior properties based on the bimolecular system of the material to obtain a set of corresponding first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X specifically includes:
[0088] Step 91: The atomic element type and the three-dimensional coordinates of the atoms of the system atoms that are identified as the first molecule identifier in the material bimolecular system are combined to form a corresponding atomic feature, and all the atomic features corresponding to the first molecule identifier are combined to form a corresponding first molecular structure M1;
[0089] Step 92: The atomic element type and the three-dimensional coordinates of the atoms of the system atoms that are identified as the second molecule identifier in the material bimolecular system are combined to form a corresponding atomic feature, and all the atomic features corresponding to the second molecule identifier are combined to form a corresponding second molecular structure M2;
[0090] Step 93: Based on the cheminformatics tool, identify the total number of aromatic rings of the first molecule corresponding to the first molecule identifier in the material bimolecular system to obtain a corresponding total number of aromatic rings of the first molecule; and based on the quantum chemical calculation tool, calculate the molecular polarity index of the first molecule in the material bimolecular system to obtain a corresponding first molecule polarity index; and based on the quantum chemical calculation tool, calculate the molecular dipole moment of the second molecule corresponding to the second molecule identifier in the material bimolecular system to obtain a corresponding second molecule dipole moment; and based on the quantum chemical calculation tool, calculate the molecular polarity index of the second molecule in the material bimolecular system to obtain a... The corresponding second molecule polarity index; and based on the quantum chemical calculation tool, the orbital overlap integral between the first molecule and the second molecule in the material bimolecular system is calculated using the semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; and based on the quantum chemical calculation tool, the intermolecular interaction energy between the first molecule and the second molecule in the material bimolecular system is calculated using the semi-empirical method; and a corresponding bimolecular a priori characteristic vector X is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current material bimolecular system;
[0091] Step 94, and output the first molecular structure M1, the second molecular structure M2, and the bimolecular prior property vector X obtained in this calculation as the result of this calculation.
[0092] A second aspect of the present invention provides an apparatus for implementing the processing method of the bimolecular transfer integral prediction model described in the first aspect above, the apparatus comprising: a model building module, a dataset building module, a model training module, and a model application module;
[0093] The model building module is used to construct a deep learning model that fuses the structure and prior properties of bimolecules and predicts bimolecule transfer integrals based on the fused features, denoted as the corresponding bimolecule transfer integral prediction model. The bimolecule transfer integral prediction model fuses the structural features of the first molecular structure M1 and the second molecular structure M2 input to the model, fuses the bimolecule fusion features with the prior property vector X input to the model, and predicts bimolecule transfer integrals based on the structure-priority fusion features, outputting the corresponding transfer integral prediction data. Both the first molecular structure M1 and the second molecular structure M2 are composed of multiple atomic features, each of which consists of an atomic element type and three-dimensional atomic coordinates. The bimolecule prior property vector X is composed of the prior properties of the first and second molecules corresponding to the first molecular structure M1 and the second molecular structure M2, specifically including the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecule orbital overlap integral calculated based on a semi-empirical method, and the intermolecular interaction energy. The semi-empirical method includes the GFN2-xTB method.
[0094] The dataset construction module is used to construct a first dataset by collecting bimolecular information of thin-film structure optoelectronic materials, and to construct a second dataset by collecting bimolecular information of crystal structure optoelectronic materials.
[0095] The model training module is used to train the bimolecular transfer integral prediction model based on the first dataset to obtain the corresponding thin-film model parameter set; and to train the bimolecular transfer integral prediction model based on the second dataset to obtain the corresponding crystal model parameter set;
[0096] The model application module receives the material structure type and material bimolecular system input by the user; and sets the model parameters of the bimolecular transfer integral prediction model using the thin film model parameter set or the crystal model parameter set corresponding to the material structure type to obtain the corresponding current prediction model; and performs bimolecular structure extraction and prior property calculation processing based on the material bimolecular system to obtain a set of corresponding first molecular structure M1, second molecular structure M2 and bimolecular prior property vector X; and inputs the current first molecular structure M1, second molecular structure M2 and bimolecular prior property vector X into the current prediction model to obtain the corresponding transfer integral prediction data and feeds it back to the current user; the material structure type includes thin film structure and crystal structure; the material bimolecular system includes multiple system atoms; the atomic properties of each system atom include the atomic element type, the atomic three-dimensional coordinates and the molecular identifier; the molecular identifier includes the first molecular identifier and the second molecular identifier.
[0097] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0098] The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above;
[0099] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0100] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.
[0101] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for processing a bimolecular transfer integral prediction model. As described above, this invention uses six types of bimolecular properties that can be quickly obtained through simple calculations as prior properties: the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral and the intermolecular interaction energy calculated based on semi-empirical methods (such as the GFN2-xTB method), and uses the bimolecular transfer integral as the prediction target. A deep learning model (bimolecular transfer integral prediction model) is thus constructed to predict the bimolecular transfer integral based on two three-dimensional molecular structures (first molecular structure M1, second molecular structure M2) and six types of prior properties (bimolecular prior property vector X). A corresponding first / second dataset is constructed by collecting bimolecular information from thin-film / crystal structure optoelectronic materials. The bimolecular transfer integral prediction model is then trained on the first / second datasets to obtain two sets of model parameters (thin-film model parameter set, crystal model parameter set) corresponding to the two types of structures. After model training, the two sets of model parameters from the bimolecular transfer integral prediction model are used to predict the bimolecular transfer integral of any bimolecular system in the field of thin-film / crystal structure optoelectronic materials. The embodiments of the present invention shorten the analysis time, improve the analysis efficiency and prediction accuracy; based on the embodiments of the present invention, high-throughput screening of optoelectronic material molecules with thin film / crystal structure can effectively save computing resources and optimize screening efficiency, providing effective assistance for accelerating the research and development iteration of new materials. Attached Figure Description
[0102] Figure 1 This is a schematic diagram of a processing method for a bimolecular transfer integral prediction model provided in Embodiment 1 of the present invention;
[0103] Figure 2 This is a block diagram of the bimolecular transfer integral prediction model provided in Embodiment 1 of the present invention;
[0104] Figure 3 This is a module structure diagram of a processing device for a bimolecular transfer integral prediction model provided in Embodiment 2 of the present invention;
[0105] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0106] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0107] Embodiment 1 of the present invention provides a method for processing a bimolecular transfer integral prediction model; such as Figure 1 The schematic diagram shows a processing method for a bimolecular transfer integral prediction model provided in Embodiment 1 of the present invention. The method mainly includes the following steps:
[0108] Step 1: Construct a deep learning model that fuses the structure and prior properties of bimolecules and predicts bimolecule transfer integrals based on the fused features. This model is denoted as the corresponding bimolecule transfer integral prediction model.
[0109] Here, the bimolecular transfer integral prediction model of this embodiment of the invention is used to perform structural feature fusion on the first molecular structure M1 and the second molecular structure M2 input to the model, and to perform feature fusion based on the bimolecular fusion features and the bimolecular prior property vector X input to the model, and to perform bimolecular transfer integral prediction based on the structure-priority fusion features and output the corresponding transfer integral prediction data. The first molecular structure M1 and the second molecular structure M2 are both composed of multiple atomic features, each atomic feature consisting of atomic element type and atomic three-dimensional coordinates. The bimolecular prior property vector X is composed of the prior properties of the first and second molecules corresponding to the first molecular structures M1 and M2, specifically including: the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral calculated based on a semi-empirical method, and the intermolecular interaction energy. Here, the semi-empirical method of this embodiment of the invention includes the GFN2-xTB method.
[0110] It should be noted that the prediction target of this embodiment of the invention is the bimolecular transfer integral. Conventionally, calculations using quantum chemical computing tools based on DFT or TD-DFT methods take hours or days. This embodiment of the invention, through an end-to-end bimolecular transfer integral prediction model, can significantly reduce the processing time of this calculation / analysis process, thereby reducing the computational / analysis difficulty and improving computational / analysis efficiency.
[0111] It should also be noted that the bimolecular prior characteristic vector X is used in this embodiment of the invention to improve the prediction accuracy of the model. The six types of prior characteristics of this bimolecular prior characteristic vector X can be quickly obtained using computational tools: 1) Total number of aromatic rings in the first molecule: This characteristic parameter can be obtained by inputting the bimolecular system into a cheminformatics tool and counting the total number of aromatic rings in the first molecule. Normally, this statistical process takes only seconds; 2) Polarity index of the first molecule, dipole moment of the second molecule, and polarity index of the second molecule: These three characteristic parameters can be obtained by inputting the bimolecular system into a quantum chemistry computational tool. Normally, the calculation process for these three characteristic parameters takes only minutes or even faster; 3) Bimolecular orbital overlap integral: This characteristic parameter can be obtained by inputting the bimolecular system into a quantum chemistry computational tool and calculating the orbital overlap integral between molecules based on the GFN2-xTB method. Normally, this calculation process takes only seconds; 4) Intermolecular interaction energy: This characteristic parameter can be obtained by inputting the bimolecular system into a quantum chemistry computational tool and calculating the intermolecular interaction energy based on the GFN2-xTB method. Normally, this calculation process takes only seconds or minutes. Therefore, although the bimolecular prior property vector X must be calculated before each use of the bimolecular transfer integral prediction model, the overall processing time (X preparation time + model processing time) is still much shorter than that of conventional quantum chemical methods. In other words, using the bimolecular prior property vector X in this embodiment of the invention can help improve the model's prediction accuracy without significantly impacting the reduction of computation / analysis time or the improvement of computation / analysis efficiency.
[0112] like Figure 2 As shown in the module structure diagram of the bimolecular transfer integral prediction model provided in Embodiment 1 of the present invention, the first model input terminal of the bimolecular transfer integral prediction model of the present invention is used to receive the first molecular structure M1, the second model input terminal is used to receive the second molecular structure M2, the third model input terminal is used to receive the bimolecular prior characteristic vector X, and the model output terminal is used to output the corresponding transfer integral prediction data.
[0113] like Figure 2As shown, the model components of the bimolecular transfer integral prediction model include: a first structure encoder, a second structure encoder, a first feature fusion module, a fusion feature encoder, a feature mapping network, a prior feature encoder, a second feature fusion module, and a prediction output layer.
[0114] like Figure 2 As shown, the connection relationships of the model components in the bimolecular transfer integral prediction model are as follows: the input end of the first structural encoder is connected to the input end of the first model, and its output end is connected to the first input end of the first feature fusion module; the input end of the second structural encoder is connected to the input end of the second model, and its output end is connected to the second input end of the first feature fusion module; the output end of the first feature fusion module is connected to the input end of the fused feature encoder; the output end of the fused feature encoder is connected to the input end of the feature mapping network; the output end of the feature mapping network is connected to the first input end of the second feature fusion module; the output end of the prior feature encoder is connected to the second input end of the second feature fusion module; the output end of the second feature fusion module is connected to the input end of the prediction output layer; and the output end of the prediction output layer is connected to the model output end.
[0115] The model component functions of the bimolecular transfer integral prediction model are shown below.
[0116] 1) First structure encoder:
[0117] The first structure encoder is implemented based on a pre-trained Uni-Mol model. The first structure encoder is used to perform atomic-level high-dimensional feature encoding on the first molecular structure M1 using the pre-trained Uni-Mol model to obtain the corresponding feature tensor E1, which is then sent to the first feature fusion module. The feature tensor E1 has a shape of N1×C. A N1 is the total number of atoms in the first molecular structure M1, C A The atomic feature dimension is preset; the feature tensor E1 consists of N1 vectors of length C. A It consists of atomic feature vectors.
[0118] Here, the Uni-Mol model for the first structure encoder has been pre-trained. Detailed model structure and pre-training methods of the Uni-Mol model can be found in the publicly available technical document "Uni-Mol: A Universal 3D Dolecul arRepresentation Learning Framework," and will not be elaborated further here. The reason this embodiment uses the Uni-Mol model as the encoder for the first molecular structure M1 is that this model can refine the precision of encoded features to the atomic level. This provides richer original feature information for the molecular-level features generated by the subsequent feature mapping network, thereby helping to further improve the final prediction accuracy.
[0119] 2) Second structure encoder:
[0120] The second structure encoder is implemented based on a pre-trained Uni-Mol model. The second structure encoder is used to perform atomic-level high-dimensional feature encoding on the second molecular structure M2 using the pre-trained Uni-Mol model, obtaining the corresponding feature tensor E2, which is then sent to the first feature fusion module. The feature tensor E2 has a shape of N²×C. A N2 represents the total number of atoms in the second molecular structure M2; the characteristic tensor E2 consists of N2 atomic characteristic vectors.
[0121] Here, the Uni-Mol model of the second structure encoder has also been pre-trained. Similar to the first structure encoder, the reason why this embodiment of the invention uses the Uni-Mol model as the encoder for the second molecular structure M2 is because this model can refine the precision of the encoded features to the atomic level, providing richer original feature information for the molecular-level features generated by the subsequent feature mapping network, which helps to improve prediction accuracy. It should be noted that although both the first and second structure encoders have a built-in pre-trained Uni-Mol model, these two Uni-Mol models are independent of each other and do not share parameters.
[0122] 3) First Feature Fusion Module
[0123] The first feature fusion module is used to merge feature tensors E1 and E2 to obtain the corresponding merged feature tensor H1, which is then sent to the fusion feature encoder.
[0124] Here, the shape of the merged feature tensor H1 is (N1+N2)×C A The merged feature tensor H1 consists of N1+N2 atomic feature vectors.
[0125] 4) Fusion Feature Encoder:
[0126] The fusion feature encoder is implemented based on the encoder model of the Transformer architecture. The fusion feature encoder is used to embed the individual atomic feature vectors of the merged feature tensor H1 into a corresponding token-based encoding vector t. i 1 ≤ index i ≤ (N1 + N2); and in ascending order of index i, embed the encoding vectors t of the N1 + N2 tokens. i Sort the data to obtain the corresponding first vector sequence {t} i}; and based on the positional embedding encoding rules of the Transformer architecture, according to the embedding encoding vector t of each token i Position embedding encoding is performed on index i to obtain a vector of length C. A Position embedding encoding vector pi ; and based on each token embedding encoding vector t i and its corresponding position embedding encoding vector p i A new token embedding encoding vector is calculated. And in ascending order of index i, embed the encoding vectors of the N1+N2 tokens. Sort to obtain the corresponding second vector sequence and the second vector sequence The encoder model of the Transformer architecture is input for encoding processing to obtain the corresponding encoded feature vector sequence H2, which is then sent to the feature mapping network.
[0127] Among them, the encoded feature vector sequence H2 consists of N1+N2 encoded feature vectors. Arranged sequentially; encoded feature vector h 2 i With token embedding encoding vector One-to-one correspondence.
[0128] 5) Feature mapping network:
[0129] Feature mapping networks are used to encode individual feature vectors of a sequence of encoded feature vectors H2 using a built-in MLP model. Do a job from C A 3D feature space to C B Vector mapping in the 3D feature space yields a vector of length C. B Mapping feature vector And consists of N1+N2 mapped feature vectors Form a shape of (N1+N2)×C B The mapping feature tensor H proj And based on the attention pooling method, the mapping feature tensor H is processed. proj The attention weights of N1+N2 feature data points in each feature dimension are calculated, and a weighted sum is calculated on the N1+N2 feature data points of the current feature dimension based on the N1+N2 attention weights of each feature dimension. The weighted sum calculation result corresponding to each feature dimension is used as the feature pooling data corresponding to the current feature dimension; and the obtained C... B Each feature pooling data is used to form a corresponding feature vector H3, which is then sent to the second feature fusion module.
[0130] Here, the mapping feature vector of the embodiment of the present invention The vector mapping process is as follows:
[0131]
[0132] Where σ1 and σ2 are the two activation functions corresponding to the feature mapping network, where σ1 is either a ReLU or Sigmoid activation function, and σ2 is either a ReLU or Sigmoid activation function; W1 and W2 are the two weight matrix parameters corresponding to the feature mapping network, and b1 and b2 are the two offset vector parameters corresponding to the feature mapping network; the encoded feature vector... The shape is C A ×1, the shape of the weight matrix parameter W1 is C B ×C A The shape of the offset vector parameter b1 is C B ×1, the shape of the weight matrix parameter W2 is C B ×C B The shape of the offset vector parameter b2 is C. B ×1, mapping feature vector The shape is C B ×1;C B The preset molecular feature dimensions.
[0133] It should be noted that the feature vector H3 output by the feature mapping network is actually the molecular-level feature of the bimolecular system.
[0134] 6) Prior Feature Encoder:
[0135] The prior feature encoder is used to normalize the prior feature data of the bimolecular prior feature vector X to obtain the corresponding normalized vector V; and based on the built-in MLP model, it performs a step from C to V. X 3D feature space to C B Vector mapping in the 3D feature space yields a vector of length C. B The feature vector H4 is sent to the second feature fusion module.
[0136] Here, the vector length of the bimolecular prior property vector X in this embodiment of the invention is C. X The length of the feature vector H4 is C. B C X The length of the pre-defined bimolecular prior property vector.
[0137] The vector mapping process of feature vector H4 in this embodiment of the invention is as follows:
[0138] H4 = σ4(W4σ3(W3V+b3)+b4);
[0139] Where σ3 and σ4 are the two activation functions corresponding to the prior feature encoder, where σ3 is either a ReLU or Sigmoid activation function, and σ4 is either a ReLU or Sigmoid activation function; W3 and W4 are the two weight matrix parameters corresponding to the prior feature encoder, and b3 and b4 are the two offset vector parameters corresponding to the prior feature encoder; the normalized vector V has a shape of C. X ×1, the shape of the weight matrix parameter W3 is C B ×C X The shape of the offset vector parameter b3 is C. B ×1, the shape of the weight matrix parameter W4 is C B ×C B The shape of the offset vector parameter b4 is C B ×1, the shape of the feature vector H4 is C B ×1.
[0140] 7) Second Feature Fusion Module:
[0141] The second feature fusion module is used to concatenate feature vectors H3 and H4 to obtain the corresponding concatenated vector H5, which is then sent to the prediction output layer.
[0142] Here, the length of the splicing vector H5 in this embodiment of the invention is 2C. B .
[0143] 8) Predicting output layer:
[0144] The prediction output layer is used to perform bimolecular transfer integral prediction based on the splicing vector H5 using the built-in fully connected layer and output the corresponding transfer integral prediction data.
[0145] Here, the prediction process for the transfer integral prediction data in this embodiment of the invention is as follows:
[0146] Y = W5H5 + b5;
[0147] Where Y represents the transfer integral prediction data, W5 and b5 are the weight matrix parameters and offset vector parameters corresponding to the prediction output layer, respectively; the shape of the concatenated vector H5 is 2C. B ×1, the shape of the weight matrix parameter W5 is 1×2C. B The offset vector parameter b5 is a scalar with a shape of 1×1, and the transfer integral prediction data Y is a scalar with a shape of 1×1.
[0148] Step 2: Construct a first dataset by collecting bimolecular information of thin-film structure optoelectronic materials; and construct a second dataset by collecting bimolecular information of crystal structure optoelectronic materials.
[0149] Specifically, this includes: Step 21, constructing a first dataset by collecting bimolecular information of thin-film structure optoelectronic materials;
[0150] The first dataset includes multiple first data records; each first data record corresponds to a bimolecular structure of a class of thin-film optoelectronic materials; each first data record includes a first training structure, a second training structure, a first training feature vector, and first label data; the data structures of the first training structure and the second training structure are consistent with the corresponding first molecular structure M1 and second molecular structure M2; the data structure of the first training feature vector is consistent with the bimolecular prior feature vector X; the first label data is the bimolecular transfer integral of the current bimolecular structure.
[0151] Specifically, it includes: Step 211, collecting big data on various molecular pair structures of thin film structure optoelectronic materials through a preset first big data channel to obtain the corresponding first molecular pair dataset;
[0152] Here, the first big data channel in this embodiment of the invention includes publicly available molecular information databases of optoelectronic materials, publicly available molecular information databases of thin-film structure optoelectronic materials, publicly available experimental information databases of thin-film structure optoelectronic materials, and publicly available technical literature on thin-film structure optoelectronic materials.
[0153] The first molecule pair dataset of this embodiment includes multiple first molecule pairs; the two molecules corresponding to the first molecule pair are denoted as molecule A1 and molecule A2; the first molecule pair includes molecule sequence S1 and molecule sequence S2; molecule sequence S1 and molecule sequence S2 are the SMILES sequences of the corresponding molecules A1 and A2, respectively;
[0154] Step 212: Take each first molecule pair as the corresponding current molecule pair;
[0155] Step 213: Based on the preset cheminformatics tools, a three-dimensional molecular conformation is created according to the molecular sequence S1 and molecular sequence S2 of the current molecular pair to obtain the corresponding first molecular conformation and second molecular conformation.
[0156] Here, the cheminformatics tools in this embodiment of the invention include Open Babel software and RDK it software;
[0157] Step 214: Based on the preset molecular dynamics simulation tool, optimize the stable conformations of the first molecular conformation and the second molecular conformation to obtain the corresponding first optimized conformation and the second optimized conformation; and based on the molecular dynamics simulation tool, simulate the thin film structure bimolecular system composed of the first optimized conformation and the second optimized conformation to obtain the corresponding first bimolecular system.
[0158] Here, the molecular dynamics simulation tools in this embodiment of the invention include Gaussian software and GROMACS software;
[0159] Step 215: Extract the element type and three-dimensional coordinates of each atom of molecule A1 in the first bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding first training structure from all the atomic features corresponding to molecule A1.
[0160] Step 216: Extract the element type and three-dimensional coordinates of each atom of molecule A2 in the first bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding second training structure from all the atomic features corresponding to molecule A2.
[0161] Step 217: Based on cheminformatics tools, identify the total number of aromatic rings of molecule A1 in the first bimolecular system to obtain a corresponding total number of aromatic rings of the first molecule; based on a preset quantum chemical calculation tool, calculate the molecular polarity index of molecule A1 in the first bimolecular system to obtain a corresponding first molecular polarity index; based on a quantum chemical calculation tool, calculate the molecular dipole moment of molecule A2 in the first bimolecular system to obtain a corresponding second molecular dipole moment; and based on a quantum chemical calculation tool, calculate the molecular polarity index of molecule A2 in the first bimolecular system to obtain a corresponding second molecular polarity index. The molecular polarity index is calculated; and based on quantum chemical calculation tools, the orbital overlap integral between molecules A1 and A2 in the first bimolecular system is calculated using a semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; and based on quantum chemical calculation tools, the intermolecular interaction energy between molecules A1 and A2 in the first bimolecular system is calculated using a semi-empirical method; and a corresponding first training characteristic vector is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current first bimolecular system;
[0162] Here, the quantum chemical calculation tools in this embodiment of the invention include Multiwfn software, Gaussian software, ORCA software, GAMESS software, NWChem software, SHARC software, GPAW software, and Newton-X software;
[0163] Step 218: Based on quantum chemical calculation tools, calculate the bimolecular transfer integral of molecule A1 and molecule A2 in the first bimolecular system according to the preset calculation method, and use the calculation result as a corresponding first tag data;
[0164] Here, the preset calculation methods in the embodiments of the present invention include the DFT method and the TD-DFT method;
[0165] Step 219: A first data record is formed by the first training structure, the second training structure, the first training feature vector, and the first label data corresponding to the current molecule pair; and a first dataset is formed by all the obtained first data records.
[0166] Step 22, and construct a second dataset by collecting bimolecular information of crystal structure optoelectronic materials;
[0167] The second dataset includes multiple second data records; each second data record corresponds to a bimolecular structure of a class of crystal structure optoelectronic materials; the second data record includes a third training structure, a fourth training structure, a second training feature vector, and second label data; the data structures of the third and fourth training structures are consistent with the corresponding first molecular structure M1 and second molecular structure M2; the data structure of the second training feature vector is consistent with the bimolecular prior feature vector X; the second label data is the bimolecular transfer integral of the current bimolecular structure.
[0168] Specifically, it includes: Step 221, collecting big data on various molecular pair structures of crystal structure optoelectronic materials through a preset second big data channel to obtain the corresponding second molecular pair dataset;
[0169] Here, the second big data channel in this embodiment of the invention includes publicly available molecular information databases of optoelectronic materials, publicly available molecular information databases of crystal structure optoelectronic materials, publicly available experimental information databases of crystal structure optoelectronic materials, and publicly available technical literature on crystal structure optoelectronic materials.
[0170] Here, the second molecular pair dataset of this invention includes multiple second molecular pairs; the two molecules corresponding to the second molecular pair are denoted as molecule A3 and molecule A4; the second molecular pair includes molecular sequence S3 and molecular sequence S4; molecular sequence S3 and molecular sequence S4 are the SMILES sequences of the corresponding molecules A3 and A4, respectively.
[0171] Step 222: Treat each second molecule pair as the corresponding current molecule pair;
[0172] Step 223: Based on cheminformatics tools, create three-dimensional molecular conformations based on the molecular sequences S3 and S4 of the current molecular pair to obtain the corresponding third and fourth molecular conformations.
[0173] Step 224: Based on molecular dynamics simulation tools, optimize the stable conformations of the third and fourth molecular conformations to obtain the corresponding third optimized conformation and fourth optimized conformation; and based on molecular dynamics simulation tools, simulate the crystal structure bimolecular system composed of the third optimized conformation and the fourth optimized conformation to obtain the corresponding second bimolecular system.
[0174] Step 225: Extract the element type and three-dimensional coordinates of each atom of molecule A3 in the second bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding third training structure from all the atomic features corresponding to molecule A3.
[0175] Step 226: Extract the element type and three-dimensional coordinates of each atom of molecule A4 in the second bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding fourth training structure from all the atomic features corresponding to molecule A4.
[0176] Step 227: Based on cheminformatics tools, the total number of aromatic rings of molecule A3 in the second bimolecular system is identified to obtain a corresponding total number of aromatic rings of the first molecule; based on quantum chemical calculation tools, the molecular polarity index of molecule A3 in the second bimolecular system is calculated to obtain a corresponding first molecule polarity index; based on quantum chemical calculation tools, the molecular dipole moment of molecule A4 in the second bimolecular system is calculated to obtain a corresponding second molecule dipole moment; and based on quantum chemical calculation tools, the molecular polarity index of molecule A4 in the second bimolecular system is calculated to obtain a corresponding second molecule... The polarity index is calculated; and based on quantum chemical calculation tools, the orbital overlap integral between molecules A3 and A4 in the second bimolecular system is calculated using a semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; and based on quantum chemical calculation tools, the intermolecular interaction energy between molecules A3 and A4 in the second bimolecular system is calculated using a semi-empirical method; and a corresponding second training characteristic vector is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current second bimolecular system;
[0177] Step 228: Based on quantum chemical calculation tools, calculate the bimolecular transfer integral of molecules A3 and A4 in the second bimolecular system according to the preset calculation method, and use the calculation result as a corresponding second tag data;
[0178] Step 229: A corresponding second data record is formed by the third training structure, the fourth training structure, the second training feature vector, and the second label data corresponding to the current molecular pair; and the corresponding second dataset is formed by all the obtained second data records.
[0179] Step 3: Train the bimolecular transfer integral prediction model based on the first dataset to obtain the corresponding thin film model parameter set; and train the bimolecular transfer integral prediction model based on the second dataset to obtain the corresponding crystal model parameter set;
[0180] Specifically, this includes: Step 31, training the bimolecular transfer integral prediction model based on the first dataset to obtain the corresponding thin-film model parameter set;
[0181] Specifically, this includes: Step 311, initializing the model parameters of the bimolecular transfer integral prediction model based on a preset initialization model parameter set;
[0182] Here, the initialization model parameter set in this embodiment of the invention can be a set of all-zero parameters or a set of random noise parameters that satisfy the Gaussian distribution characteristics;
[0183] Step 312: Based on the preset first segmentation ratio, the first dataset is randomly divided into two subsets, denoted as the first training set and the first evaluation set.
[0184] Here, the first segmentation ratio in this embodiment of the invention is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio;
[0185] Step 313: Take each first data record of the first training set as the corresponding current training record; and take the first training structure, second training structure, and first training feature vector of the current training record as the current first molecular structure M1, second molecular structure M2, and bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained from this prediction as the corresponding first prediction data; and take the first prediction data corresponding to the current training record and the first label data as a corresponding first prediction-label pair.
[0186] Step 314: Input all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value;
[0187] Here, the first model loss function in this embodiment of the invention is implemented based on the L1 loss function, the Smooth L1 loss function, or the L2 loss function;
[0188] Step 315: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 316; if it does not, based on the preset first model optimizer, perform a round of modulation on the model parameters of the bimolecular transfer integral prediction model in the direction of minimizing the first model loss function, and return to step 313 to continue training when this round of modulation is combined.
[0189] Here, the first loss value range in this embodiment of the invention is a pre-set numerical range; the first model optimizer includes the Adam optimizer and the SGD optimizer;
[0190] Step 316: Take each first data record of the first evaluation set as the corresponding current evaluation record; and take the first training structure, second training structure, and first training feature vector of the current evaluation record as the current first molecular structure M1, second molecular structure M2, and bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding second prediction data; and take the second prediction data corresponding to the current evaluation record and the first label data as a corresponding second prediction-label pair; and take all the obtained second prediction-label pairs into the preset first model evaluation function to calculate the corresponding first evaluation value.
[0191] Here, the first model evaluation function in this embodiment of the invention is implemented based on the MAE function, MSE function, or RMSE function;
[0192] Step 317: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 312 to continue training; if it does, extract the current model parameters of the bimolecular transfer integral prediction model as the corresponding thin film version model parameter set and save it.
[0193] Here, the first evaluation value range of this embodiment of the invention is a preset numerical range;
[0194] Step 32, and train the bimolecular transfer integral prediction model based on the second dataset to obtain the corresponding crystal model parameter set;
[0195] Specifically, this includes: Step 321, initializing the model parameters of the bimolecular transfer integral prediction model based on a preset initialization model parameter set;
[0196] Step 322: Based on the preset second segmentation ratio, the second dataset is randomly divided into two subsets, which are denoted as the corresponding second training set and second evaluation set;
[0197] Here, the second segmentation ratio in this embodiment of the invention is a preset ratio parameter, such as 8:2; both the second training set and the second evaluation set consist of multiple second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set satisfies the second segmentation ratio;
[0198] Step 323: Take each second data record of the second training set as the corresponding current training record; and take the third training structure, fourth training structure, and second training feature vector of the current training record as the current first molecular structure M1, second molecular structure M2, and bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding third prediction data; and take the third prediction data and second label data corresponding to the current training record as a corresponding third prediction-label pair.
[0199] Step 324: Input all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value;
[0200] Here, the second model loss function in this embodiment of the invention is implemented based on the L1 loss function, the Smooth L1 loss function, or the L2 loss function;
[0201] Step 325: Identify whether the second loss value meets the preset range of the second loss value; if it does, proceed to step 326; if it does not, based on the preset second model optimizer, perform a round of modulation on the model parameters of the bimolecular transfer integral prediction model in the direction of minimizing the second model loss function, and return to step 323 to continue training when this round of modulation is combined.
[0202] Here, the second loss value range in this embodiment of the invention is a pre-set numerical range; the second model optimizer includes the Adam optimizer and the SGD optimizer;
[0203] Step 326: Take each second data record of the second evaluation set as the corresponding current evaluation record; and take the third training structure, fourth training structure, and second training feature vector of the current evaluation record as the current first molecular structure M1, second molecular structure M2, and bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding fourth prediction data; and take the fourth prediction data and second label data corresponding to the current evaluation record as a corresponding fourth prediction-label pair; and take all the obtained fourth prediction-label pairs into the preset second model evaluation function to calculate the corresponding second evaluation value.
[0204] Here, the second model evaluation function in this embodiment of the invention is implemented based on the MAE function, MSE function, or RMSE function;
[0205] Step 327: Identify whether the second evaluation value meets the preset range of the second evaluation value; if not, return to step 322 to continue training; if it does, extract the current model parameters of the bimolecular transfer integral prediction model as the corresponding crystal version model parameter set and save it.
[0206] Here, the second evaluation value range in this embodiment of the invention is a pre-set numerical range.
[0207] Step 4: Receive the user's input of the material structure type and the material bimolecular system; use the thin film model parameter set or crystal model parameter set corresponding to the material structure type to set the model parameters of the bimolecular transfer integral prediction model to obtain the corresponding current prediction model; and perform bimolecular structure extraction and prior property calculation based on the material bimolecular system to obtain a set of corresponding first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X; input the current first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X into the current prediction model to make predictions and obtain the corresponding transfer integral prediction data, which is then fed back to the current user.
[0208] Specifically, this includes: Step 41, receiving user input of the material structure type and material bimolecular system;
[0209] The material structure types include thin film structures and crystal structures; the material bimolecular system includes multiple system atoms; the atomic properties of each system atom include atomic element type, atomic three-dimensional coordinates and molecular identifier; the molecular identifier includes first molecular identifier and second molecular identifier;
[0210] Step 42, and use the thin film model parameter set or crystal model parameter set corresponding to the material structure type to set the model parameters of the bimolecular transfer integral prediction model to obtain the corresponding current prediction model;
[0211] Specifically, this includes: identifying the material structure type; if the material structure type is a thin film structure, setting the model parameters of the bimolecular transfer integral prediction model based on the thin film model parameter set; if the material structure type is a crystal structure, setting the model parameters of the bimolecular transfer integral prediction model based on the crystal model parameter set; and using the bimolecular transfer integral prediction model with the completed parameter settings as the corresponding current prediction model.
[0212] Step 43, and based on the material bimolecular system, extract the bimolecular structure and calculate the prior properties to obtain a set of corresponding first molecular structure M1, second molecular structure M2 and bimolecular prior property vector X;
[0213] Specifically, it includes: Step 431, forming a corresponding atomic feature by combining the atomic element type and three-dimensional coordinates of the system atoms that belong to each molecule in the material bimolecular system as the first molecule identifier, and forming a corresponding first molecular structure M1 by combining all the atomic features corresponding to the first molecule identifier.
[0214] Step 432: The atomic element type and three-dimensional coordinates of the system atoms that belong to each molecule in the material bimolecular system are identified as the second molecule identifier to form a corresponding atomic feature, and all the atomic features corresponding to the second molecule identifier are combined to form a corresponding second molecular structure M2.
[0215] Step 433: Based on cheminformatics tools, identify the total number of aromatic rings of the first molecule corresponding to the first molecule identifier in the bimolecular system of the material to obtain a corresponding total number of aromatic rings of the first molecule; and based on quantum chemical calculation tools, calculate the molecular polarity index of the first molecule in the bimolecular system of the material to obtain a corresponding first molecule polarity index; and based on quantum chemical calculation tools, calculate the molecular dipole moment of the second molecule corresponding to the second molecule identifier in the bimolecular system of the material to obtain a corresponding second molecule dipole moment; and based on quantum chemical calculation tools, calculate the molecular polarity index of the second molecule in the bimolecular system of the material. A corresponding second molecule polarity index is obtained; and based on quantum chemical calculation tools, the orbital overlap integral between the first and second molecules in the bimolecular system is calculated using a semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; and based on quantum chemical calculation tools, the intermolecular interaction energy between the first and second molecules in the bimolecular system is calculated using a semi-empirical method; and a corresponding bimolecular a priori characteristic vector X is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current bimolecular system.
[0216] Step 434, and output the first molecular structure M1, the second molecular structure M2, and the bimolecular prior property vector X obtained in this calculation as the result of this calculation;
[0217] Step 44: Input the current first molecular structure M1, second molecular structure M2 and bimolecular prior characteristic vector X into the current prediction model to make predictions and obtain the corresponding transfer integral prediction data, which is then fed back to the current user.
[0218] Figure 3This is a block diagram of a processing device for a bimolecular transfer integral prediction model provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 3 As shown, the device includes: a model building module 201, a dataset building module 202, a model training module 203, and a model application module 204.
[0219] The model building module 201 is used to build a deep learning model that fuses the structure and prior properties of bimolecules and predicts bimolecule transfer integrals based on the fused features, denoted as the corresponding bimolecule transfer integral prediction model. The bimolecule transfer integral prediction model is used to fuse the structural features of the first molecular structure M1 and the second molecular structure M2 input to the model, fuse the features based on the bimolecule fusion features and the bimolecule prior property vector X input to the model, and predict the bimolecule transfer integral based on the structure-priority fusion features and output the corresponding transfer integral prediction data. The first molecular structure M1 and the second molecular structure M2 are both composed of multiple atomic features, each atomic feature is composed of atomic element type and atomic three-dimensional coordinates. The bimolecule prior property vector X is composed of the prior properties of the first molecule and the second molecule corresponding to the first molecular structure M1 and the second molecular structure M2, specifically including the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecule orbital overlap integral calculated based on the semi-empirical method, and the intermolecular interaction energy. The semi-empirical method includes the GFN2-xTB method.
[0220] The dataset construction module 202 is used to construct a first dataset by collecting bimolecular information of thin-film structure optoelectronic materials; and to construct a second dataset by collecting bimolecular information of crystal structure optoelectronic materials.
[0221] The model training module 203 is used to train the bimolecular transfer integral prediction model based on the first dataset to obtain the corresponding thin film version model parameter set; and to train the bimolecular transfer integral prediction model based on the second dataset to obtain the corresponding crystal version model parameter set.
[0222] The model application module 204 receives the material structure type and material bimolecular system input by the user; it sets the model parameters of the bimolecular transfer integral prediction model using the thin film model parameter set or crystal model parameter set corresponding to the material structure type to obtain the corresponding current prediction model; it extracts the bimolecular structure and calculates prior properties based on the material bimolecular system to obtain a set of corresponding first molecular structure M1, second molecular structure M2 and bimolecular prior property vector X; it inputs the current first molecular structure M1, second molecular structure M2 and bimolecular prior property vector X into the current prediction model to obtain the corresponding transfer integral prediction data and feeds it back to the current user; the material structure type includes thin film structure and crystal structure; the material bimolecular system includes multiple system atoms; the atomic properties of each system atom include atomic element type, atomic three-dimensional coordinates and molecular identifier; the molecular identifier includes first molecular identifier and second molecular identifier.
[0223] The processing device for a bimolecular transfer integral prediction model provided in this embodiment of the invention can execute the method steps in the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.
[0224] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing elements; they can be fully implemented in hardware; or some modules can be implemented by processing elements calling software, while others are implemented in hardware. For example, the model building module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0225] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).
[0226] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0227] Figure 4 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 4As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.
[0228] exist Figure 4 The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include Non-Volatile Memory, such as at least one disk storage device.
[0229] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0230] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.
[0231] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for processing a bimolecular transfer integral prediction model. As described above, this invention uses six types of bimolecular properties that can be quickly obtained through simple calculations as prior properties: the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral and the intermolecular interaction energy calculated based on semi-empirical methods (such as the GFN2-xTB method), and uses the bimolecular transfer integral as the prediction target. A deep learning model (bimolecular transfer integral prediction model) is thus constructed to predict the bimolecular transfer integral based on two three-dimensional molecular structures (first molecular structure M1, second molecular structure M2) and six types of prior properties (bimolecular prior property vector X). A corresponding first / second dataset is constructed by collecting bimolecular information from thin-film / crystal structure optoelectronic materials. The bimolecular transfer integral prediction model is then trained on the first / second datasets to obtain two sets of model parameters (thin-film model parameter set, crystal model parameter set) corresponding to the two types of structures. After model training, the two sets of model parameters from the bimolecular transfer integral prediction model are used to predict the bimolecular transfer integral of any bimolecular system in the field of thin-film / crystal structure optoelectronic materials. The embodiments of the present invention shorten the analysis time, improve the analysis efficiency and prediction accuracy; based on the embodiments of the present invention, high-throughput screening of optoelectronic material molecules with thin film / crystal structure can effectively save computing resources and optimize screening efficiency, providing effective assistance for accelerating the research and development iteration of new materials.
[0232] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0233] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for processing a bimolecular transfer integral prediction model, characterized in that, The method includes: A deep learning model is constructed to fuse the structure and prior properties of bimolecules and predict bimolecule transfer integrals based on the fused features, denoted as the corresponding bimolecule transfer integral prediction model. This model fuses structural features of the first molecular structure M1 and the second molecular structure M2 input to the model, fuses features based on the bimolecule fusion features and the prior property vector X input to the model, and predicts bimolecule transfer integrals based on the structure-priority fusion features, outputting the corresponding transfer integral prediction data. Both the first molecular structure M1 and the second molecular structure M2 are composed of multiple atomic features, each consisting of an atomic element type and three-dimensional atomic coordinates. The bimolecule prior property vector X is composed of the prior properties of the first and second molecules corresponding to the first molecular structure M1 and the second molecular structure M2, specifically including the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecule orbital overlap integral calculated based on a semi-empirical method, and the intermolecular interaction energy. The semi-empirical method includes the GFN2-xTB method. The first dataset was constructed by collecting bimolecular information of thin-film structure optoelectronic materials; and the second dataset was constructed by collecting bimolecular information of crystal structure optoelectronic materials. The bimolecular transfer integral prediction model is trained based on the first dataset to obtain the corresponding thin-film model parameter set; and the bimolecular transfer integral prediction model is trained based on the second dataset to obtain the corresponding crystalline model parameter set. The system receives user input regarding the material structure type and the bimolecular system. It then uses the parameter set of the thin-film model or the parameter set of the crystal model corresponding to the material structure type to set the model parameters of the bimolecular transfer integral prediction model, obtaining the corresponding current prediction model. Based on the bimolecular system, it performs bimolecular structure extraction and prior property calculation to obtain a set of corresponding first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X. The system then inputs the current first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X into the current prediction model for prediction, obtaining the corresponding transfer integral prediction data, which is then fed back to the current user. The material structure type includes thin-film structure and crystal structure; the bimolecular system includes multiple system atoms; the atomic properties of each system atom include the atomic element type, the atomic three-dimensional coordinates, and the molecular identifier; the molecular identifier includes a first molecular identifier and a second molecular identifier. The first model input terminal of the bimolecular transfer integral prediction model is used to receive the first molecular structure M1, the second model input terminal is used to receive the second molecular structure M2, the third model input terminal is used to receive the bimolecular prior characteristic vector X, and the model output terminal is used to output the corresponding transfer integral prediction data. The bimolecular transfer integral prediction model includes a first structural encoder, a second structural encoder, a first feature fusion module, a fused feature encoder, a feature mapping network, a prior feature encoder, a second feature fusion module, and a prediction output layer. The input of the first structural encoder is connected to the input of the first model, and its output is connected to the first input of the first feature fusion module; the input of the second structural encoder is connected to the input of the second model, and its output is connected to the second input of the first feature fusion module; the output of the first feature fusion module is connected to the input of the fused feature encoder; the output of the fused feature encoder is connected to the input of the feature mapping network; the output of the feature mapping network is connected to the first input of the second feature fusion module; the output of the prior feature encoder is connected to the second input of the second feature fusion module; the output of the second feature fusion module is connected to the input of the prediction output layer; and the output of the prediction output layer is connected to the model output.
2. The processing method for the bimolecular transfer integral prediction model according to claim 1, characterized in that, The first structure encoder is implemented based on a pre-trained Uni-Mol model; the first structure encoder is used to perform atomic-level high-dimensional feature encoding on the first molecular structure M1 using the pre-trained Uni-Mol model to obtain the corresponding feature tensor E1, which is then sent to the first feature fusion module; the feature tensor E1 has a shape of N1×C. A N1 is the total number of atoms in the first molecular structure M1, C A The preset atomic feature dimension; The feature tensor E1 consists of N1 vectors of length C. A Composed of atomic feature vectors; The second structure encoder is implemented based on a pre-trained Uni-Mol model; the second structure encoder is used to perform atomic-level high-dimensional feature encoding on the second molecular structure M2 using the pre-trained Uni-Mol model to obtain the corresponding feature tensor E2, which is then sent to the first feature fusion module; the feature tensor E2 has a shape of N2×C. A N2 is the total number of atoms in the second molecular structure M2; the feature tensor E2 is composed of N2 atomic feature vectors; The first feature fusion module is used to merge the feature tensor E1 and the feature tensor E2 to obtain the corresponding merged feature tensor H1 and send it to the fusion feature encoder; the shape of the merged feature tensor H1 is (N1+N2)×C A The merged feature tensor H1 is composed of N1+N2 atomic feature vectors. The fusion feature encoder is implemented based on the encoder model of the Transformer architecture; the fusion feature encoder is used to treat each atomic feature vector of the merged feature tensor H1 as a corresponding token embedding encoding vector t. i 1 ≤ index i ≤ (N1 + N2); and in ascending order of index i, the N1 + N2 token embedding encoding vectors t are processed. i Sort the data to obtain the corresponding first vector sequence {t} i }; and based on the position embedding encoding rules of the Transformer architecture, according to each of the token embedding encoding vectors t i Position embedding encoding is performed on index i to obtain a vector of length C. A Position embedding encoding vector p i ; and based on each of the token embedding encoding vectors t i and its corresponding position embedding encoding vector p i A new token embedding encoding vector is calculated. =t i +p i And in ascending order of index i, the N1+N2 token embedding encoding vectors are processed. Sort the results to obtain the corresponding second vector sequence { }; and the second vector sequence { The encoder model of the input Transformer architecture is processed to obtain the corresponding encoded feature vector sequence H2, which is then sent to the feature mapping network; the encoded feature vector sequence H2 consists of N1+N2 encoded feature vectors. The encoded feature vector is formed by sequential sorting; With the token embedding encoding vector One-to-one correspondence; The feature mapping network is used to map each of the encoded feature vectors in the encoded feature vector sequence H2 using a built-in MLP model. Do a job from C A 3D feature space to C B Vector mapping in the 3D feature space yields a vector of length C. B Mapping feature vector And consisting of N1+N2 mapped feature vectors Form a shape of (N1+N2)×C B The mapping feature tensor H proj And based on attention pooling, the mapping feature tensor H is processed. proj The attention weights of N1+N2 feature data points in each feature dimension are calculated, and a weighted sum is calculated on the N1+N2 feature data points of the current feature dimension based on the N1+N2 attention weights of each feature dimension. The weighted sum calculation result corresponding to each feature dimension is used as the feature pooling data corresponding to the current feature dimension; and the obtained C... B Each of the aforementioned feature pooling data forms a corresponding feature vector H3, which is then sent to the second feature fusion module. Among them, C B These are preset molecular feature dimensions; The mapping feature vector The vector mapping process is as follows: ; σ1 and σ2 are two activation functions corresponding to the feature mapping network, where σ1 is either a ReLU or a Sigmoid activation function, and σ2 is either a ReLU or a Sigmoid activation function; W1 and W2 are two weight matrix parameters corresponding to the feature mapping network, and b1 and b2 are two offset vector parameters corresponding to the feature mapping network; the encoded feature vector The shape is C A ×1, the shape of the weight matrix parameter W1 is C B ×C A The shape of the offset vector parameter b1 is C. B ×1, the shape of the weight matrix parameter W2 is C B ×C B The shape of the offset vector parameter b2 is C. B ×1, the mapping feature vector The shape is C B ×1; The prior feature encoder is used to normalize the prior feature data of the bimolecular prior feature vector X to obtain the corresponding normalized vector V; and based on the built-in MLP model, it performs a step from C to V on the normalized vector V. X 3D feature space to C B Vector mapping in the 3D feature space yields a vector of length C. B The feature vector H4 is sent to the second feature fusion module; Wherein, the vector length of the bimolecular prior property vector X is C. X The length of the feature vector H4 is C. B C X The preset length of the bimolecular prior property vector; The vector mapping process of the feature vector H4 is as follows: ; σ3 and σ4 are two activation functions of the prior feature encoder, where σ3 is either a ReLU or a Sigmoid activation function, and σ4 is either a ReLU or a Sigmoid activation function; W3 and W4 are two weight matrix parameters corresponding to the prior feature encoder, and b3 and b4 are two offset vector parameters corresponding to the prior feature encoder; the normalized vector V has a shape of C. X ×1, the shape of the weight matrix parameter W3 is C B ×C X The shape of the offset vector parameter b3 is C. B ×1, the shape of the weight matrix parameter W4 is C B ×C B The shape of the offset vector parameter b4 is C. B ×1, the shape of the feature vector H4 is C B ×1; The second feature fusion module is used to concatenate the feature vectors H3 and H4 to obtain the corresponding concatenated vector H5, which is then sent to the prediction output layer; the length of the concatenated vector H5 is 2C. B ; The prediction output layer is used to perform bimolecular transfer integral prediction based on the splicing vector H5 using the built-in fully connected layer and output the corresponding transfer integral prediction data. The prediction process for the transfer integral prediction data is as follows: ; Y represents the transfer integral prediction data, W5 and b5 are the weight matrix parameters and offset vector parameters corresponding to the prediction output layer, respectively; the shape of the concatenation vector H5 is 2C. B ×1, the shape of the weight matrix parameter W5 is 1×2C B The offset vector parameter b5 is a scalar with a shape of 1×1, and the transfer integral prediction data Y is a scalar with a shape of 1×1.
3. The processing method for the bimolecular transfer integral prediction model according to claim 1, characterized in that, The first dataset includes multiple first data records; each first data record corresponds to a bimolecular structure of a class of thin-film optoelectronic materials; the first data record includes a first training structure, a second training structure, a first training feature vector, and first label data; the data structures of the first training structure and the second training structure are consistent with the corresponding first molecular structure M1 and second molecular structure M2; the data structure of the first training feature vector is consistent with the bimolecular prior feature vector X; the first label data is the bimolecular transfer integral of the current bimolecular structure. The second dataset includes multiple second data records; each second data record corresponds to a bimolecular structure of a class of crystal structure optoelectronic materials; the second data record includes a third training structure, a fourth training structure, a second training feature vector, and second label data; The data structures of the third training structure and the fourth training structure are consistent with the corresponding first molecular structure M1 and second molecular structure M2; the data structure of the second training feature vector is consistent with the bimolecular prior feature vector X; the second label data is the bimolecular transfer integral of the current bimolecular structure.
4. The processing method for the bimolecular transfer integral prediction model according to claim 3, characterized in that, The construction of the first dataset by collecting bimolecular information of thin-film optoelectronic materials specifically includes: Step 41: Collect big data on various molecular pair structures of thin-film optoelectronic materials through a preset first big data channel to obtain the corresponding first molecular pair dataset; The first big data channel includes publicly available molecular information databases for optoelectronic materials, publicly available molecular information databases for thin-film structure optoelectronic materials, publicly available experimental information databases for thin-film structure optoelectronic materials, and publicly available technical literature on thin-film structure optoelectronic materials. The first molecular pair dataset includes multiple first molecular pairs; the two molecules corresponding to the first molecular pair are denoted as molecule A1 and molecule A2; the first molecular pair includes molecular sequence S1 and molecular sequence S2; the molecular sequence S1 and the molecular sequence S2 are respectively the SMILES sequences of the corresponding molecules A1 and A2; Step 42: Take each of the first molecule pairs as the corresponding current molecule pair; Step 43: Based on a preset cheminformatics tool, a three-dimensional molecular conformation is created according to the molecular sequence S1 and the molecular sequence S2 of the current molecular pair to obtain the corresponding first molecular conformation and second molecular conformation. The cheminformatics tools mentioned include Open Babel software and RDKit software; Step 44: Based on a preset molecular dynamics simulation tool, optimize the stable conformations of the first molecular conformation and the second molecular conformation to obtain the corresponding first optimized conformation and the second optimized conformation; and based on the molecular dynamics simulation tool, simulate the thin film structure bimolecular system composed of the first optimized conformation and the second optimized conformation to obtain the corresponding first bimolecular system. The molecular dynamics simulation tools include Gaussian software and GROMACS software; Step 45: Extract the element type and three-dimensional coordinates of each atom of molecule A1 in the first bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding first training structure from all the atomic features corresponding to molecule A1. Step 46: Extract the element type and three-dimensional coordinates of each atom of molecule A2 in the first bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding second training structure from all the atomic features corresponding to molecule A2. Step 47: Based on the cheminformatics tool, identify the total number of aromatic rings of molecule A1 in the first bimolecular system to obtain a corresponding total number of aromatic rings of the first molecule; and based on a preset quantum chemical calculation tool, calculate the molecular polarity index of molecule A1 in the first bimolecular system to obtain a corresponding first molecular polarity index; and based on the quantum chemical calculation tool, calculate the molecular dipole moment of molecule A2 in the first bimolecular system to obtain a corresponding second molecular dipole moment; and based on the quantum chemical calculation tool, calculate the molecular polarity index of molecule A2 in the first bimolecular system to obtain a corresponding second molecular polarity index. Based on the quantum chemical calculation tool, the orbital overlap integral between molecule A1 and molecule A2 in the first bimolecular system is calculated using the semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; based on the quantum chemical calculation tool, the intermolecular interaction energy between molecule A1 and molecule A2 in the first bimolecular system is calculated using the semi-empirical method; and a corresponding first training characteristic vector is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current first bimolecular system; The quantum chemical calculation tools include Multiwfn software, Gaussian software, ORCA software, GAMESS software, NWChem software, SHARC software, GPAW software, and Newton-X software. Step 48: Based on the quantum chemical calculation tool, calculate the bimolecular transfer integral of molecule A1 and molecule A2 in the first bimolecular system according to the preset calculation method, and use the calculation result as a corresponding first tag data; The preset calculation methods include the DFT method and the TD-DFT method; Step 49: The first training structure, the second training structure, the first training feature vector, and the first label data corresponding to the current molecular pair are combined to form a corresponding first data record; and all the obtained first data records are combined to form the corresponding first dataset.
5. The processing method for the bimolecular transfer integral prediction model according to claim 4, characterized in that, The construction of the second dataset by collecting bimolecular information of crystal structure optoelectronic materials specifically includes: Step 51: Collect big data on various molecular pair structures of crystal structure optoelectronic materials through a preset second big data channel to obtain the corresponding second molecular pair dataset; The second big data channel includes publicly available molecular information databases for optoelectronic materials, publicly available molecular information databases for crystal structure optoelectronic materials, publicly available experimental information databases for crystal structure optoelectronic materials, and publicly available technical literature on crystal structure optoelectronic materials. The second molecular pair dataset includes multiple second molecular pairs; the two molecules corresponding to the second molecular pairs are denoted as molecules A3 and A4; the second molecular pair includes molecular sequence S3 and molecular sequence S4; the molecular sequence S3 and the molecular sequence S4 are the SMILES sequences of the corresponding molecules A3 and A4, respectively; Step 52: Take each of the second molecular pairs as the corresponding current molecular pairs; Step 53: Based on the cheminformatics tool, a three-dimensional molecular conformation is created according to the molecular sequence S3 and the molecular sequence S4 of the current molecular pair to obtain the corresponding third molecular conformation and fourth molecular conformation. Step 54: Based on the molecular dynamics simulation tool, optimize the stable conformations of the third and fourth molecular conformations to obtain the corresponding third optimized conformation and fourth optimized conformation; and based on the molecular dynamics simulation tool, simulate the crystal structure bimolecular system composed of the third optimized conformation and the fourth optimized conformation to obtain the corresponding second bimolecular system. Step 55: Extract the element type and three-dimensional coordinates of each atom of molecule A3 in the second bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding third training structure from all the atomic features corresponding to molecule A3. Step 56: Extract the element type and three-dimensional coordinates of each atom of molecule A4 in the second bimolecular system as a set of corresponding atomic element types and atomic three-dimensional coordinates to form a corresponding atomic feature, and form a corresponding fourth training structure from all the atomic features corresponding to molecule A4. Step 57: Based on the cheminformatics tool, identify the total number of aromatic rings of molecule A3 in the second bimolecular system to obtain a corresponding total number of aromatic rings of the first molecule; and based on the quantum chemical calculation tool, calculate the molecular polarity index of molecule A3 in the second bimolecular system to obtain a corresponding first molecular polarity index; and based on the quantum chemical calculation tool, calculate the molecular dipole moment of molecule A4 in the second bimolecular system to obtain a corresponding second molecular dipole moment; and based on the quantum chemical calculation tool, calculate the molecular polarity index of molecule A4 in the second bimolecular system to obtain a corresponding second molecular polarity index. Based on the quantum chemical calculation tool, the orbital overlap integral between molecule A3 and molecule A4 in the second bimolecular system is calculated using the semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; based on the quantum chemical calculation tool, the intermolecular interaction energy between molecule A3 and molecule A4 in the second bimolecular system is calculated using the semi-empirical method; and a corresponding second training characteristic vector is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current second bimolecular system; Step 58: Based on the quantum chemical calculation tool, calculate the bimolecular transfer integral of molecule A3 and molecule A4 in the second bimolecular system according to the preset calculation method, and use the calculation result as a corresponding second tag data; Step 59: The third training structure, the fourth training structure, the second training feature vector, and the second label data corresponding to the current molecular pair are combined to form a corresponding second data record; and all the obtained second data records are combined to form the corresponding second dataset.
6. The processing method for the bimolecular transfer integral prediction model according to claim 3, characterized in that, The step of training the bimolecular transfer integral prediction model based on the first dataset to obtain the corresponding thin-film model parameter set specifically includes: Step 61: Initialize the model parameters of the bimolecular transfer integral prediction model based on the preset initialization model parameter set; Step 62: Based on a preset first segmentation ratio, the first dataset is randomly divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio; Step 63: Take each of the first data records in the first training set as the corresponding current training record; and take the first training structure, the second training structure, and the first training feature vector of the current training record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding first prediction data; and take the first prediction data corresponding to the current training record and the first label data as a corresponding first prediction-label pair. Step 64: Input all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value; The first model loss function is implemented based on the L1 loss function, the Smooth L1 loss function, or the L2 loss function. Step 65: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 66; if it does not, modulate the model parameters of the bimolecular transfer integral prediction model in one round based on the preset first model optimizer in the direction of minimizing the first model loss function, and return to step 63 to continue training when this round of modulation is combined. The first model optimizer includes the Adam optimizer and the SGD optimizer; Step 66: Take each of the first data records in the first evaluation set as the corresponding current evaluation record; and take the first training structure, the second training structure, and the first training feature vector of the current evaluation record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding second prediction data; and take the second prediction data corresponding to the current evaluation record and the first label data as a corresponding second prediction-label pair; and take all the obtained second prediction-label pairs into the preset first model evaluation function to calculate the corresponding first evaluation value; The first model evaluation function is implemented based on the MAE function, MSE function, or RMSE function. Step 67: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 62 to continue training; if yes, extract the current model parameters of the bimolecular transfer integral prediction model as the corresponding thin film model parameter set and save it.
7. The processing method for the bimolecular transfer integral prediction model according to claim 3, characterized in that, The step of training the bimolecular transfer integral prediction model based on the second dataset to obtain the corresponding crystal model parameter set specifically includes: Step 71: Initialize the model parameters of the bimolecular transfer integral prediction model based on the preset initialization model parameter set; Step 72: Based on the preset second segmentation ratio, the second dataset is randomly divided into two subsets, which are denoted as the corresponding second training set and second evaluation set. The second training set and the second evaluation set are both composed of multiple second data records; the ratio of the total number of records in the second training set to the total number of records in the second evaluation set satisfies the second segmentation ratio. Step 73: Take each of the second data records in the second training set as the corresponding current training record; and take the third training structure, the fourth training structure, and the second training feature vector of the current training record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding third prediction data; and take the third prediction data corresponding to the current training record and the second label data as a corresponding third prediction-label pair. Step 74: Input all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value; The second model loss function is implemented based on the L1 loss function, the Smooth L1 loss function, or the L2 loss function. Step 75: Identify whether the second loss value meets the preset second loss value range; if it does, proceed to step 76; if it does not, based on the preset second model optimizer, perform a round of modulation on the model parameters of the bimolecular transfer integral prediction model in the direction of minimizing the second model loss function, and return to step 73 to continue training when this round of modulation is combined. The second model optimizer includes the Adam optimizer and the SGD optimizer; Step 76: Take each of the second data records in the second evaluation set as the corresponding current evaluation record; and take the third training structure, the fourth training structure, and the second training feature vector of the current evaluation record as the current first molecular structure M1, the second molecular structure M2, and the bimolecular prior feature vector X as input to the bimolecular transfer integral prediction model for prediction, and take the transfer integral prediction data obtained in this prediction as the corresponding fourth prediction data; and take the fourth prediction data corresponding to the current evaluation record and the second label data as a corresponding fourth prediction-label pair; and take all the obtained fourth prediction-label pairs into the preset second model evaluation function to calculate the corresponding second evaluation value. The second model evaluation function is implemented based on the MAE function, MSE function, or RMSE function. Step 77: Identify whether the second evaluation value meets the preset range of the second evaluation value; if not, return to step 72 to continue training; if so, extract the current model parameters of the bimolecular transfer integral prediction model as the corresponding crystal version model parameter set and save it.
8. The processing method for the bimolecular transfer integral prediction model according to claim 1, characterized in that, The step of setting the model parameters of the bimolecular transfer integral prediction model using the thin film model parameter set or the crystal model parameter set corresponding to the material structure type to obtain the corresponding current prediction model specifically includes: The material structure type is identified; if the material structure type is a thin film structure, the model parameters of the bimolecular transfer integral prediction model are set based on the thin film model parameter set; if the material structure type is a crystal structure, the model parameters of the bimolecular transfer integral prediction model are set based on the crystal model parameter set; and the bimolecular transfer integral prediction model whose parameters have been set is used as the corresponding current prediction model.
9. The processing method for the bimolecular transfer integral prediction model according to claim 4, characterized in that, The step of extracting the bimolecular structure and calculating the prior properties of the material bimolecular system to obtain a set of corresponding first molecular structure M1, second molecular structure M2, and bimolecular prior property vector X specifically includes: Step 91: The atomic element type and the three-dimensional coordinates of the atoms of the system atoms that are identified as the first molecule identifier in the material bimolecular system are combined to form a corresponding atomic feature, and all the atomic features corresponding to the first molecule identifier are combined to form a corresponding first molecular structure M1; Step 92: The atomic element type and the three-dimensional coordinates of the atoms of the system atoms that are identified as the second molecule identifier in the material bimolecular system are combined to form a corresponding atomic feature, and all the atomic features corresponding to the second molecule identifier are combined to form a corresponding second molecular structure M2; Step 93: Based on the cheminformatics tool, identify the total number of aromatic rings of the first molecule corresponding to the first molecule identifier in the material bimolecular system to obtain a corresponding total number of aromatic rings of the first molecule; and based on the quantum chemical calculation tool, calculate the molecular polarity index of the first molecule in the material bimolecular system to obtain a corresponding first molecule polarity index; and based on the quantum chemical calculation tool, calculate the molecular dipole moment of the second molecule corresponding to the second molecule identifier in the material bimolecular system to obtain a corresponding second molecule dipole moment; and based on the quantum chemical calculation tool, calculate the molecular polarity index of the second molecule in the material bimolecular system to obtain a... The corresponding second molecule polarity index; and based on the quantum chemical calculation tool, the orbital overlap integral between the first molecule and the second molecule in the material bimolecular system is calculated using the semi-empirical method to obtain the corresponding bimolecular orbital overlap integral; and based on the quantum chemical calculation tool, the intermolecular interaction energy between the first molecule and the second molecule in the material bimolecular system is calculated using the semi-empirical method; and a corresponding bimolecular a priori characteristic vector X is formed by the total number of aromatic rings of the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecular orbital overlap integral, and the intermolecular interaction energy corresponding to the current material bimolecular system; Step 94, and output the first molecular structure M1, the second molecular structure M2, and the bimolecular prior property vector X obtained in this calculation as the result of this calculation.
10. An apparatus for performing a processing method for a bimolecular transfer integral prediction model according to any one of claims 1-9, characterized in that, The device includes: a model building module, a dataset building module, a model training module, and a model application module; The model building module is used to construct a deep learning model that fuses the structure and prior properties of bimolecules and predicts bimolecule transfer integrals based on the fused features, denoted as the corresponding bimolecule transfer integral prediction model. The bimolecule transfer integral prediction model fuses the structural features of the first molecular structure M1 and the second molecular structure M2 input to the model, fuses the bimolecule fusion features with the prior property vector X input to the model, and predicts bimolecule transfer integrals based on the structure-priority fusion features, outputting the corresponding transfer integral prediction data. Both the first molecular structure M1 and the second molecular structure M2 are composed of multiple atomic features, each of which consists of an atomic element type and three-dimensional atomic coordinates. The bimolecule prior property vector X is composed of the prior properties of the first and second molecules corresponding to the first molecular structure M1 and the second molecular structure M2, specifically including the total number of aromatic rings in the first molecule, the polarity index of the first molecule, the dipole moment of the second molecule, the polarity index of the second molecule, the bimolecule orbital overlap integral calculated based on a semi-empirical method, and the intermolecular interaction energy. The semi-empirical method includes the GFN2-xTB method. The dataset construction module is used to construct a first dataset by collecting bimolecular information of thin-film structure optoelectronic materials, and to construct a second dataset by collecting bimolecular information of crystal structure optoelectronic materials. The model training module is used to train the bimolecular transfer integral prediction model based on the first dataset to obtain the corresponding thin-film model parameter set; and to train the bimolecular transfer integral prediction model based on the second dataset to obtain the corresponding crystal model parameter set; The model application module receives the material structure type and bimolecular system input by the user; it sets the model parameters of the bimolecular transfer integral prediction model using the thin film model parameter set or the crystal model parameter set corresponding to the material structure type to obtain the corresponding current prediction model; it extracts the bimolecular structure and calculates prior properties based on the material bimolecular system to obtain a set of corresponding first molecular structure M1, second molecular structure M2 and bimolecular prior property vector X; and it inputs the current first molecular structure M1, second molecular structure M2 and bimolecular prior property vector X into the current prediction model to obtain the corresponding transfer integral prediction data and feeds it back to the current user; the material structure type includes thin film structure and crystal structure; the material bimolecular system includes multiple system atoms; the atomic properties of each system atom include the atomic element type, the atomic three-dimensional coordinates and the molecular identifier; the molecular identifier includes the first molecular identifier and the second molecular identifier; The first model input terminal of the bimolecular transfer integral prediction model is used to receive the first molecular structure M1, the second model input terminal is used to receive the second molecular structure M2, the third model input terminal is used to receive the bimolecular prior characteristic vector X, and the model output terminal is used to output the corresponding transfer integral prediction data. The bimolecular transfer integral prediction model includes a first structural encoder, a second structural encoder, a first feature fusion module, a fused feature encoder, a feature mapping network, a prior feature encoder, a second feature fusion module, and a prediction output layer. The input of the first structural encoder is connected to the input of the first model, and its output is connected to the first input of the first feature fusion module; the input of the second structural encoder is connected to the input of the second model, and its output is connected to the second input of the first feature fusion module; the output of the first feature fusion module is connected to the input of the fused feature encoder; the output of the fused feature encoder is connected to the input of the feature mapping network; the output of the feature mapping network is connected to the first input of the second feature fusion module; the output of the prior feature encoder is connected to the second input of the second feature fusion module; the output of the second feature fusion module is connected to the input of the prediction output layer; and the output of the prediction output layer is connected to the model output.
11. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-9; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-9.