Method and apparatus for predicting optical energy characteristics of organic photovoltaic material molecules

By constructing an optical energy characteristic prediction model, the problem of high complexity in predicting the ionization energy, emission energy, and absorption energy of organic photovoltaic materials in existing technologies has been solved, achieving efficient one-step prediction and reducing computation time.

CN120183563BActive Publication Date: 2025-11-18BEIJING DP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510294704.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-11-18
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

Existing technologies for predicting the ionization energy, emission energy, and absorption energy of organic photovoltaic material molecules are computationally complex, time-consuming, and inefficient, and the three calculation processes are not universally applicable.

Method used

An optical energy characteristic prediction model is constructed by collecting data on the molecular structure, ionization energy, emission energy, and absorption energy curves of various organic photovoltaic materials under temperature and pressure conditions, building a dataset, and training the model to achieve one-step prediction of ionization energy, emission energy, and absorption energy.

Benefits of technology

It reduces prediction complexity, shortens computation time, and improves prediction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183563B_ABST
    Figure CN120183563B_ABST
Patent Text Reader

Abstract

The embodiment of the present application relates to a kind of organic photovoltaic material molecule optical energy characteristic prediction method and device, the method comprises: constructing a kind of optical energy characteristic prediction model for according to model input molecular structure, temperature, pressure and wavelength sequence carry out ionization energy, emission energy and absorption energy prediction processing;And by data acquisition, first data set is constructed;And based on first data set, optical energy characteristic prediction model is trained;And after training, the first molecular structure, the first temperature, the first pressure and the first wavelength sequence input optical energy characteristic prediction model are handled as corresponding molecular structure M, temperature x, pressure y and wavelength sequence Z after training end, and the ionization energy u IE , emission energy sequence S EE And absorption energy sequence S AE As corresponding first ionization energy, first emission energy sequence and first absorption energy sequence are fed back to the current user.The prediction efficiency can be improved by the present application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a method and device for predicting optical energy characteristics of organic photovoltaic material molecules. BACKGROUND

[0002] An organic photovoltaic (OPV) material molecule refers to an organic molecule that can be used as an OPV material. Ionization energy (IE), emission energy (EE), and absorption energy (AE) are three important optical energy characteristic parameters of the organic photovoltaic material molecule. The emission / absorption energy is related to the wavelength of light, and the ionization energy is independent of the wavelength of light. Conventionally, the three optical energy characteristic parameters of the organic photovoltaic material molecule are predicted based on quantum chemistry calculation methods (such as density functional theory calculation), and the three calculation processes cannot be used for each other. This means that the conventional prediction method has the problems of high calculation complexity, long calculation time, and low calculation efficiency. SUMMARY

[0003] The present application is to solve the defects of the prior art, and provides a method and device for predicting optical energy characteristics of organic photovoltaic material molecules, an electronic device, and a computer readable storage medium. The present application customizes an optical energy characteristic prediction model for predicting ionization energy, emission energy, and absorption energy based on the model input of the molecular structure, temperature, pressure, and wavelength sequence. A first data set for model training is constructed by collecting the molecular structures of a plurality of types of organic photovoltaic material molecules and the ionization energy, emission energy curves, and absorption energy curves of each molecule under different temperature and pressure conditions. The optical energy characteristic prediction model is trained based on the first data set. After the training is completed, the user input molecular structure, temperature, pressure, and wavelength sequence are input into the optical energy characteristic prediction model for processing, and the ionization energy, emission energy sequence, and absorption energy sequence output by the model processing are fed back to the current user. The present application realizes one-step prediction of three optical energy characteristic parameters (ionization energy, emission energy, and absorption energy) through a customized optical energy characteristic prediction model, which can effectively reduce the prediction complexity, shorten the prediction time, and improve the prediction efficiency.

[0004] To achieve the above-mentioned purpose, the first aspect of the present application provides a method for predicting optical energy characteristics of organic photovoltaic material molecules, which comprises:

[0005] An optical energy characteristic prediction model is constructed to predict three optical energy characteristic parameters of organic photovoltaic material molecules: ionization energy, emission energy, and absorption energy. The optical energy characteristic prediction model is used to predict the ionization energy, emission energy, and absorption energy based on the molecular structure M, temperature x, pressure y, and wavelength sequence Z input to the model, and outputs the corresponding ionization energy u. IE Emission energy sequence S EE and absorption energy sequence S AE ;

[0006] A dataset for model training is constructed by collecting data on the molecular structure of various organic photovoltaic materials and the ionization energy, emission energy curves, and absorption energy curves of each molecule under different temperature and pressure conditions. This dataset is denoted as the first dataset. The optical energy characteristic prediction model is then trained based on the first dataset.

[0007] After training, the system receives the user's input of the first molecular structure, first temperature, first pressure, and first wavelength sequence; and inputs the first molecular structure, first temperature, first pressure, and first wavelength sequence as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z into the optical energy characteristic prediction model for processing. The ionization energy u output by this model processing is then used. IE The emission energy sequence S EE and the absorption energy sequence S AE The corresponding first ionization energy, first emission energy sequence, and first absorption energy sequence are fed back to the current user.

[0008] Preferably, the molecular structure M comprises multiple atoms m i The atom m i Atomic parameters include atom type and atom three-dimensional coordinates; 1 ≤ atom index i ≤ N M N M The total number of atoms in the molecular structure M;

[0009] The wavelength sequence Z consists of multiple wavelengths z j Arranged in ascending order of wavelength; 1 ≤ wavelength index j ≤ N Z N Z The sequence length of the wavelength sequence Z;

[0010] The ionization energy u IE Let it be a real number;

[0011] The emission energy sequence S EE Composed of multiple emission energies EE,j The emission energy sequence S is formed by sequential sorting; EE The emission energy sEE,j the wavelength z of the wavelength sequence Z j correspond one-to-one;

[0012] the absorption energy sequence S AE by a plurality of absorption energies s AE,j sequentially sorted; the absorption energy sequence S AE of the absorption energy s AE,j correspond one-to-one; j correspond one-to-one;

[0013] The first data set includes a plurality of first data records; the first data record includes a first training molecular structure, a first training temperature, a first training pressure, a first training wavelength sequence, a first label ionization energy, a first label emission energy sequence, and a first label absorption energy sequence; the data structure of the first training molecular structure, the first training wavelength sequence, the first label emission energy sequence, and the first label absorption energy sequence are consistent with the corresponding molecular structure M, wavelength sequence Z, emission energy sequence S EE , and absorption energy sequence S AE ; the sequence length of the first training wavelength sequence, the first label emission energy sequence, and the first label absorption energy sequence is the same, and the first training wavelength of the first training wavelength sequence corresponds one-to-one with the first label emission energy of the first label emission energy sequence and the first label absorption energy of the first label absorption energy sequence.

[0014] Preferably, the first model input end of the optical energy characteristic prediction model is used to receive the molecular structure M of the model input, the second model input end is used to receive the temperature x, the pressure y, and the wavelength sequence Z of the model input, the first model output end is used to output the corresponding ionization energy u IE , the second model output end is used to output the corresponding emission energy sequence S EE , and the third model output end is used to output the corresponding absorption energy sequence S AE .

[0015] The optical energy characteristic prediction model includes a first embedding encoding module, a molecular structure encoding module, a second embedding encoding module, a vector space mapping module, a feature fusion module, an ionization energy prediction module, an emission energy prediction module, and an absorption energy prediction module.

[0016] The input end of the first embedding coding module is connected with the first model input end, and the output end is connected with the input end of the molecular structure coding module; the output end of the molecular structure coding module is connected with the first input end of the feature fusion module; the input end of the second embedding coding module is connected with the second model input end, and the output end is connected with the input end of the vector space mapping module; the output end of the vector space mapping module is connected with the second input end of the feature fusion module; the first output end of the feature fusion module is connected with the input end of the ionization energy prediction module, the second output end is connected with the input end of the emission energy prediction module, and the third output end is connected with the input end of the absorption energy prediction module; the output end of the ionization energy prediction module is connected with the first model output end; the output end of the emission energy prediction module is connected with the second model output end; and the output end of the absorption energy prediction module is connected with the third model output end;

[0017] The first embedding coding module is used for atom type one-hot coding of the molecular structure M according to the atom type one-hot coding rule of the Uni-Mol model i to obtain a corresponding atom coding vector, and the obtained N M atom coding vectors form a corresponding atom coding tensor; and according to all three-dimensional atomic coordinates of the molecular structure M, a pair coding tensor with a tensor shape of N M ×N M is obtained according to the pair feature initialization coding rule of the Uni-Mol model; and the obtained atom coding tensor and pair coding tensor are sent to the molecular structure coding module;

[0018] The molecular structure coding module is realized based on the Uni-Mol model; the molecular structure coding module is used for atom feature and pair feature extraction processing according to the input atom coding tensor and pair coding tensor to obtain a corresponding hidden feature tensor H, which is sent to the feature fusion module; the hidden feature tensor H is composed of N M hidden feature vectors h i , the hidden feature vector h i corresponds to the atom m i ;

[0019] The second embedding coding module is used for pre-setting a plurality of first Gaussian kernel functions, a plurality of second Gaussian kernel functions and a plurality of third Gaussian kernel functions in the module;

[0020] The second embedding encoding module is further configured to perform corresponding kernel function calculations on the temperature x input to the model according to each of the first Gaussian kernel functions, and use the calculated numerical results as the corresponding first result values; and form a corresponding temperature encoding vector from all the obtained first result values; and perform corresponding kernel function calculations on the pressure y input to the model according to each of the second Gaussian kernel functions, and use the calculated numerical results as the corresponding second result values; and form a corresponding pressure encoding vector from all the obtained second result values; and perform kernel function calculations on each of the wavelengths z of the wavelength sequence Z input to the model according to each of the third Gaussian kernel functions. j Perform the corresponding kernel function calculation and use the calculated numerical result as the corresponding third result value, and then use each wavelength z j All the corresponding third result values ​​form a corresponding wavelength encoding vector, and the obtained N Z The wavelength encoding vectors form a corresponding wavelength encoding tensor; and the temperature encoding vector, the pressure encoding vector, and the wavelength encoding tensor are sent to the vector space mapping module; the wavelength encoding vector and the wavelength z j One-to-one correspondence;

[0021] The vector space mapping module is used to pre-set three MLP models within the module;

[0022] The vector space mapping module is further used to map the latent feature vector h. i The corresponding vector space is used as the target vector space; and based on the first built-in MLP model, the temperature encoding vector is mapped to the target vector space to obtain a corresponding temperature mapping vector; and based on the second built-in MLP model, the pressure encoding vector is mapped to the target vector space to obtain a corresponding pressure mapping vector; and based on the third built-in MLP model, each wavelength encoding vector of the wavelength encoding tensor is mapped to the target vector space to obtain a corresponding wavelength mapping vector, and the obtained N... Z The wavelength mapping vectors are combined to form a corresponding wavelength mapping tensor; and the temperature mapping vector, the pressure mapping vector, and the wavelength mapping tensor are sent to the feature fusion module;

[0023] The feature fusion module is used to fuse the latent feature vectors h of the latent feature tensor H. i The temperature mapping vector and the pressure mapping vector are sequentially concatenated to obtain a corresponding atom-temperature-pressure feature fusion vector; and the obtained N MThe atom-temperature-pressure feature fusion vectors are composed of a corresponding first fusion tensor; and a corresponding wavelength-atom-temperature-pressure feature fusion vector is obtained by concatenating each wavelength encoding vector with each of the first fusion vectors; and a wavelength-atom-temperature-pressure feature fusion vector is obtained by concatenating each of the wavelength encoding vectors .... j The corresponding N M The wavelength-atom-temperature-pressure feature fusion vectors are concatenated to form a corresponding wavelength-molecule-temperature-pressure feature fusion vector, and the resulting N... Z The wavelength-molecule-temperature-pressure feature fusion vectors are used to form a corresponding second fusion tensor; and the first fusion tensor is sent to the ionization energy prediction module, and the second fusion tensor is sent to the emission energy prediction module and the absorption energy prediction module;

[0024] The ionization energy prediction module is implemented based on an MLP model; the ionization energy prediction module is used to predict the corresponding ionization energy u based on the input first fusion tensor. IE ;

[0025] The emission energy prediction module is implemented based on an MLP model; the emission energy prediction module is used to predict the corresponding emission energy sequence S based on the input second fusion tensor. EE ;

[0026] The absorption energy prediction module is implemented based on an MLP model; the absorption energy prediction module is used to predict the corresponding absorption energy sequence S based on the input second fusion tensor. AE .

[0027] Preferably, the step of constructing a dataset for model training by collecting data on the molecular structures of various organic photovoltaic material molecules and the ionization energy, emission energy curves, and absorption energy curves of each molecule under different temperature and pressure conditions is denoted as the corresponding first dataset, specifically including:

[0028] Step 41: Collect data on the molecular structure of various organic photovoltaic materials and the ionization energy, emission energy curves and absorption energy curves of each molecule under different temperature and pressure conditions to obtain multiple corresponding first data collections.

[0029] The first collected data includes the first collected molecular structure, the first collected temperature, the first collected pressure, the first collected ionization energy, the first collected emission energy curve, and the first collected absorption energy curve.

[0030] Step 42, set a wavelength range as a first training wavelength range; extract the curve part of the first acquisition emission energy curve of each first acquisition data in the first training wavelength range as a corresponding first extraction curve; extract the curve part of the first acquisition absorption energy curve of each first acquisition data in the first training wavelength range as a corresponding second extraction curve; sample the curve points of each first extraction curve based on a preset sampling wavelength interval Δλ to obtain a plurality of first sampling points corresponding to each first extraction curve to form a corresponding first sampling point sequence; sample the curve points of each second extraction curve based on the sampling wavelength interval Δλ to obtain a plurality of second sampling points corresponding to each second extraction curve to form a corresponding second sampling point sequence;

[0031] Wherein, the sampling point parameter of the first sampling point is composed of a sampling point wavelength and a corresponding sampling point emission energy; the sampling point parameter of the second sampling point is composed of a sampling point wavelength and a corresponding sampling point absorption energy; the total number of the first and second sampling points of the corresponding first and second sampling point sequences of each first acquisition data is the same, that is, the first and second sampling points in the corresponding first and second sampling point sequences of each first acquisition data are one-to-one corresponding, and the sampling point wavelengths of each group of corresponding first and second sampling points are consistent;

[0032] Step 43, iterate all the first acquisition data for one round; and in this round of iteration, the first acquisition data currently iterated is taken as a corresponding current acquisition data; the sampling point wavelength of each sampling point of the first sampling point sequence corresponding to the current acquisition data is taken as a corresponding first training wavelength, and all the first training wavelengths obtained are sequentially sorted to form a corresponding first training wavelength sequence; the sampling point emission energy of each sampling point of the first sampling point sequence corresponding to the current acquisition data is taken as a corresponding first label emission energy, and all the first label emission energies obtained are sequentially sorted to form a corresponding first label emission energy sequence; the sampling point absorption energy of each sampling point of the second sampling point sequence corresponding to the current acquisition data is taken as a corresponding first label absorption energy, and all the first label absorption energies obtained are sequentially sorted to form a corresponding first label absorption energy sequence; the first acquisition molecular structure, the first acquisition temperature, the first acquisition pressure and the first acquisition ionization energy of the current acquisition data are taken as the corresponding first training molecular structure, the first training temperature, the first training pressure and the first label ionization energy; and the first training molecular structure, the first training temperature, the first training pressure, the first training wavelength sequence, the first label ionization energy, the first label emission energy sequence and the first label absorption energy sequence corresponding to the current acquisition data form a corresponding first data record;

[0033] Step 44, after the end of the current round of traversal of all the first collection data, the corresponding first data set is composed of all the first data records obtained.

[0034] Preferably, the model training of the optical energy characteristic prediction model based on the first data set specifically includes:

[0035] Step 51, the first data set is divided into two sub-data sets according to a preset first segmentation ratio, which are recorded as a corresponding first training set and a first evaluation set;

[0036] Wherein, the first training set and the first evaluation set are both composed of a plurality of first data records; the total number of records of the first training set and the first evaluation set meets the first segmentation ratio;

[0037] Step 52, the first first data record of the first training set is taken as a corresponding current training record;

[0038] Step 53, the first label ionization energy, the first label emission energy sequence and the first label absorption energy sequence of the current training record are recorded as a corresponding first label ionization energy u IE,tag , a first label emission energy sequence S EE,tag and a first label absorption energy sequence S AE,tag ;

[0039] Step 54, the first training molecular structure, the first training temperature, the first training pressure and the first training wavelength sequence of the current training record are taken as the corresponding molecular structure M, the temperature x, the pressure y and the wavelength sequence Z, which are input into the optical energy characteristic prediction model for processing, and the ionization energy u IE , the emission energy sequence S EE and the absorption energy sequence S AE output by the current model processing are taken as the corresponding first predicted ionization energy u IE,pre , the first predicted emission energy sequence S EE,pre and the first predicted absorption energy sequence S AE,pre ;

[0040] Step 55, the first label ionization energy u IE,tag and the first predicted ionization energy u IE,pre are taken into a preset first model loss function L M1 ; and the first label emission energy sequence S EE,tag and the first predicted emission energy sequence S EE,pre are taken into a preset second model loss function L M2 ; and the first label absorption energy sequence SAE,tag and the first predicted absorption energy sequence S AE,pre into a preset third model loss function L M3 ; and the first, second, and third model loss functions L M1 , L M2 , and L M3 are added to form a corresponding overall model loss function L total = L M1 + L M2 + L M3 ; and based on a preset first model optimizer, the model parameters of the optical energy characteristic prediction model are modulated in a way that the first, second, and third model loss functions L M1 , L M2 , and L M3 and the overall model loss function L total all reach a minimum value;

[0041] The first model loss function L M1 is implemented based on an L1 loss function, a smooth L1 loss function, or an L2 loss function; the second model loss function L M2 is implemented based on an L1 loss function or an L2 loss function; and the third model loss function L M3 is implemented based on an L1 loss function or an L2 loss function.

[0042] Step 56: identifying whether the current training record is the last first data record of the first training set, if yes, going to step 57, if not, taking the next first data record of the first training set as a new current training record and returning to step 53 to continue training;

[0043] Step 57: performing a round of traversal on all the first data records of the first evaluation set; and in the round of traversal, taking the first data record currently traversed as a corresponding current evaluation record; and taking the first label ionization energy, the first label emission energy sequence, and the first label absorption energy sequence of the current evaluation record as corresponding first label ionization energy u IE,tag , first label emission energy sequence S EE,tag , and first label absorption energy sequence S AE,tag ; and taking the first training molecular structure, the first training temperature, the first training pressure, and the first training wavelength sequence of the current evaluation record as corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z, inputting the optical energy characteristic prediction model for processing and taking the ionization energy u IE , the emission energy sequence S EE , and the absorption energy sequence SAE a corresponding first predicted ionization energy u IE,pre a first predicted emission energy sequence S EE,pre a first predicted absorption energy sequence S AE,pre ; and the first label ionization energy u IE,tag and the first predicted ionization energy u IE,pre ; and the first label emission energy sequence S EE,tag and the first predicted emission energy sequence S EE,pre ; and the first label absorption energy sequence S AE,tag and the first predicted absorption energy sequence S AE,pre ; and the first predicted-label pair is obtained; and the second predicted-label pair is obtained; and the third predicted-label pair is obtained; and at the end of the current iteration, all the first predicted-label pairs are brought into a preset first model evaluation function to obtain a corresponding first evaluation value; all the second predicted-label pairs are brought into a preset second model evaluation function to obtain a corresponding second evaluation value; and all the third predicted-label pairs are brought into a preset third model evaluation function to obtain a corresponding third evaluation value; and the first, second and third evaluation values are weighted and summed to obtain a corresponding fourth evaluation value.

[0044] The first, second and third model evaluation functions are each based on an MAE function, an MSE function or an RMSE function.

[0045] Step 58, the first, second, third and fourth evaluation values are identified; if the first evaluation value does not satisfy a preset first evaluation value range, or the second evaluation value does not satisfy a preset second evaluation value range, or the third evaluation value does not satisfy a preset third evaluation value range, or the third evaluation value satisfies a preset third evaluation value range, or the fourth evaluation value does not satisfy a preset fourth evaluation value range, then return to step 52 for continuous training; if the first evaluation value satisfies the first evaluation value range, the second evaluation value satisfies the second evaluation value range, the third evaluation value satisfies the third evaluation value range, the third evaluation value satisfies the third evaluation value range, and the fourth evaluation value satisfies the fourth evaluation value range, then stop training and confirm that the model training is completed.

[0046] The second aspect of the embodiment of the application provides a device for implementing the optical energy characteristic prediction method of the organic photovoltaic material molecule in the first aspect.

[0047] The model construction module is configured to construct an optical energy characteristic prediction model for predicting three optical energy characteristic parameters of organic photovoltaic material molecules, the three optical energy characteristic parameters including ionization energy, emission energy and absorption energy, and the optical energy characteristic prediction model is configured to perform ionization energy, emission energy and absorption energy prediction processing according to a model input of a molecular structure M, a temperature x, a pressure y and a wavelength sequence Z and output corresponding ionization energy u IE , emission energy sequence S EE and absorption energy sequence S AE .

[0048] The model training module is configured to construct a data set for model training, denoted as a corresponding first data set, by collecting data of molecular structures of multiple types of organic photovoltaic material molecules and ionization energy, emission energy curve and absorption energy curve of each molecule under different temperature and pressure conditions, and perform model training on the optical energy characteristic prediction model based on the first data set.

[0049] The model application module is configured to, after the training is completed, receive a first molecular structure, a first temperature, a first pressure and a first wavelength sequence input by a user, and input the first molecular structure, the first temperature, the first pressure and the first wavelength sequence as the corresponding molecular structure M, the temperature x, the pressure y and the wavelength sequence Z into the optical energy characteristic prediction model for processing and output the ionization energy u IE , the emission energy sequence S EE and the absorption energy sequence S AE output by the model processing this time as the corresponding first ionization energy, first emission energy sequence and first absorption energy sequence to the current user for feedback.

[0050] The third aspect of the embodiment of the application provides an electronic device, including a memory, a processor and a transceiver.

[0051] The processor is configured to be coupled with the memory, read and execute instructions in the memory to realize the method steps of the first aspect.

[0052] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transceiving.

[0053] The fourth aspect of the embodiment of the application provides a computer readable storage medium, the computer readable storage medium stores computer instructions, when the computer instructions are executed by a computer, the computer instructions make the computer execute the instructions of the method of the first aspect.

[0054] The embodiment of the present application provides a kind of organic photovoltaic material molecule optical energy characteristic prediction method, device, electronic equipment and computer readable storage medium.The above content can know, the optical energy characteristic prediction model for ionization energy, emission energy and absorption energy prediction is customized according to the model input molecular structure, temperature, pressure and wavelength sequence;And by the data acquisition of the ionization energy, emission energy curve and absorption energy curve of the molecular structure of multiple organic photovoltaic material molecules and each molecule under different temperature and pressure conditions, a first data set for model training is constructed;And based on the first data set, the optical energy characteristic prediction model is model trained;And after training, the molecular structure, temperature, pressure and wavelength sequence input by user are input into optical energy characteristic prediction model for processing, and ionization energy, emission energy sequence and absorption energy sequence output by this model processing are fed back to current user.The embodiment of the present application realizes one-step prediction of three optical energy characteristic parameters (ionization energy, emission energy, absorption energy) by a customized optical energy characteristic prediction model, and effectively reduces the prediction complexity, shortens the prediction time and improves the prediction efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0055] Figure 1 The optical energy characteristic prediction method of the organic photovoltaic material molecule provided by the embodiment of the present application is shown in the schematic diagram.

[0056] Figure 2 The module structure diagram of the optical energy characteristic prediction model provided by the embodiment of the present application is shown.

[0057] Figure 3 The module structure diagram of the optical energy characteristic prediction device of the organic photovoltaic material molecule provided by the embodiment of the present application is shown.

[0058] Figure 4 The structure schematic diagram of the electronic equipment provided by the embodiment of the present application is shown. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be described in further detail below with reference to the drawings.Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0060] The embodiment of the present application provides a kind of organic photovoltaic material molecule optical energy characteristic prediction method, as Figure 1 The optical energy characteristic prediction method of the organic photovoltaic material molecule provided by the embodiment of the present application is shown in the schematic diagram, and the method mainly includes the following steps:

[0061] Step 1: Construct an optical energy characteristic prediction model for predicting three optical energy characteristic parameters of organic photovoltaic material molecules.

[0062] Here, the three optical energy characteristic parameters of this invention include ionization energy, emission energy, and absorption energy.

[0063] The optical energy characteristic prediction model of this invention is used to predict ionization energy, emission energy, and absorption energy based on the molecular structure M, temperature x, pressure y, and wavelength sequence Z input to the model, and outputs the corresponding ionization energy u. IE Emission energy sequence S EE and absorption energy sequence S AE The molecular structure M comprises multiple atoms m. i ; atom m i Atomic parameters include atom type and atom three-dimensional coordinates; 1 ≤ atom index i ≤ N M N M The total number of atoms in the molecular structure M; the wavelength sequence Z consists of multiple wavelengths z. j Arranged in ascending order of wavelength; 1 ≤ wavelength index j ≤ N Z N Z The wavelength sequence Z is the sequence length; the ionization energy u IE Let S be a real number; the emission energy sequence is S. EE Composed of multiple emission energies EE,j Arranged in sequence; emission energy sequence S EE emission energy s EE,j With wavelength z of wavelength sequence Z j One-to-one correspondence; absorption energy sequence S AE Consists of multiple absorption energies s AE,j Arranged in sequence; Absorption energy sequence S AE Absorbed energy s AE,j With wavelength z of wavelength sequence Z j One-to-one correspondence.

[0064] like Figure 2 As shown in the module structure diagram of the optical energy characteristic prediction model provided in Embodiment 1 of the present invention, the first model input terminal of the optical energy characteristic prediction model is used to receive the molecular structure M input by the model, the second model input terminal is used to receive the temperature x, pressure y and wavelength sequence Z input by the model, and the first model output terminal is used to output the corresponding ionization energy u. IE The second model output is used to output the corresponding emission energy sequence S. EE The third model output is used to output the corresponding absorption energy sequence S. AE .

[0065] like Figure 2As shown, the model components of the optical energy property prediction model include: a first embedding coding module, a molecular structure coding module, a second embedding coding module, a vector space mapping module, a feature fusion module, an ionization energy prediction module, an emission energy prediction module, and an absorption energy prediction module.

[0066] like Figure 2 As shown, the connection relationships of the components of the optical energy characteristic prediction model are as follows: the input end of the first embedding coding module is connected to the input end of the first model, and the output end is connected to the input end of the molecular structure coding module; the output end of the molecular structure coding module is connected to the first input end of the feature fusion module; the input end of the second embedding coding module is connected to the input end of the second model, and the output end is connected to the input end of the vector space mapping module; the output end of the vector space mapping module is connected to the second input end of the feature fusion module; the first output end of the feature fusion module is connected to the input end of the ionization energy prediction module, the second output end is connected to the input end of the emission energy prediction module, and the third output end is connected to the input end of the absorption energy prediction module; the output end of the ionization energy prediction module is connected to the output end of the first model; the output end of the emission energy prediction module is connected to the output end of the second model; and the output end of the absorption energy prediction module is connected to the output end of the third model.

[0067] The functions of each component in the optical energy property prediction model are shown below.

[0068] 1) First Embedded Encoding Module:

[0069] The first embedding encoding module in this embodiment of the invention is used to encode each atom m of the molecular structure M according to the one-hot encoding rule of the atom type in the Uni-Mol model. i Perform one-hot encoding of the atom type to obtain the corresponding atom encoding vector, and then use the obtained N... M Each atom encoding vector forms a corresponding atom encoding tensor; and according to the pairwise feature initialization encoding rule of the Uni-Mol model, encoding is performed based on the coordinates of all three-dimensional atoms of the molecular structure M to obtain a tensor with shape N. M ×N M The resulting pairwise encoded tensors are then sent to the molecular structure encoding module.

[0070] The Uni-Mol model mentioned above is an encoder model for atom-level feature coding of a molecular structure, and the detailed model structure, model reasoning principle, atom type one-hot coding rule, pair feature initialization coding rule, and model pre-training scheme of the model are explicitly described in the published technical document A "Uni-Mol: A Universal 3D Molecular Representation Learning Framework". Here, the coding details of the first embedding coding module based on the atom type one-hot coding rule and the pair feature initialization coding rule of the Uni-Mol model will not be further repeated.

[0071] 2) Molecular structure coding module:

[0072] The molecular structure coding module of the embodiment of the application is implemented based on the Uni-Mol model. The molecular structure coding module is used to perform atom feature and pair feature extraction processing according to the input atom coding tensor and pair coding tensor to obtain a corresponding hidden feature tensor H and send it to the feature fusion module. The hidden feature tensor H is composed of N M hidden feature vectors h i , and the hidden feature vector h i corresponds to the atom m i one-to-one.

[0073] 3) Second embedding coding module:

[0074] The second embedding coding module of the embodiment of the application is used to preset a plurality of first Gaussian kernel functions, a plurality of second Gaussian kernel functions, and a plurality of third Gaussian kernel functions in the module. The second embedding coding module is also used to perform corresponding kernel function calculation on the temperature x of the model input according to each first Gaussian kernel function and take the calculated numerical result as the corresponding first result value; and a corresponding temperature coding vector is composed of all the first result values obtained; and perform corresponding kernel function calculation on the pressure y of the model input according to each second Gaussian kernel function and take the calculated numerical result as the corresponding second result value; and a corresponding pressure coding vector is composed of all the second result values obtained; and perform corresponding kernel function calculation on each wavelength z j of the wavelength sequence Z of the model input according to each third Gaussian kernel function and take the calculated numerical result as the corresponding third result value, and a corresponding wavelength coding vector is composed of all the third result values corresponding to each wavelength z j , and N Z wavelength coding vectors obtained are used to compose a corresponding wavelength coding tensor; and the temperature coding vector, the pressure coding vector, and the wavelength coding tensor are sent to the vector space mapping module.

[0075] Here, the wavelength encoding vector in the wavelength encoding tensor corresponds to the wavelength z of the wavelength sequence Z j One-to-one correspondence.

[0076] 4) Vector space mapping module:

[0077] The vector space mapping module of the embodiment of the application is used to preset three MLP models in the module. The vector space mapping module is also used to map the hidden feature vector h i to the target vector space based on the first MLP model, to obtain a corresponding temperature mapping vector; and map the pressure encoding vector to the target vector space based on the second MLP model, to obtain a corresponding pressure mapping vector; and map each wavelength encoding vector of the wavelength encoding tensor to the target vector space based on the third MLP model, to obtain a corresponding wavelength mapping vector, and the N Z wavelength mapping vectors form a corresponding wavelength mapping tensor; and the temperature mapping vector, the pressure mapping vector and the wavelength mapping tensor are sent to the feature fusion module.

[0078] 5) Feature fusion module:

[0079] The feature fusion module of the embodiment of the application is used to sequentially splice each hidden feature vector h i of the hidden feature tensor H and the temperature mapping vector and the pressure mapping vector to obtain a corresponding atom-temperature-pressure feature fusion vector; and the N M atom-temperature-pressure feature fusion vectors form a corresponding first fusion tensor; and each wavelength encoding vector and each first fusion vector are spliced to obtain a corresponding wavelength-atom-temperature-pressure feature fusion vector; and the N j corresponding wavelength-atom-temperature-pressure feature fusion vectors corresponding to each wavelength z M are spliced to obtain a corresponding wavelength-molecule-temperature-pressure feature fusion vector, and the N Z wavelength-molecule-temperature-pressure feature fusion vectors form a corresponding second fusion tensor; and the first fusion tensor is sent to the ionization energy prediction module, and the second fusion tensor is sent to the emission energy prediction module and the absorption energy prediction module.

[0080] 6) Ionization energy prediction module:

[0081] The ionization energy prediction module of the embodiment of the application is implemented based on the MLP model. The ionization energy prediction module is used to perform ionization energy prediction according to the input first fusion tensor to obtain a corresponding ionization energy u IE .

[0082] 7) Emission energy prediction module:

[0083] The emission energy prediction module of the embodiment of the present application is implemented based on an MLP model. The emission energy prediction module is used to perform emission energy sequence prediction according to an input second fusion tensor to obtain a corresponding emission energy sequence S EE .

[0084] 8) Absorption energy prediction module:

[0085] The absorption energy prediction module of the embodiment of the present application is implemented based on an MLP model. The absorption energy prediction module is used to perform absorption energy sequence prediction according to an input second fusion tensor to obtain a corresponding absorption energy sequence S AE .

[0086] Step 2, a data set for model training is constructed by collecting data of molecular structures of a plurality of types of organic photovoltaic material molecules and ionization energy curves, emission energy curves and absorption energy curves of each molecule under different temperature and pressure conditions, denoted as a corresponding first data set; and the optical energy characteristic prediction model is trained based on the first data set;

[0087] Specifically, step 21, a data set for model training is constructed by collecting data of molecular structures of a plurality of types of organic photovoltaic material molecules and ionization energy curves, emission energy curves and absorption energy curves of each molecule under different temperature and pressure conditions, denoted as a corresponding first data set;

[0088] The first data set includes a plurality of first data records; the first data record includes a first training molecular structure, a first training temperature, a first training pressure, a first training wavelength sequence, a first label ionization energy, a first label emission energy sequence and a first label absorption energy sequence; the data structures of the first training molecular structure, the first training wavelength sequence, the first label emission energy sequence and the first label absorption energy sequence are consistent with the corresponding molecular structure M, wavelength sequence Z, emission energy sequence S EE and absorption energy sequence S AE The sequence lengths of the first training wavelength sequence, the first label emission energy sequence and the first label absorption energy sequence are the same, and the first training wavelength of the first training wavelength sequence corresponds to the first label emission energy of the first label emission energy sequence one by one and corresponds to the first label absorption energy of the first label absorption energy sequence one by one;

[0089] Specifically, step 211, a plurality of first collection data are obtained by collecting data of molecular structures of a plurality of types of organic photovoltaic material molecules and ionization energy curves, emission energy curves and absorption energy curves of each molecule under different temperature and pressure conditions;

[0090] The first collection data include a first collection molecular structure, a first collection temperature, a first collection pressure, a first collection ionization energy, a first collection emission energy curve and a first collection absorption energy curve;

[0091] Step 212: Set a wavelength range as the first training wavelength range; extract the curve portion of the first acquisition emission energy curve of each first acquisition data within the first training wavelength range as the corresponding first extraction curve; extract the curve portion of the first acquisition absorption energy curve of each first acquisition data within the first training wavelength range as the corresponding second extraction curve; sample curve points of each first extraction curve based on a preset sampling wavelength interval Δλ to obtain multiple corresponding first sampling points to form a corresponding first sampling point sequence; and sample curve points of each second extraction curve based on the sampling wavelength interval Δλ to obtain multiple corresponding second sampling points to form a corresponding second sampling point sequence.

[0092] The sampling point parameters of the first sampling point consist of a sampling point wavelength and a corresponding sampling point emission energy; the sampling point parameters of the second sampling point consist of a sampling point wavelength and a corresponding sampling point absorption energy; the total number of sampling points in the first and second sampling point sequences corresponding to each first data acquisition is the same, that is, the first and second sampling points in the first and second sampling point sequences corresponding to each first data acquisition are one-to-one, and the sampling point wavelengths of each group of corresponding first and second sampling points are consistent.

[0093] Step 213: Perform a round of traversal on all the first acquired data; during this round of traversal, take the currently traversed first acquired data as the corresponding current acquired data; take the wavelength of each sampling point in the first sampling point sequence corresponding to the current acquired data as a corresponding first training wavelength, and sort all the obtained first training wavelengths in order to form the corresponding first training wavelength sequence; take the emission energy of each sampling point in the first sampling point sequence corresponding to the current acquired data as a corresponding first tag emission energy, and sort all the obtained first tag emission energies in order to form the corresponding first tag emission energy sequence; and take the second sampling point corresponding to the current acquired data... The absorption energy of each sampling point in the sequence is taken as a corresponding first tag absorption energy, and all the obtained first tag absorption energies are ordered to form a corresponding first tag absorption energy sequence; the first acquisition molecular structure, first acquisition temperature, first acquisition pressure, and first acquisition ionization energy of the currently acquired data are taken as the corresponding first training molecular structure, first training temperature, first training pressure, and first tag ionization energy; and a corresponding first data record is formed by the first training molecular structure, first training temperature, first training pressure, first training wavelength sequence, first tag ionization energy, first tag emission energy sequence, and first tag absorption energy sequence corresponding to the currently acquired data.

[0094] Step 214: After this round of traversal of all the first collected data is completed, the corresponding first dataset is composed of all the first data records obtained;

[0095] Step 22, and train the optical energy property prediction model based on the first dataset;

[0096] Specifically, it includes: Step 221, dividing the first dataset into two sub-datasets based on a preset first segmentation ratio, denoted as the corresponding first training set and first evaluation set;

[0097] Here, the first segmentation ratio is a pre-set ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio;

[0098] Step 222: Take the first data record of the first training set as the corresponding current training record;

[0099] Step 223: Record the first tag ionization energy, first tag emission energy sequence, and first tag absorption energy sequence currently recorded during training as the corresponding first tag ionization energy u. IE,tag First tag emission energy sequence S EE,tag and the first tag absorption energy sequence S AE,tag ;

[0100] Step 224: Input the first training molecular structure, first training temperature, first training pressure, and first training wavelength sequence recorded in the current training record as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z into the optical energy characteristic prediction model for processing, and then input the ionization energy u output by the model processing. IE Emission energy sequence S EE and absorption energy sequence S AE As the corresponding first predicted ionization energy u IE,pre First predicted emission energy sequence S EE,pre and the first predicted absorption energy sequence S AE,pre ;

[0101] Step 225, set the ionization energy of the first tag u IE,tag And the first predicted ionization energy u IE,pre Substitute the preset first model loss function L M1 ; and the first tag emits the energy sequence S EE,tag and the first predicted emission energy sequence S EE,pre Substitute the preset second model loss function L M2 ; and the first tag absorption energy sequence S AE,tag and the first predicted absorption energy sequence S AE,pre Substitute the preset third model loss function L M3 ; and the loss functions L of the first, second, and third models M1 L M2 L M3The summation constitutes the overall model loss function L. total =L M1 +L M2 +L M3 ; and based on the preset first model optimizer, it moves towards making the loss functions L of the first, second, and third models more efficient. M1 L M2 L M3 and the overall model loss function L total The model parameters of the optical energy characteristic prediction model are modulated in one round by ensuring that all values ​​reach their minimum values.

[0102] Here, the first model loss function L in this embodiment of the invention M1 Implemented based on L1 loss function, smoothed L1 loss function, or L2 loss function; second model loss function L M2 Implemented based on L1 or L2 loss function; third model loss function L M3 Implemented based on L1 loss function or L2 loss function; the first model optimizer in this embodiment of the invention includes at least Adam optimizer and SGD optimizer;

[0103] Step 226: Identify whether the current training record is the last first data record of the first training set. If yes, proceed to step 227; otherwise, take the next first data record of the first training set as the new current training record and return to step 223 to continue training.

[0104] Step 227: Perform a traversal of all first data records in the first evaluation set; during this traversal, the first data record currently being traversed is taken as the corresponding current evaluation record; and the first tag ionization energy, first tag emission energy sequence, and first tag absorption energy sequence of the current evaluation record are recorded as the corresponding first tag ionization energy u. IE,tag First tag emission energy sequence S EE,tag and the first tag absorption energy sequence S AE,tag The first training molecular structure, first training temperature, first training pressure, and first training wavelength sequence recorded in the current evaluation are used as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z, which are then input into the optical energy property prediction model for processing. The ionization energy u output by this model processing is then used. IE Emission energy sequence S EE and absorption energy sequence S AE As the corresponding first predicted ionization energy u IE,pre First predicted emission energy sequence S EE,pre and the first predicted absorption energy sequence S AE,pre ; and the ionization energy u corresponding to the first tag recorded in the current assessment. IE,tag And the first predicted ionization energy u IE,preA corresponding first prediction-tag pair is formed; and the first tag emission energy sequence S corresponding to the current evaluation record is formed. EE,tag and the first predicted emission energy sequence S EE,pre A corresponding second prediction-tag pair is formed; and the first tag absorption energy sequence S corresponding to the current evaluation record is used to form a second prediction-tag pair. AE,tag and the first predicted absorption energy sequence S AE,pre A corresponding third prediction-label pair is formed; at the end of this round of traversal, all the obtained first prediction-label pairs are input into a preset first model evaluation function to calculate the corresponding first evaluation value; all the obtained second prediction-label pairs are input into a preset second model evaluation function to calculate the corresponding second evaluation value; all the obtained third prediction-label pairs are input into a preset third model evaluation function to calculate the corresponding third evaluation value; and the obtained first, second and third evaluation values ​​are weighted and summed to calculate the corresponding fourth evaluation value.

[0105] Here, the first, second, and third model evaluation functions in this embodiment of the invention are each implemented based on the MAE function, the MSE function, or the RMSE function;

[0106] Step 228: Identify the first, second, third, and fourth evaluation values; if the first evaluation value does not meet the preset first evaluation value range, or the second evaluation value does not meet the preset second evaluation value range, or the third evaluation value does not meet the preset third evaluation value range, or the third evaluation value meets the preset third evaluation value range, or the fourth evaluation value does not meet the preset fourth evaluation value range, then return to step 222 to continue training; if the first evaluation value meets the first evaluation value range, the second evaluation value meets the second evaluation value range, the third evaluation value meets the third evaluation value range, and the fourth evaluation value meets the fourth evaluation value range, then stop training and confirm that the model training has ended.

[0107] Here, the first, second, third, and fourth evaluation value ranges in this embodiment of the invention are four preset numerical ranges.

[0108] Step 3: After training, receive the user's input of the first molecular structure, first temperature, first pressure, and first wavelength sequence; and input the first molecular structure, first temperature, first pressure, and first wavelength sequence as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z into the optical energy characteristic prediction model for processing. Then, output the ionization energy u from this model processing. IE Emission energy sequence S EE and absorption energy sequence S AE The corresponding first ionization energy, first emission energy sequence, and first absorption energy sequence are fed back to the current user.

[0109] Here, the data structures of the first molecular structure, the first wavelength sequence, the first emission energy sequence, and the first absorption energy sequence are respectively associated with the corresponding molecular structure M, wavelength sequence Z, and emission energy sequence S. EE and absorption energy sequence S AE Maintain consistency; the first wavelength sequence is formed by sorting multiple first wavelengths, the first emission energy sequence is formed by sorting multiple first emission energies, and the first absorption energy sequence is formed by sorting multiple first absorption energies; the first wavelength of the first wavelength sequence corresponds one-to-one with the first emission energy of the first emission energy sequence and one-to-one with the first absorption energy of the first absorption energy sequence.

[0110] Figure 3 This is a module structure diagram of an optical energy characteristic prediction device for organic photovoltaic material molecules provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 3 As shown, the device includes: a model building module 201, a model training module 202, and a model application module 203.

[0111] Model building module 201 is used to construct an optical energy characteristic prediction model for predicting three optical energy characteristic parameters of organic photovoltaic material molecules; the three optical energy characteristic parameters include ionization energy, emission energy, and absorption energy; the optical energy characteristic prediction model is used to predict the ionization energy, emission energy, and absorption energy based on the molecular structure M, temperature x, pressure y, and wavelength sequence Z input to the model, and outputs the corresponding ionization energy u. IE Emission energy sequence S EE and absorption energy sequence S AE .

[0112] The model training module 202 is used to construct a dataset for model training by collecting data on the molecular structure of various organic photovoltaic material molecules and the ionization energy, emission energy curves and absorption energy curves of each molecule under different temperature and pressure conditions. This dataset is denoted as the first dataset. The optical energy characteristic prediction model is then trained based on the first dataset.

[0113] The model application module 203 is used to receive the first molecular structure, first temperature, first pressure, and first wavelength sequence input by the user after training; and to input the first molecular structure, first temperature, first pressure, and first wavelength sequence as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z into the optical energy characteristic prediction model for processing, and to output the ionization energy u from this model processing. IE Emission energy sequence S EE and absorption energy sequence S AEThe corresponding first ionization energy, first emission energy sequence, and first absorption energy sequence are fed back to the current user.

[0114] The optical energy characteristic prediction device for organic photovoltaic material molecules provided in this embodiment of the invention can execute the method steps in the above method embodiment. Its implementation principle and technical effect are similar, and will not be repeated here.

[0115] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing elements; they can be fully implemented in hardware; or some modules can be implemented by processing elements calling software, while others are implemented in hardware. For example, the model building module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0116] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-on-a-Chip (SOC).

[0117] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0118] Figure 4 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 4 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.

[0119] exist Figure 4The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.

[0120] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0121] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.

[0122] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for predicting the optical energy properties of organic photovoltaic material molecules. As described above, this invention customizes an optical energy property prediction model for predicting ionization energy, emission energy, and absorption energy based on the molecular structure, temperature, pressure, and wavelength sequence input to the model. A first dataset for model training is constructed by collecting data on the molecular structures of various organic photovoltaic material molecules and their ionization energy, emission energy curves, and absorption energy curves under different temperature and pressure conditions. The optical energy property prediction model is then trained based on this first dataset. After training, the user-input molecular structure, temperature, pressure, and wavelength sequence are input into the optical energy property prediction model for processing, and the ionization energy, emission energy, and absorption energy sequences output by the model are fed back to the user. This invention achieves one-step prediction of three optical energy property parameters (ionization energy, emission energy, and absorption energy) through a customized optical energy property prediction model, effectively reducing prediction complexity, shortening prediction time, and improving prediction efficiency.

[0123] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0124] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for predicting the optical energy properties of organic photovoltaic material molecules, characterized in that, The method includes: An optical energy characteristic prediction model is constructed to predict three optical energy characteristic parameters of organic photovoltaic material molecules: ionization energy, emission energy, and absorption energy. The optical energy characteristic prediction model is used to predict the ionization energy, emission energy, and absorption energy based on the molecular structure M, temperature x, pressure y, and wavelength sequence Z input to the model, and outputs the corresponding ionization energy u. IE Emission energy sequence S EE and absorption energy sequence S AE ; A dataset for model training is constructed by collecting data on the molecular structure of various organic photovoltaic materials and the ionization energy, emission energy curves, and absorption energy curves of each molecule under different temperature and pressure conditions. This dataset is denoted as the first dataset. The optical energy characteristic prediction model is then trained based on the first dataset. After training, the system receives the user's input of the first molecular structure, first temperature, first pressure, and first wavelength sequence; and inputs the first molecular structure, first temperature, first pressure, and first wavelength sequence as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z into the optical energy characteristic prediction model for processing. The ionization energy u output by this model processing is then used. IE The emission energy sequence S EE and the absorption energy sequence S AE The corresponding first ionization energy, first emission energy sequence, and first absorption energy sequence are fed back to the current user.

2. The method for predicting the optical energy properties of organic photovoltaic material molecules according to claim 1, characterized in that, The molecular structure M comprises multiple atoms m i The atom m i Atomic parameters include atom type and atom three-dimensional coordinates; 1 ≤ atom index i ≤ N M N M The total number of atoms in the molecular structure M; The wavelength sequence Z consists of multiple wavelengths z j Arranged in ascending order of wavelength; 1 ≤ wavelength index j ≤ N Z N Z The sequence length of the wavelength sequence Z; The ionization energy u IE Let it be a real number; The emission energy sequence S EE Composed of multiple emission energies EE,j The emission energy sequence S is formed by sequential sorting; EE The emission energy s EE,j The wavelength z of the wavelength sequence Z j One-to-one correspondence; The absorption energy sequence S AE Consists of multiple absorption energies s AE,j The absorption energy sequence S is formed by sequentially sorting the data; AE The absorbed energy s AE,j The wavelength z of the wavelength sequence Z j One-to-one correspondence; The first dataset includes multiple first data records; the first data record includes a first training molecular structure, a first training temperature, a first training pressure, a first training wavelength sequence, a first tag ionization energy, a first tag emission energy sequence, and a first tag absorption energy sequence; The data structures of the first training molecular structure, the first training wavelength sequence, the first tag emission energy sequence, and the first tag absorption energy sequence are respectively associated with the corresponding molecular structure M, the wavelength sequence Z, and the emission energy sequence S. EE and the absorption energy sequence S AE The first training wavelength sequence, the first tag emission energy sequence, and the first tag absorption energy sequence have the same sequence length. The first training wavelength of the first training wavelength sequence corresponds one-to-one with the first tag emission energy of the first tag emission energy sequence and one-to-one with the first tag absorption energy of the first tag absorption energy sequence.

3. The method for predicting the optical energy properties of organic photovoltaic material molecules according to claim 2, characterized in that, The first model input terminal of the optical energy characteristic prediction model is used to receive the molecular structure M as input to the model; the second model input terminal is used to receive the temperature x, the pressure y, and the wavelength sequence Z as input to the model; and the first model output terminal is used to output the corresponding ionization energy u. IE The second model output terminal is used to output the corresponding emission energy sequence S. EE The third model output terminal is used to output the corresponding absorption energy sequence S. AE ; The optical energy characteristic prediction model includes a first embedding coding module, a molecular structure coding module, a second embedding coding module, a vector space mapping module, a feature fusion module, an ionization energy prediction module, an emission energy prediction module, and an absorption energy prediction module. The input terminal of the first embedding coding module is connected to the input terminal of the first model, and its output terminal is connected to the input terminal of the molecular structure coding module; the output terminal of the molecular structure coding module is connected to the first input terminal of the feature fusion module; the input terminal of the second embedding coding module is connected to the input terminal of the second model, and its output terminal is connected to the input terminal of the vector space mapping module; the output terminal of the vector space mapping module is connected to the second input terminal of the feature fusion module; the first output terminal of the feature fusion module is connected to the input terminal of the ionization energy prediction module, the second output terminal is connected to the input terminal of the emission energy prediction module, and the third output terminal is connected to the input terminal of the absorption energy prediction module; the output terminal of the ionization energy prediction module is connected to the output terminal of the first model; the output terminal of the emission energy prediction module is connected to the output terminal of the second model; and the output terminal of the absorption energy prediction module is connected to the output terminal of the third model. The first embedding encoding module is used to encode each atom m of the molecular structure M according to the one-hot encoding rule of the atom type of the Uni-Mol model. i Perform one-hot encoding of the atom type to obtain the corresponding atom encoding vector, and then use the obtained N... M The atomic encoding vectors form a corresponding atomic encoding tensor; and according to the pairwise feature initialization encoding rule of the Uni-Mol model, the encoding is performed based on all three-dimensional atomic coordinates of the molecular structure M to obtain a tensor with shape N. M ×N M The pairwise encoded tensor; and the obtained atomic encoded tensor and the pairwise encoded tensor are sent to the molecular structure encoding module; The molecular structure encoding module is implemented based on the Uni-Mol model; the molecular structure encoding module is used to perform atomic feature and pairwise feature extraction processing on the input atomic encoding tensor and the pairwise encoding tensor to obtain the corresponding latent feature tensor H and send it to the feature fusion module; the latent feature tensor H is composed of N M Hidden feature vectors h i Composition, the hidden feature vector h i With the atom m i One-to-one correspondence; The second embedding encoding module is used to preset multiple first Gaussian kernel functions, multiple second Gaussian kernel functions, and multiple third Gaussian kernel functions within the module; The second embedding encoding module is further configured to perform corresponding kernel function calculations on the temperature x input to the model according to each of the first Gaussian kernel functions and use the calculated numerical results as the corresponding first result values; and to form a corresponding temperature encoding vector by combining all the obtained first result values; The pressure y input to the model is calculated using the corresponding kernel function according to each of the second Gaussian kernel functions, and the calculated numerical results are used as the corresponding second result values; a corresponding pressure encoding vector is formed by all the obtained second result values; and each wavelength z of the wavelength sequence Z input to the model is calculated using the corresponding third Gaussian kernel function. j Perform the corresponding kernel function calculation and use the calculated numerical result as the corresponding third result value, and then use each wavelength z j All the corresponding third result values ​​form a corresponding wavelength encoding vector, and the obtained N Z The wavelength encoding vectors form a corresponding wavelength encoding tensor; and the temperature encoding vector, the pressure encoding vector, and the wavelength encoding tensor are sent to the vector space mapping module; the wavelength encoding vector and the wavelength z j One-to-one correspondence; The vector space mapping module is used to pre-set three MLP models within the module; The vector space mapping module is further used to map the latent feature vector h. i The corresponding vector space is used as the target vector space; and based on the first built-in MLP model, the temperature encoding vector is mapped to the target vector space to obtain a corresponding temperature mapping vector; and based on the second built-in MLP model, the pressure encoding vector is mapped to the target vector space to obtain a corresponding pressure mapping vector; and based on the third built-in MLP model, each wavelength encoding vector of the wavelength encoding tensor is mapped to the target vector space to obtain a corresponding wavelength mapping vector, and the obtained N... Z The wavelength mapping vectors are combined to form a corresponding wavelength mapping tensor; and the temperature mapping vector, the pressure mapping vector, and the wavelength mapping tensor are sent to the feature fusion module; The feature fusion module is used to fuse the latent feature vectors h of the latent feature tensor H. i The temperature mapping vector and the pressure mapping vector are sequentially concatenated to obtain a corresponding atom-temperature-pressure feature fusion vector; and the obtained N M The atom-temperature-pressure feature fusion vectors are composed of a corresponding first fusion tensor; and a corresponding wavelength-atom-temperature-pressure feature fusion vector is obtained by concatenating each wavelength encoding vector with each of the first fusion vectors; and a wavelength-atom-temperature-pressure feature fusion vector is obtained by concatenating each of the wavelength encoding vectors .... j The corresponding N M The wavelength-atom-temperature-pressure feature fusion vectors are concatenated to form a corresponding wavelength-molecule-temperature-pressure feature fusion vector, and the resulting N... Z The wavelength-molecule-temperature-pressure feature fusion vectors are used to form a corresponding second fusion tensor; and the first fusion tensor is sent to the ionization energy prediction module, and the second fusion tensor is sent to the emission energy prediction module and the absorption energy prediction module; The ionization energy prediction module is implemented based on an MLP model; the ionization energy prediction module is used to predict the corresponding ionization energy u based on the input first fusion tensor. IE ; The emission energy prediction module is implemented based on an MLP model; the emission energy prediction module is used to predict the corresponding emission energy sequence S based on the input second fusion tensor. EE ; The absorption energy prediction module is implemented based on an MLP model; the absorption energy prediction module is used to predict the corresponding absorption energy sequence S based on the input second fusion tensor. AE .

4. The method for predicting the optical energy properties of organic photovoltaic material molecules according to claim 2, characterized in that, The process involves collecting data on the molecular structures of various organic photovoltaic materials and the ionization energy, emission energy curves, and absorption energy curves of each molecule under different temperature and pressure conditions to construct a dataset for model training, denoted as the first dataset. This dataset specifically includes: Step 41: Collect data on the molecular structure of various organic photovoltaic materials and the ionization energy, emission energy curves and absorption energy curves of each molecule under different temperature and pressure conditions to obtain multiple corresponding first data collections. The first collected data includes the first collected molecular structure, the first collected temperature, the first collected pressure, the first collected ionization energy, the first collected emission energy curve, and the first collected absorption energy curve. Step 42: Set a wavelength range as the first training wavelength range; extract the curve portion of the first acquisition emission energy curve of each of the first acquisition data within the first training wavelength range as the corresponding first extraction curve; extract the curve portion of the first acquisition absorption energy curve of each of the first acquisition data within the first training wavelength range as the corresponding second extraction curve; and perform curve point sampling on each of the first extraction curves based on a preset sampling wavelength interval Δλ to obtain multiple corresponding first sampling points forming a corresponding first sampling point sequence; and perform curve point sampling on each of the second extraction curves based on the sampling wavelength interval Δλ to obtain multiple corresponding second sampling points forming a corresponding second sampling point sequence. The sampling point parameters of the first sampling point consist of a sampling point wavelength and a corresponding sampling point emission energy; the sampling point parameters of the second sampling point consist of a sampling point wavelength and a corresponding sampling point absorption energy; the total number of sampling points in the first and second sampling point sequences corresponding to each first acquired data is the same, that is, the first and second sampling points in the first and second sampling point sequences corresponding to each first acquired data correspond one-to-one, and the sampling point wavelengths of each group of corresponding first and second sampling points are consistent. Step 43: Perform a traversal of all the first acquired data; during this traversal, the currently traversed first acquired data is taken as the corresponding current acquired data; the wavelengths of each sampling point in the first sampling point sequence corresponding to the current acquired data are taken as a corresponding first training wavelength, and all the obtained first training wavelengths are ordered to form a corresponding first training wavelength sequence; the emission energy of each sampling point in the first sampling point sequence corresponding to the current acquired data is taken as a corresponding first tag emission energy, and all the obtained first tag emission energy are ordered to form a corresponding first tag emission energy sequence; and the emission energy of each sampling point in the second sampling point sequence corresponding to the current acquired data is taken as a corresponding first tag emission energy. The point absorption energy is used as a corresponding first tag absorption energy, and all the obtained first tag absorption energies are ordered to form a corresponding first tag absorption energy sequence; the first acquisition molecular structure, first acquisition temperature, first acquisition pressure, and first acquisition ionization energy of the currently acquired data are used as the corresponding first training molecular structure, first training temperature, first training pressure, and first tag ionization energy; and the first training molecular structure, first training temperature, first training pressure, first training wavelength sequence, first tag ionization energy, first tag emission energy sequence, and first tag absorption energy sequence corresponding to the currently acquired data are used to form a corresponding first data record; Step 44: After this round of traversal of all the first collected data is completed, the first dataset is composed of all the first data records obtained.

5. The method for predicting the optical energy properties of organic photovoltaic material molecules according to claim 2, characterized in that, The step of training the optical energy characteristic prediction model based on the first dataset specifically includes: Step 51: Based on a preset first segmentation ratio, the first dataset is divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 52: Take the first data record of the first training set as the corresponding current training record; Step 53: Record the first tag ionization energy, the first tag emission energy sequence, and the first tag absorption energy sequence recorded in the current training as the corresponding first tag ionization energy u. IE,tag First tag emission energy sequence S EE,tag and the first tag absorption energy sequence S AE,tag ; Step 54: Input the first training molecular structure, first training temperature, first training pressure, and first training wavelength sequence recorded in the current training as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z into the optical energy characteristic prediction model for processing, and output the ionization energy u from this model processing. IE The emission energy sequence S EE and the absorption energy sequence S AE As the corresponding first predicted ionization energy u IE,pre First predicted emission energy sequence S EE,pre and the first predicted absorption energy sequence S AE,pre ; Step 55, the ionization energy u of the first tag is... IE,tag and the first predicted ionization energy u IE,pre Substitute the preset first model loss function L M1 ; and the first tag emits the energy sequence S EE,tag and the first predicted emission energy sequence S EE,pre Substitute the preset second model loss function L M2 ; and the first tag absorbs the energy sequence S AE,tag and the first predicted absorption energy sequence S AE,pre Substitute the preset third model loss function L M3 ; and the loss functions L of the first, second, and third models M1 L M2 L M3 The summation constitutes the overall model loss function L. total =L M1 +L M2 +L M3 ; and based on a preset first model optimizer, it moves towards making the loss functions L of the first, second, and third models more efficient. M1 L M2 L M3 and the overall model loss function L total The model parameters of the optical energy characteristic prediction model are modulated in one round in a manner that minimizes all values. Wherein, the first model loss function L M1 Implemented based on L1 loss function, smoothed L1 loss function, or L2 loss function; the second model loss function L M2 Implemented based on L1 or L2 loss function; the third model loss function L M3 Implemented based on L1 or L2 loss function; Step 56: Identify whether the current training record is the last first data record of the first training set. If yes, proceed to step 57. Otherwise, take the next first data record of the first training set as the new current training record and return to step 53 to continue training. Step 57: Perform a traversal of all the first data records in the first evaluation set; and during this traversal, take the currently traversed first data record as the corresponding current evaluation record; and record the first tag ionization energy, the first tag emission energy sequence, and the first tag absorption energy sequence of the current evaluation record as the corresponding first tag ionization energy u. IE,tag First tag emission energy sequence S EE,tag and the first tag absorption energy sequence S AE,tag The first training molecular structure, first training temperature, first training pressure, and first training wavelength sequence recorded in the current evaluation are used as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z, and input into the optical energy characteristic prediction model for processing. The ionization energy u output by this model processing is then used as the model's input. IE The emission energy sequence S EE and the absorption energy sequence S AE As the corresponding first predicted ionization energy u IE,pre First predicted emission energy sequence S EE,pre and the first predicted absorption energy sequence S AE,pre ; and the first tag ionization energy u corresponding to the current evaluation record. IE,tag and the first predicted ionization energy u IE,pre A corresponding first prediction-tag pair is formed; and the first tag emission energy sequence S corresponding to the current evaluation record is formed. EE,tag and the first predicted emission energy sequence S EE,pre A corresponding second prediction-tag pair is formed; and the first tag absorption energy sequence S corresponding to the current evaluation record is used to form a second prediction-tag pair. AE,tag and the first predicted absorption energy sequence S AE,pre A corresponding third prediction-label pair is formed; and at the end of this round of traversal, all the obtained first prediction-label pairs are input into a preset first model evaluation function to calculate the corresponding first evaluation value; all the obtained second prediction-label pairs are input into a preset second model evaluation function to calculate the corresponding second evaluation value; all the obtained third prediction-label pairs are input into a preset third model evaluation function to calculate the corresponding third evaluation value; and the obtained first, second and third evaluation values ​​are weighted and summed to calculate the corresponding fourth evaluation value. Among them, the first, second and third model evaluation functions are each implemented based on the MAE function, MSE function or RMSE function; Step 58: Identify the first, second, third, and fourth evaluation values; if the first evaluation value does not meet the preset first evaluation value range, or the second evaluation value does not meet the preset second evaluation value range, or the third evaluation value does not meet the preset third evaluation value range, or the third evaluation value meets the preset third evaluation value range, or the fourth evaluation value does not meet the preset fourth evaluation value range, then return to step 52 to continue training; if the first evaluation value meets the first evaluation value range, the second evaluation value meets the second evaluation value range, the third evaluation value meets the third evaluation value range, and the fourth evaluation value meets the fourth evaluation value range, then stop training and confirm that the model training has ended.

6. An apparatus for performing the method for predicting the optical energy properties of organic photovoltaic material molecules according to any one of claims 1-5, characterized in that, The device includes: a model building module, a model training module, and a model application module; The model building module is used to construct an optical energy characteristic prediction model for predicting three optical energy characteristic parameters of organic photovoltaic material molecules; the three optical energy characteristic parameters include ionization energy, emission energy, and absorption energy; the optical energy characteristic prediction model is used to predict the ionization energy, emission energy, and absorption energy based on the molecular structure M, temperature x, pressure y, and wavelength sequence Z input to the model, and outputs the corresponding ionization energy u. IE Emission energy sequence S EE and absorption energy sequence S AE ; The model training module is used to construct a dataset for model training by collecting data on the molecular structure of various organic photovoltaic material molecules and the ionization energy, emission energy curves and absorption energy curves of each molecule under different temperature and pressure conditions. This dataset is denoted as the first dataset. The optical energy characteristic prediction model is then trained based on the first dataset. The model application module is used to receive the first molecular structure, first temperature, first pressure, and first wavelength sequence input by the user after training; and to input the first molecular structure, first temperature, first pressure, and first wavelength sequence as the corresponding molecular structure M, temperature x, pressure y, and wavelength sequence Z into the optical energy characteristic prediction model for processing, and output the ionization energy u from this model processing. IE The emission energy sequence S EE and the absorption energy sequence S AE The corresponding first ionization energy, first emission energy sequence, and first absorption energy sequence are fed back to the current user.

7. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-5; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method for predicting energy level of organic solar cell material

    CN118412067A

  • Method for predicting bulk heterojunction organic photovoltaics performance

    WO2024043565A1