Processing method and device of a prediction model for the maximum absorption peak of an organic molecule

By constructing a deep learning model based on MLP and Uni-Mol models, and combining experiments and big data collection to train a model for predicting the maximum absorption peak of organic molecules, the problems of long processing time and poor real-time performance in existing technologies are solved, and fast and personalized prediction of the maximum absorption peak is achieved.

CN120183561BActive Publication Date: 2025-12-12BEIJING DP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510287895.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-12-12
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Existing technologies suffer from long processing times and poor real-time performance when identifying the maximum absorption peak of organic molecules, and cannot effectively combine environmental parameters for personalized prediction.

Method used

We constructed deep learning models based on MLP and Uni-Mol models, collected organic molecular structure data through experiments, simulations, and big data acquisition, trained the models to predict the maximum absorption peak, and made accurate predictions under different environmental conditions.

Benefits of technology

It shortens the processing cycle of the maximum absorption peak identification task, improves the real-time performance of task processing, and enhances the personalized prediction level of the maximum absorption peak parameter by combining environmental parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183561B_ABST
    Figure CN120183561B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to a kind of organic molecule maximum absorption peak prediction model processing method and device, the method includes: with MLP model and Uni-Mol model as core component, construct a maximum absorption peak prediction model for the maximum absorbance and maximum absorption wavelength of maximum absorption peak are predicted according to the molecular structure, environmental temperature, environmental illumination intensity and environmental PH value of model input;And by data acquisition, construct first data set, and based on first data set, train maximum absorption peak prediction model;And after model training ends, receive the first molecular structure and first environmental parameter set input by user, and utilize maximum absorption peak prediction model to predict the maximum absorption peak information of first molecular structure under all first environmental parameters to obtain corresponding first prediction vector set Feedback to current user.The present application can improve the task processing real-time of maximum absorption peak identification task, and can improve individualized prediction level.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a processing method and device of an organic molecule maximum absorption peak prediction model. BACKGROUND

[0002] Organic photovoltaics (OPV) material refers to an organic material that converts solar energy or other light energy into electrical energy through photovoltaic effect, and the organic photovoltaic material is composed of organic molecules. The absorption spectrum curve is an important optical property curve of such organic molecules, and the vertical coordinate is the absorbance and the horizontal coordinate is the light wavelength in the conventional case. The maximum absorption peak is the maximum peak position on the absorption spectrum curve, and the vertical coordinate of the point is the maximum absorbance and the horizontal coordinate is the maximum absorption wavelength. At present, in the case of known organic molecule structure, the maximum absorption peak is identified by either experimental method or simulation calculation method. The experimental method refers to obtaining the absorption spectrum curve through a series of experimental methods, and then further extracting the maximum absorption peak from the curve. The simulation calculation method refers to obtaining the maximum absorption peak through a series of quantum chemical calculations (such as density functional theory calculation, time-dependent density functional theory calculation, quantum mechanics calculation, molecular mechanics calculation, etc.). However, according to practical experience, both of the two conventional processing methods have the problems of long processing period and poor real-time performance. SUMMARY

[0003] The purpose of the present application is to provide a processing method and device of an organic molecule maximum absorption peak prediction model, electronic equipment and computer readable storage medium, aiming at the defects of the prior art. The present application takes MLP model and Uni-Mol model as the core components to construct a deep learning model for predicting the maximum absorption peak parameters (maximum absorbance, maximum absorption wavelength) of the current molecule according to the model input of the molecular structure, environmental temperature, environmental light intensity and environmental PH value, which is called maximum absorption peak prediction model. And through experiments, simulation calculations and big data collection methods, the molecular structure data of a plurality of types of organic molecules of organic photovoltaic materials and the maximum absorption peak information of each organic molecule under different environmental conditions (temperature, light intensity, PH value) are collected to obtain a first data set for model training. And based on the first data set, the maximum absorption peak prediction model is trained. After the model training is completed, the maximum absorption peak prediction model is used to predict the maximum absorption peak information of the molecular structure input by the user under the user-specified multiple sets of environmental parameters (temperature, light intensity, PH value). The technical scheme of the present application can shorten the task processing period of the maximum absorption peak identification task and improve the task processing real-time performance based on the maximum absorption peak prediction model, and can also improve the individualized prediction level of the maximum absorption peak parameters (maximum absorbance, maximum absorption wavelength) in combination with the environmental parameters.

[0004] To achieve the above object, the first aspect of the embodiment of the present application provides a processing method of an organic molecule maximum absorption peak prediction model, which comprises:

[0005] a deep learning model is constructed with the MLP model and the Uni-Mol model as core components, denoted as a corresponding maximum absorption peak prediction model; the maximum absorption peak prediction model is used to predict the maximum absorbance and the maximum absorption wavelength of the maximum absorption peak according to the model input of the molecular structure X, the environmental temperature, the environmental light intensity and the environmental PH value and output the corresponding prediction vector Y; the prediction vector Y comprises a predicted absorbance and a predicted wavelength, which correspond to the maximum absorbance and the maximum absorption wavelength, respectively;

[0006] Through experiments, simulation calculations and big data collection methods, the molecular structure data of multiple types of organic molecules of organic photovoltaic materials and the maximum absorption peak information of each organic molecule under different environmental conditions are collected to form a corresponding first data set based on the collected data;

[0007] The maximum absorption peak prediction model is trained based on the first data set;

[0008] After the model training is completed, a first molecular structure and a first environmental parameter set input by a user are received; and the maximum absorption peak prediction model is used to predict the maximum absorption peak information of the first molecular structure under all first environmental parameters in the first environmental parameter set to obtain a corresponding first prediction vector set to feedback to the current user; the first environmental parameter set comprises multiple first environmental parameters; the first environmental parameters comprise temperature, light intensity and PH value; the first prediction vector set comprises multiple first prediction vectors; the first prediction vector is composed of a first absorbance and a first wavelength, and corresponds to the first environmental parameter one by one.

[0009] Preferably, the molecular structure X comprises multiple atoms a i , 1≤atomic index i≤A X , A X is the total number of atoms of the molecular structure X; the atomic parameters of each atom a i include the atomic type t i and the three-dimensional atomic coordinates c i (x, y, z);

[0010] The first data set includes a plurality of first data records; the first data record includes a first training molecule structure, a first training environment parameter, and a first maximum absorption peak label; the first training molecule structure includes a plurality of first atoms, and the atomic parameters of each first atom include an atomic type and three-dimensional atomic coordinates; the first training environment parameter includes a first temperature, a first light intensity, and a first pH value; and the first maximum absorption peak label includes a first maximum absorbance label and a first maximum absorption wavelength label.

[0011] Preferably, the maximum absorption peak prediction model includes a preprocessing module, a first MLP model, a Uni-Mol model, a feature fusion module, a second MLP model, a third MLP model, and a prediction output module.

[0012] The input end of the preprocessing module is connected with the model input end of the maximum absorption peak prediction model, and the first and second output ends are respectively connected with the input ends of the first MLP model and the Uni-Mol model; the output ends of the first MLP model and the Uni-Mol model are respectively connected with the first and second input ends of the feature fusion module; the output end of the feature fusion module is connected with the input ends of the second MLP model and the third MLP model; the output ends of the second MLP model and the third MLP model are respectively connected with the first and second input ends of the prediction output module; and the output end of the prediction output module is connected with the model output end of the maximum absorption peak prediction model.

[0013] The preprocessing module is configured to embed and encode the environment temperature, the environment light intensity, and the environment pH value of the model input according to a preset temperature, light intensity, and pH value embedding and encoding mode to obtain corresponding temperature encoding, light intensity encoding, and pH value encoding to form a corresponding first environment encoding vector; and based on an atomic type one-hot encoding rule of the Uni-Mol model, according to all atomic types of the molecule structure X, to obtain the one-hot encoding of each atomic type of the molecule structure X. iThe atomic type one-hot encoding is performed to obtain a corresponding first atomic one-hot encoding vector, and all the obtained first atomic one-hot encoding vectors form a corresponding first atomic encoding tensor; and an encoding rule is initialized based on an atomic pair feature of the Uni-Mol model, and a corresponding first atomic pair encoding tensor is obtained by initializing the pair atomic feature according to all three-dimensional atomic coordinates of the molecular structure X for each pair of atoms; and the first environment encoding vector is sent to the first MLP model; and the first atomic encoding tensor and the first atomic pair encoding tensor are sent to the Uni-Mol model; the first environment encoding vector is a real number vector with a length of 3, which is composed of the temperature encoding, the illumination intensity encoding and the PH value encoding; the first atomic one-hot encoding vector is a one-dimensional one-hot encoding vector, and the vector length is denoted as D1, which matches the total number of atomic types of the molecular structure X; the one-hot encoding corresponding to the atomic type t of the current atom a in the first atomic one-hot encoding vector is set to 1, and the remaining D1-1 one-hot encodings are set to 0; the tensor shape of the first atomic encoding tensor is D1×A; the tensor shape of the first atomic pair encoding tensor is D2×AxA, and D2 is a preset atomic pair encoding feature dimension. i i X X X

[0014] The first MLP model is sequentially connected by multiple layers of fully connected layers; the first MLP model is used for environment feature extraction processing according to the input first environment encoding vector to obtain a corresponding first environment feature tensor, which is sent to the feature fusion module; the tensor shape of the first environment feature tensor is D3×3, and D3 is a preset environment feature dimension; the first environment feature tensor is specifically composed of three environment feature vectors with a vector length of D3, which are a first temperature feature vector, a first illumination intensity feature vector and a first PH value feature vector.

[0015] The Uni-Mol model is used for atomic and atomic pair feature extraction processing according to the input first atomic encoding tensor and the first atomic pair encoding tensor to obtain a corresponding first atomic feature tensor and a first atomic pair feature tensor; and the first atomic feature tensor and the first atomic pair feature tensor are fused to obtain a corresponding first molecular feature tensor, which is sent to the feature fusion module; the tensor shape of the first atomic feature tensor is D4×A, and D4 is a preset atomic feature dimension; the first atomic feature tensor is specifically composed of A first sub-feature vectors with a vector length of D4; the tensor shape of the first atomic pair feature tensor is D5×AxA, and D5 is a preset atomic pair feature dimension. X X X X ​​​​​​​​D5 is a preset atomic pair feature dimension; the first atomic pair feature tensor is specifically composed of A X first sub-feature tensors with a shape of D5 x A X , each of the first sub-feature tensors is specifically composed of A X second sub-feature vectors with a length of D5; the tensor shape of the first molecular feature tensor is (D4+D5 x A X ) x A X ; the first molecular feature tensor is specifically composed of A X first atomic fusion vectors with a length of (D4+D5 x A X ); each of the first atomic fusion vectors is sequentially spliced from a corresponding first sub-feature vector and A X second sub-feature vectors;

[0016] The feature fusion module is configured to sequentially splice each of the first atomic fusion vectors of the first molecular feature tensor and three environment feature vectors of the first environment feature tensor to form a corresponding second atomic fusion vector, and to form a corresponding first fusion feature tensor from all the obtained second atomic fusion vectors and send the first fusion feature tensor to the second and third MLP models; the tensor shape of the first fusion feature tensor is (D3 x 3+D4+D5 x A X ) x A X ;

[0017] The second MLP model is sequentially connected by multiple fully connected layers; the second MLP model is configured to perform maximum absorbance prediction processing according to the input first fusion feature tensor to obtain a corresponding predicted absorbance and send the predicted absorbance to the prediction output module;

[0018] The third MLP model is sequentially connected by multiple fully connected layers; the third MLP model is configured to perform maximum absorption wavelength prediction processing according to the input first fusion feature tensor to obtain a corresponding predicted wavelength and send the predicted wavelength to the prediction output module;

[0019] The prediction output module is configured to form a corresponding prediction vector Y from the predicted absorbance and the predicted wavelength and output the prediction vector Y.

[0020] Preferably, the model training of the maximum absorption peak prediction model based on the first data set specifically includes:

[0021] Step 41, based on a preset first segmentation ratio, the first data set is segmented to obtain two sub-data sets, denoted as a corresponding first training set and a first evaluation set;

[0022] The first training set and the first evaluation set are composed of a plurality of first data records; the proportion of the total number of records of the first training set and the first evaluation set meets the first split proportion;

[0023] Step 42, taking the first first data record of the first training set as a corresponding current training record;

[0024] Step 43, taking the first training molecular structure of the current training record and the first temperature, the first light intensity and the first PH value of the first training environment parameter as the corresponding molecular structure X, the environmental temperature, the environmental light intensity and the environmental PH value, inputting the maximum absorption peak prediction model for prediction processing to obtain the corresponding prediction vector Y;

[0025] Step 44, extracting the first maximum absorption peak label of the current training record as a corresponding current label vector, taking the current prediction vector Y as a corresponding current prediction vector; and taking the current prediction vector and the current label vector into a preset first model loss function to obtain a corresponding first loss value;

[0026] The first model loss function is realized based on an L1 loss function or an L2 loss function;

[0027] Step 45, identifying whether the first loss value meets a preset first loss value range; if the first loss value meets the first loss value range, identifying whether the current training record is the last first data record of the first training set, if yes, turning to step 46, if not, taking the next first data record of the first training set as a new current training record and returning to step 43 for continuous training; if the first loss value does not meet the first loss value range, based on a preset first model optimizer, a round of modulation is performed on the model parameters of the first, second and third MLP models and the Uni-Mol model of the maximum absorption peak prediction model in the direction of minimizing the first model loss function, and returning to step 43 for continuous training at the end of the round of modulation;

[0028] The first model optimizer at least includes an Adam optimizer and an SGD optimizer;

[0029] Step 46, a round of traversal is performed on all the first data records of the first evaluation set; and during the round of traversal, the first data record currently traversed is taken as a corresponding current evaluation record; the first training molecular structure of the current evaluation record and the first temperature, the first light intensity and the first PH value of the first training environment parameter are taken as the molecular structure X, the environment temperature, the environment light intensity and the environment PH value respectively, and are input into the maximum absorption peak prediction model for prediction processing to obtain a corresponding prediction vector Y; the first maximum absorption peak label of the current evaluation record is extracted as a corresponding current label vector, and the current prediction vector Y is taken as a corresponding current prediction vector, and a corresponding prediction-label pair is formed by the current prediction vector and the current label vector; and when the round of traversal ends, all the prediction-label pairs obtained are brought into a preset first model evaluation function to obtain a corresponding first evaluation value;

[0030] The first model evaluation function is implemented based on an MAE function, an MSE function or an RMSE function.

[0031] Step 47, whether the first evaluation value meets a preset first evaluation value range is identified; if not, the step 42 is returned to continue training; if yes, the training is stopped and it is confirmed that the model training is ended.

[0032] Preferably, the first prediction vector set corresponding to the maximum absorption peak information of the first molecular structure under all the first environment parameters in the first environment parameter set is fed back to the current user by using the maximum absorption peak prediction model to predict the maximum absorption peak information, and specifically includes:

[0033] A round of traversal is performed on all the first environment parameters of the first environment parameter set; and during the round of traversal, the first environment parameter currently traversed is taken as a corresponding current environment parameter; the first molecular structure and the temperature, the light intensity and the PH value of the current environment parameter are taken as the molecular structure X, the environment temperature, the environment light intensity and the environment PH value respectively, and are input into the maximum absorption peak prediction model for prediction processing to obtain a corresponding prediction vector Y; the predicted absorbance and the predicted wavelength of the current prediction vector Y are taken as the first absorbance and the first wavelength respectively to form a corresponding first prediction vector; and when the round of traversal ends, all the first prediction vectors obtained form a corresponding first prediction vector set which is fed back to the current user.

[0034] The second aspect of the embodiment of the present application provides a device for implementing the processing method of the organic molecule maximum absorption peak prediction model in the first aspect, and the device comprises a model construction module, a data acquisition module, a model training module and a model application module.

[0035] The model construction module is configured to construct a deep learning model taking the MLP model and the Uni-Mol model as core components, and the deep learning model is denoted as a corresponding maximum absorption peak prediction model; the maximum absorption peak prediction model is configured to predict the maximum absorbance and the maximum absorption wavelength of the maximum absorption peak according to the model input of the molecular structure X, the environmental temperature, the environmental illumination intensity and the environmental PH value, and output a corresponding prediction vector Y; the prediction vector Y comprises a predicted absorbance and a predicted wavelength, and the predicted absorbance and the predicted wavelength correspond to the maximum absorbance and the maximum absorption wavelength, respectively;

[0036] The data acquisition module is configured to collect the molecular structure data of a plurality of types of organic molecules of an organic photovoltaic material and the maximum absorption peak information of each organic molecule under different environmental conditions through experiments, simulation calculations and big data acquisition, and form a corresponding first data set based on the collected data;

[0037] The model training module is configured to train the maximum absorption peak prediction model based on the first data set;

[0038] The model application module is configured to receive a first molecular structure and a first environmental parameter set input by a user after the model training is completed, and predict the maximum absorption peak information of the first molecular structure under all first environmental parameters in the first environmental parameter set by using the maximum absorption peak prediction model to obtain a corresponding first prediction vector set and feed back to the current user; the first environmental parameter set comprises a plurality of first environmental parameters; the first environmental parameters comprise temperature, illumination intensity and PH value; the first prediction vector set comprises a plurality of first prediction vectors; and each first prediction vector is composed of a first absorbance and a first wavelength and corresponds to one of the first environmental parameters.

[0039] The third aspect of the embodiment of the present application provides an electronic device, comprising a memory, a processor and a transceiver.

[0040] The processor is configured to read and execute instructions in the memory to implement the method steps in the first aspect;

[0041] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transceiving.

[0042] The fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores computer instructions, and when the computer instructions are executed by a computer, the computer instructions make the computer execute the method of the first aspect.

[0043] The embodiment of the present application provides a processing method and device of an organic molecule maximum absorption peak prediction model, an electronic device and a computer readable storage medium. According to the above content, the embodiment of the present application takes the MLP model and the Uni-Mol model as core components to construct a deep learning model for predicting the maximum absorption peak parameters (maximum absorbance and maximum absorption wavelength) of a current molecule according to the molecular structure, environmental temperature, environmental light intensity and environmental PH value input by the model, which is referred to as a maximum absorption peak prediction model. The molecular structure data of a plurality of types of organic molecules of an organic photovoltaic material and the maximum absorption peak information of each organic molecule under different environmental conditions (temperature, light intensity and PH value) are collected through experiments, simulation calculations and big data collection, so as to obtain a first data set for model training. The maximum absorption peak prediction model is trained based on the first data set. After the model training is completed, the maximum absorption peak prediction model is used to predict the maximum absorption peak information of the molecular structure input by a user under a plurality of environmental parameters (temperature, light intensity and PH value) specified by the user. The embodiment of the present application shortens the task processing period of the maximum absorption peak identification task and improves the real-time performance of task processing on the one hand. On the other hand, the personalized prediction level of the maximum absorption peak parameters is improved by combining environmental parameters during model prediction. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 A processing method of an organic molecule maximum absorption peak prediction model provided by the first embodiment of the present application is shown in the figure.

[0045] Figure 2 A module structure diagram of the maximum absorption peak prediction model provided by the first embodiment of the present application is shown in the figure.

[0046] Figure 3 A module structure diagram of a processing device of an organic molecule maximum absorption peak prediction model provided by the second embodiment of the present application is shown in the figure.

[0047] Figure 4 A structure schematic diagram of an electronic device provided by the third embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0048] In order to make the objects, technical solutions and advantages of the present application clearer, the following will further describe the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0049] The embodiment one of the present application provides a processing method of an organic molecule maximum absorption peak prediction model, which is shown in a schematic diagram of the processing method of the organic molecule maximum absorption peak prediction model provided by the embodiment one of the present application. Figure 1 The embodiment one of the present application provides a processing method of an organic molecule maximum absorption peak prediction model, which is shown in a schematic diagram of the processing method of the organic molecule maximum absorption peak prediction model provided by the embodiment one of the present application.

[0050] Step 1, constructing a deep learning model with MLP model and Uni-Mol model as core components, denoted as corresponding maximum absorption peak prediction model.

[0051] Here, the maximum absorption peak prediction model of the embodiment of the present application is used to predict the maximum absorbance and the maximum absorption wavelength of the maximum absorption peak according to the model input of the molecular structure X, the environmental temperature, the environmental light intensity and the environmental PH value, and output the corresponding prediction vector Y.

[0052] The model input of the molecular structure X includes a plurality of atoms a i , 1≤atomic index i≤A X , A X is the total number of atoms of the molecular structure X; the atomic parameters of each atom a i include the atomic type t i and the three-dimensional atomic coordinates c i (x, y, z). The model input of the environmental temperature is a temperature value, and the unit is Celsius degree by default. It needs to be explained that the unit of all temperatures mentioned in the embodiment of the present application is Celsius degree by default. The model input of the environmental light intensity is the light intensity of a preset light type, such as ultraviolet light, infrared light, etc. It needs to be explained that all light intensities mentioned in the embodiment of the present application are the light intensities of the preset light type. The model input of the environmental PH value is the acid-base degree parameter of the chemical environment where the organic molecule corresponding to the current molecular structure X is located, and its value is any integer between 0-14 in the conventional PH value domain. The model output of the prediction vector Y includes the predicted absorbance and the predicted wavelength. The predicted absorbance corresponds to the maximum absorbance of the maximum absorption peak, and the predicted wavelength corresponds to the maximum absorption wavelength of the maximum absorption peak.

[0053] The model structure of the maximum absorption peak prediction model is as follows Figure 2A module structure diagram of the maximum absorption peak prediction model provided for the first embodiment of the present application is shown, comprising: a preprocessing module, a first MLP model, a Uni-Mol model, a feature fusion module, a second MLP model, a third MLP model, and a prediction output module.

[0054] The connection relationship of each component of the maximum absorption peak prediction model is as follows: the input end of the preprocessing module is connected with the model input end of the maximum absorption peak prediction model, and the first and second output ends are respectively connected with the input ends of the first MLP model and the Uni-Mol model; the output ends of the first MLP model and the Uni-Mol model are respectively connected with the first and second input ends of the feature fusion module; the output end of the feature fusion module is respectively connected with the input ends of the second MLP model and the third MLP model; the output ends of the second MLP model and the third MLP model are respectively connected with the first and second input ends of the prediction output module; and the output end of the prediction output module is connected with the model output end of the maximum absorption peak prediction model.

[0055] The functions of each component of the maximum absorption peak prediction model are as follows.

[0056] 1) Preprocessing module:

[0057] The preprocessing module of the embodiment of the present application is used for embedding and coding the environmental temperature, the environmental illumination intensity and the environmental PH value of the model input respectively according to the preset temperature, illumination intensity and PH value embedding and coding mode to obtain the corresponding temperature coding, illumination intensity coding and PH value coding to form a corresponding first environmental coding vector; and based on the atom type one-hot coding rule of the Uni-Mol model, according to all atom types of the molecular structure X, the atom type one-hot coding of each atom a i is performed to obtain a corresponding first atom one-hot coding vector, and all the obtained first atom one-hot coding vectors form a corresponding first atom coding tensor; and based on the atom pair feature initialization coding rule of the Uni-Mol model, according to all three-dimensional atom coordinates of the molecular structure X, the paired atom feature initialization of each pair of atoms is performed to obtain a corresponding first atom pair coding tensor; the first environmental coding vector is sent to the first MLP model; and the first atom coding tensor and the first atom pair coding tensor are sent to the Uni-Mol model.

[0058] Here, the preset temperature, light intensity and PH value embedding coding mode of the embodiment of the application is three kinds of preset embedding coding modes, which can be customized based on specific application requirements. One of the conventional customization methods is: based on the normalization method, the temperature and light intensity are processed and the normalized values obtained are used as the corresponding temperature encoding and light intensity encoding, and the value of the PH value is directly used as the corresponding PH value encoding. In addition, the Uni-Mol model used by the embodiment of the application is an encoder model implemented based on the Encoder sub-model of the Transformer model and used for atomic-level feature encoding of molecular structure. The detailed model structure, model inference principle, atomic type one-hot encoding rule, atomic pair feature initialization encoding rule and model pre-training scheme of the Uni-Mol model are described in detail in the public technical document A <Uni-Mol: A Universal 3D Molecular Representation Learning Framework>, so the atomic type one-hot encoding rule and atomic pair feature initialization encoding rule used by the preprocessing module will not be further described here.

[0059] It should be noted that the first environment encoding vector generated in the processing process of the preprocessing module of the embodiment of the application is a real number vector with a length of 3, which is composed of temperature encoding, light intensity encoding and PH value encoding. The first atomic one-hot encoding vector generated by the preprocessing module is a one-dimensional one-hot encoding vector, the vector length is denoted as D1, D1 matches the total number of atomic types of the molecular structure X; the one-hot encoding corresponding to the atomic type t i of the current atom a i is set to 1, and the remaining D1-1 one-hot encodings are set to 0. The tensor shape of the first atomic encoding tensor generated by the preprocessing module is D1xA X . The tensor shape of the first atomic pair encoding tensor generated by the preprocessing module is D2xA X xA X , and D2 is the preset atomic pair encoding feature dimension. If the preprocessing module only calculates and normalizes the Euclidean distance of the atomic pair, D2 can be set to 1 at the minimum.

[0060] 2) First MLP model:

[0061] The first MLP model of the embodiment of the application is sequentially connected by multiple fully connected layers. The first MLP model is used to perform environment feature extraction processing according to the input first environment encoding vector to obtain the corresponding first environment feature tensor and send it to the feature fusion module.

[0062] Here, the tensor shape of the first environmental feature tensor of the embodiment of the application is D3x3, and D3 is a preset environmental feature dimension; the first environmental feature tensor can be specifically regarded as being composed of three environmental feature vectors with a length of D3, which are a first temperature feature vector, a first illumination intensity feature vector and a first PH value feature vector.

[0063] 3) Uni-Mol model:

[0064] The Uni-Mol model of the embodiment of the application is used for performing atomic and atomic pair feature extraction processing on the input first atomic encoding tensor and the first atomic pair encoding tensor to obtain a corresponding first atomic feature tensor and a first atomic pair feature tensor; and performing feature fusion on the first atomic feature tensor and the first atomic pair feature tensor to obtain a corresponding first molecular feature tensor.

[0065] Here, the tensor shape of the first atomic feature tensor of the embodiment of the application is D4xA X , and D4 is a preset atomic feature dimension; the first atomic feature tensor can be specifically regarded as being composed of A X first sub-feature vectors with a length of D4.

[0066] The tensor shape of the first atomic pair feature tensor of the embodiment of the application is D5xA X xA X , and D5 is a preset atomic pair feature dimension; the first atomic pair feature tensor can be specifically regarded as being composed of A X first sub-feature tensors with a shape of D5xA X , and each first sub-feature tensor can be specifically regarded as being composed of A X second sub-feature vectors with a length of D5.

[0067] The tensor shape of the first molecular feature tensor of the embodiment of the application is (D4+D5xA X )xA X ; the first molecular feature tensor can be specifically regarded as being composed of A X first atomic fusion vectors with a length of (D4+D5xA X ); and each first atomic fusion vector can be regarded as being sequentially spliced by a corresponding first sub-feature vector and A X corresponding second sub-feature vectors.

[0068] As can be known from the foregoing, the detailed model structure, model reasoning principle and model pre-training scheme of the Uni-Mol model used in the embodiment of the present application have been published in the public technical document A, so the encoding process of the Uni-Mol model will not be further described here. It should be noted that the Uni-Mol model used in the embodiment of the present application has been pre-trained according to the model pre-training scheme given in the public technical document A.

[0069] 4) Feature fusion module:

[0070] The feature fusion module of the embodiment of the present application is used to sequentially splice each first atomic fusion vector of the first molecular feature tensor and the three environmental feature vectors of the first environmental feature tensor to form a corresponding second atomic fusion vector; and all the obtained second atomic fusion vectors form a corresponding first fusion feature tensor which is sent to the second and third MLP models.

[0071] Here, the tensor shape of the first fusion feature tensor of the embodiment of the present application is (D3x3+D4+D5xA X )xA X .

[0072] 5) Second MLP model:

[0073] The second MLP model of the embodiment of the present application is sequentially connected by multiple fully connected layers. The second MLP model is used to perform maximum absorbance prediction processing according to the input first fusion feature tensor to obtain a corresponding predicted absorbance which is sent to the prediction output module.

[0074] 6) Third MLP model:

[0075] The third MLP model of the embodiment of the present application is sequentially connected by multiple fully connected layers. The third MLP model is used to perform maximum absorption wavelength prediction processing according to the input first fusion feature tensor to obtain a corresponding predicted wavelength which is sent to the prediction output module.

[0076] 7) Prediction output module:

[0077] The prediction output module of the embodiment of the present application is composed of the predicted absorbance and the predicted wavelength to form a corresponding prediction vector Y and output.

[0078] Step 2, collect the molecular structure data of multiple types of organic molecules of the organic photovoltaic material and the maximum absorption peak information of each organic molecule under different environmental conditions through experiments, simulation calculations and big data collection, and form a corresponding first data set based on the collected data;

[0079] The environmental conditions include temperature, light intensity, and pH value; the first data set includes a plurality of first data records; the first data record includes a first training molecular structure, a first training environmental parameter, and a first maximum absorption peak label; the first training molecular structure includes a plurality of first atoms, and the atomic parameters of each first atom include atomic type and three-dimensional atomic coordinates; the first training environmental parameter includes a first temperature, a first light intensity, and a first pH value; and the first maximum absorption peak label includes a first maximum absorbance label and a first maximum absorption wavelength label.

[0080] Specifically comprising: step 21, determining the maximum absorption peak information of a plurality of organic molecules for organic photovoltaic materials and known molecular structure under different environmental conditions by an experimental method; taking the molecular structure of each organic molecule processed in the experimental process as a corresponding first training molecular structure, taking the temperature, light intensity, and pH value of each environmental condition set in the experimental process as a corresponding first temperature, first light intensity, and first pH value to form a corresponding first training environmental parameter, and taking the maximum absorbance and maximum absorption wavelength of the maximum absorption peak information of each organic molecule measured under each environmental condition in the experimental process as a corresponding first maximum absorbance label and first maximum absorption wavelength label to form a corresponding first maximum absorption peak label; and finally, each first training molecular structure and its corresponding first training environmental parameter and first maximum absorption peak label obtained by the experimental method form a corresponding first data record;

[0081] Step 22, simulating and calculating the maximum absorption peak information of a plurality of organic molecules for organic photovoltaic materials and known molecular structure under different environmental conditions by a simulation calculation method; taking the molecular structure of each organic molecule processed in the simulation calculation process as a corresponding first training molecular structure, taking the temperature, light intensity, and pH value of each environmental condition set in the simulation calculation process as a corresponding first temperature, first light intensity, and first pH value to form a corresponding first training environmental parameter, and taking the maximum absorbance and maximum absorption wavelength of the maximum absorption peak information of each organic molecule measured under each environmental condition in the simulation calculation process as a corresponding first maximum absorbance label and first maximum absorption wavelength label to form a corresponding first maximum absorption peak label; and finally, each first training molecular structure and its corresponding first training environmental parameter and first maximum absorption peak label obtained by the simulation calculation method form a corresponding first data record;

[0082] Step 23, through a plurality of preset big data collection channels, maximum absorption peak information of a plurality of organic molecules for organic photovoltaic materials and known molecular structure under different environmental conditions is collected; the molecular structure of each organic molecule collected is taken as a corresponding first training molecular structure, the temperature, light intensity and PH value of each environmental condition corresponding to each organic molecule collected are taken as a corresponding first temperature, first light intensity and first PH value to form a corresponding first training environmental parameter, and the maximum absorbance and maximum absorption wavelength of the maximum absorption peak information of each organic molecule under each corresponding environmental condition are taken as a corresponding first maximum absorbance label and first maximum absorption wavelength label to form a corresponding first maximum absorption peak label; finally, each first training molecular structure and the corresponding first training environmental parameter and first maximum absorption peak label obtained by big data collection and processing form a corresponding first data record;

[0083] Here, the plurality of big data collection channels of the embodiment of the application at least include a molecular structure information library of all disclosed organic photovoltaic material molecules, a molecular optical property information library, all disclosed scientific literature, periodicals, magazines, papers, technical solutions and experimental reports, etc.

[0084] Step 24, all first data records obtained by experimental methods, simulation calculation methods and big data collection methods form a corresponding first record set; the first data records in the first record set are subjected to de-duplication and abnormal data filtering processing; finally, the first record set after de-duplication and abnormal data filtering is taken as a corresponding first data set.

[0085] Step 3, model training is performed on the maximum absorption peak prediction model based on the first data set;

[0086] Specifically, step 31, the first data set is segmented based on a preset first segmentation ratio to obtain two sub-data sets, which are taken as a corresponding first training set and a first evaluation set;

[0087] Here, the first segmentation ratio of the embodiment of the application is a pre-set ratio parameter, for example, 8:2; the first training set and the first evaluation set are both composed of a plurality of first data records; the total number ratio of the first training set and the first evaluation set satisfies the first segmentation ratio;

[0088] Step 32, the first first data record of the first training set is taken as a corresponding current training record;

[0089] Step 33, input the first training molecular structure of the current training record and the first temperature, the first light intensity and the first PH value of the first training environment parameter as the corresponding molecular structure X, the environmental temperature, the environmental light intensity and the environmental PH value into the maximum absorption peak prediction model for prediction processing to obtain the corresponding prediction vector Y;

[0090] Step 34, extract the first maximum absorption peak label of the current training record as the corresponding current label vector, take the current prediction vector Y as the corresponding current prediction vector, and bring the current prediction vector and the current label vector into the preset first model loss function to obtain the corresponding first loss value;

[0091] Here, the first model loss function of the embodiment of the application is realized based on the L1 loss function or the L2 loss function;

[0092] Step 35, identify whether the first loss value meets the preset first loss value range; if the first loss value meets the first loss value range, identify whether the current training record is the last first data record of the first training set, if yes, go to step 36, if not, take the next first data record of the first training set as a new current training record and return to step 33 for continuous training; if the first loss value does not meet the first loss value range, based on the preset first model optimizer, one round of modulation is performed on the model parameters of the first, second and third MLP models and the Uni-Mol model of the maximum absorption peak prediction model in the direction of making the first model loss function reach the minimum value, and returns to step 33 for continuous training at the end of the round of modulation;

[0093] Here, the first loss value range of the embodiment of the application is a pre-set loss value range; the first model optimizer at least includes the Adam optimizer and the SGD optimizer;

[0094] Step 36, perform one round of traversal on all first data records of the first evaluation set; and in the round of traversal, take the currently traversed first data record as the corresponding current evaluation record; input the first training molecular structure of the current evaluation record and the first temperature, the first light intensity and the first PH value of the first training environment parameter as the corresponding molecular structure X, the environmental temperature, the environmental light intensity and the environmental PH value into the maximum absorption peak prediction model for prediction processing to obtain the corresponding prediction vector Y; extract the first maximum absorption peak label of the current evaluation record as the corresponding current label vector, take the current prediction vector Y as the corresponding current prediction vector, and form a corresponding prediction-label pair by the current prediction vector and the current label vector; and at the end of the round of traversal, bring all the obtained prediction-label pairs into the preset first model evaluation function to obtain the corresponding first evaluation value;

[0095] Here, the first model evaluation function of the embodiment of the present application is implemented based on the MAE function, the MSE function or the RMSE function.

[0096] Step 37, whether the first evaluation value satisfies a preset first evaluation value range is identified; if not, returning to step 32 for continuing training; if yes, stopping training and confirming that the model training is ended.

[0097] Here, the first evaluation value range of the embodiment of the present application is a preset evaluation value range.

[0098] Step 4, after the model training is ended, a first molecular structure and a first environmental parameter set input by a user are received; and the maximum absorption peak prediction model is used to predict the maximum absorption peak information of the first molecular structure under all first environmental parameters in the first environmental parameter set to obtain a corresponding first prediction vector set for feeding back to the current user.

[0099] Specifically, it includes: step 41, after the model training is ended, a first molecular structure and a first environmental parameter set input by a user are received;

[0100] Here, the first molecular structure of the embodiment of the present application is a molecular structure of an organic molecule for an organic photovoltaic material, which is specifically composed of multiple atoms, and the atomic parameters of each atom are composed of a corresponding atomic type and atomic three-dimensional coordinates; the first environmental parameter set includes multiple first environmental parameters; each first environmental parameter includes temperature, light intensity and PH value;

[0101] Step 42, the maximum absorption peak prediction model is used to predict the maximum absorption peak information of the first molecular structure under all first environmental parameters in the first environmental parameter set to obtain a corresponding first prediction vector set for feeding back to the current user.

[0102] Wherein, the first prediction vector set of the embodiment of the present application includes multiple first prediction vectors; the first prediction vector is composed of a first absorbance and a first wavelength; the first prediction vectors of the first prediction vector set correspond to the first environmental parameters of the first environmental parameter set one by one.

[0103] Specifically comprising: a round of traversal is performed on all first environment parameters of the first environment parameter set; and in the current round of traversal, the first environment parameter currently traversed is taken as the corresponding current environment parameter; the first molecular structure and the temperature, the light intensity and the PH value of the current environment parameter are taken as the corresponding molecular structure X, the environmental temperature, the environmental light intensity and the environmental PH value, and are input into the maximum absorption peak prediction model for prediction processing to obtain the corresponding prediction vector Y; the predicted absorbance and the predicted wavelength of the current prediction vector Y are taken as the corresponding first absorbance and first wavelength to form a corresponding first prediction vector; and at the end of the current round of traversal, all the first prediction vectors obtained form a corresponding first prediction vector set, which is fed back to the current user.

[0104] Figure 3 A module structure diagram of a processing device of an organic molecule maximum absorption peak prediction model is provided for the second embodiment of the application. The device is a terminal device or a server for implementing the method embodiments, or a device capable of enabling the terminal device or the server to implement the method embodiments, such as a device or a chip system of the terminal device or the server. As shown in the figure, the device comprises a model construction module 201, a data acquisition module 202, a model training module 203 and a model application module 204. Figure 3 The model construction module 201 is configured to construct a deep learning model taking the MLP model and the Uni-Mol model as core components, which is denoted as a corresponding maximum absorption peak prediction model.

[0105] The model construction module 201 is configured to construct a deep learning model taking the MLP model and the Uni-Mol model as core components, which is denoted as a corresponding maximum absorption peak prediction model.

[0106] The data acquisition module 202 is configured to collect the molecular structure data of multiple types of organic molecules of organic photovoltaic materials and the maximum absorption peak information of each organic molecule under different environmental conditions through experiments, simulation calculations and big data acquisition methods, and form a corresponding first data set based on the collected data.

[0107] The model training module 203 is configured to perform model training on the maximum absorption peak prediction model based on the first data set.

[0108] The model application module 204 is configured to receive a first molecular structure and a first set of environmental parameters input by a user after the model training is completed, and predict the maximum absorption peak information of the first molecular structure under all first environmental parameters in the first set of environmental parameters by using the maximum absorption peak prediction model to obtain a corresponding first set of prediction vectors, and feed back to the current user; the first set of environmental parameters comprises a plurality of first environmental parameters; the first environmental parameters comprise temperature, light intensity and pH value; the first set of prediction vectors comprises a plurality of first prediction vectors; the first prediction vector is composed of first absorbance and first wavelength, and corresponds to the first environmental parameter one by one.

[0109] The processing device of the organic molecular maximum absorption peak prediction model provided by the embodiment of the present application can execute the method steps in the method embodiment, and has similar implementation principles and technical effects, which will not be described here.

[0110] It should be noted that the division of each module of the above device is only a logical division of functions, and all or part of the modules can be integrated into one physical entity, or can be physically separated. These modules can all be implemented in the form of software called by a processing element; all can be implemented in the form of hardware; some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the model construction module can be a separately set processing element, or can be integrated in a chip of the above device, in addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determination module is called and executed by a processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together, or can be independently implemented. The processing element described herein can be an integrated circuit having signal processing capability. In the implementation process, each step of the above method or each module can be completed by the integrated logic circuit of the hardware or the instruction of the software in the processor element.

[0111] For example, the above modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. For another example, when a certain module above is implemented in the form of a processing element scheduling code, the processing element can be a general purpose processor, such as a Central Processing Unit (CPU) or other processor that can invoke code. For another example, the modules can be integrated together to implement in the form of a System-on-a-chip (SOC).

[0112] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed by a computer, the computer program instructions generate the processes or functions described in the above method embodiments, all or part of the computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.) mode. The above computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The above available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0113] Figure 4 A structural schematic diagram of an electronic device is provided for Embodiment Three of the present application. The electronic device can be a terminal device or a server that implements the method of the above embodiments, or a terminal device or a server connected to the terminal device or the server that implements the method of the above embodiments. As shown in FIG. 3, the electronic device can include a processor 301, a memory 302, a transceiver 303, and a communication interface 304. The processor 301 can be configured to implement the method of the above embodiments. The memory 302 can be configured to store the computer instructions executed by the processor 301. The transceiver 303 can be configured to transmit and receive data. The communication interface 304 can be configured to connect to another electronic device. Figure 4As shown, the electronic device can include a processor 301 (such as a CPU), a memory 302, a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiving action of the transceiver 303. The memory 302 can store various instructions for completing various processing functions and implementing the processing steps described in the foregoing embodiment method description. Preferably, the electronic device related to the embodiments of the present application further includes a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize the communication connection between elements. The above-mentioned communication port 306 is used for connection communication between the electronic device and other peripherals.

[0114] In Figure 4 The system bus 305 mentioned in the above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize the communication between the database access device and other devices (such as the client, the read-write library and the read-only library). The memory can contain a Random Access Memory (RAM), and can also include a Non-Volatile Memory, such as at least one disk memory.

[0115] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0116] It should be noted that the embodiments of the present application also provide a computer-readable storage medium, which stores instructions, when running on a computer, causes the computer to execute the method and processing procedure provided in the above embodiments.

[0117] The embodiment of the present application provides a processing method and device of an organic molecule maximum absorption peak prediction model, an electronic equipment and a computer readable storage medium. According to the above content, the embodiment of the present application takes the MLP model and the Uni-Mol model as core components to construct a deep learning model for predicting the maximum absorption peak parameters (maximum absorbance, maximum absorption wavelength) of the current molecule according to the model input of the molecular structure, the environmental temperature, the environmental light intensity and the environmental PH value, which is recorded as the maximum absorption peak prediction model. And through experiments, simulation calculation and big data collection, the molecular structure data of a plurality of types of organic molecules of the organic photovoltaic material and the maximum absorption peak information of each organic molecule under different environmental conditions (temperature, light intensity, PH value) are collected to obtain a first data set for model training. And based on the first data set, the maximum absorption peak prediction model is trained. After the model training is completed, the maximum absorption peak prediction model is used to predict the maximum absorption peak information of the molecular structure input by the user under a plurality of environmental parameters (temperature, light intensity, PH value) specified by the user. The embodiment of the present application shortens the task processing period of the maximum absorption peak identification task and improves the task processing real-time performance on the one hand. On the other hand, the personalized prediction level of the maximum absorption peak parameters is improved by combining the environmental parameters during model prediction.

[0118] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), memory, flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0119] The above specific embodiments further specifically describe the purposes, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.

Claims

1. A processing method of an organic molecule maximum absorption peak prediction model, characterized by, The method comprises: A deep learning model is constructed with MLP model and Uni-Mol model as core components, which is recorded as a corresponding maximum absorption peak prediction model; the maximum absorption peak prediction model is used to predict the maximum absorbance and maximum absorption wavelength of the maximum absorption peak according to the model input molecular structure X, environmental temperature, environmental light intensity and environmental PH value, and output the corresponding prediction vector Y; the prediction vector Y includes predicted absorbance and predicted wavelength, which correspond to maximum absorbance and maximum absorption wavelength respectively; Through experiments, simulation calculations and big data collection methods, the molecular structure data of multiple types of organic molecules of organic photovoltaic materials and the maximum absorption peak information of each organic molecule under different environmental conditions are collected to form a corresponding first data set based on the collected data; the environmental conditions include temperature, light intensity and PH value; The maximum absorption peak prediction model is trained based on the first data set; After the model training is completed, the first molecular structure and the first environmental parameter set input by the user are received; and the maximum absorption peak prediction model is used to predict the maximum absorption peak information of the first molecular structure under all first environmental parameters in the first environmental parameter set to obtain a corresponding first prediction vector set to feedback to the current user; the first environmental parameter set includes multiple first environmental parameters; the first environmental parameters include temperature, light intensity and PH value; the first prediction vector set includes multiple first prediction vectors; the first prediction vector is composed of first absorbance and first wavelength, and corresponds to the first environmental parameter one by one.

2. The processing method of the organic molecule maximum absorption peak prediction model according to claim 1, characterized in that: The molecular structure X includes a plurality of atoms a i , 1 ≤ atom index i ≤ A X , A X is the total number of atoms of the molecular structure X; each atom parameter of the atom a i includes an atom type t i and three-dimensional atom coordinates c i (x, y, z); The first data set includes multiple first data records; the first data record includes a first training molecular structure, a first training environmental parameter and a first maximum absorption peak label; The first training molecular structure includes multiple first atoms, and the atomic parameters of each first atom include atomic type and three-dimensional atomic coordinates; the first training environmental parameter includes first temperature, first light intensity and first PH value; the first maximum absorption peak label includes first maximum absorbance label and first maximum absorption wavelength label.

3. The processing method of the organic molecule maximum absorption peak prediction model according to claim 2, characterized in that: The maximum absorption peak prediction model includes a preprocessing module, a first MLP model, a Uni-Mol model, a feature fusion module, a second MLP model, a third MLP model and a prediction output module; The input end of the preprocessing module is connected with the model input end of the maximum absorption peak prediction model, and the first and second output ends are respectively connected with the input ends of the first MLP model and the Uni-Mol model; the output ends of the first MLP model and the Uni-Mol model are respectively connected with the first and second input ends of the feature fusion module; the output end of the feature fusion module is connected with the input ends of the second MLP model and the third MLP model; the output ends of the second MLP model and the third MLP model are respectively connected with the first and second input ends of the prediction output module; and the output end of the prediction output module is connected with the model output end of the maximum absorption peak prediction model. The preprocessing module is configured to embed and encode the environment temperature, the environment illumination intensity and the environment PH value of the model input respectively according to a preset temperature, illumination intensity and PH value embedding and encoding mode to obtain corresponding temperature encoding, illumination intensity encoding and PH value encoding to form a corresponding first environment encoding vector; and based on an atomic type one-hot encoding rule of the Uni-Mol model, each atomic type of the molecular structure X is subjected to atomic type one-hot encoding to obtain a corresponding first atomic one-hot encoding vector, and all the obtained first atomic one-hot encoding vectors form a corresponding first atomic encoding tensor; and based on an atomic pair feature initialization encoding rule of the Uni-Mol model, all three-dimensional atomic coordinates of the molecular structure X are used to perform paired atomic feature initialization on each pair of atoms to obtain a corresponding first atomic pair encoding tensor; the first environment encoding vector is sent to the first MLP model; and the first atomic encoding tensor and the first atomic pair encoding tensor are sent to the Uni-Mol model; the first environment encoding vector is a real number vector with a length of 3, which is composed of the temperature encoding, the illumination intensity encoding and the PH value encoding; the first atomic one-hot encoding vector is a one-dimensional one-hot encoding vector, and the vector length is denoted as D1, which matches the total number of atomic types of the molecular structure X; the one-hot encoding corresponding to the atomic type t of the current atomic a in the first atomic one-hot encoding vector is set to 1, and the remaining D1-1 one-hot encodings are set to 0; the tensor shape of the first atomic encoding tensor is D1xA X ; the tensor shape of the first atomic pair encoding tensor is D2xA X xA X , and D2 is a preset atomic pair encoding feature dimension. i i i X X X The first MLP model is sequentially connected by multiple full connection layers; the first MLP model is used for performing environment feature extraction processing on the input first environment encoding vector to obtain a corresponding first environment feature tensor, and sending the first environment feature tensor to the feature fusion module; the tensor shape of the first environment feature tensor is D3×3, and D3 is a preset environment feature dimension; the first environment feature tensor is specifically composed of three environment feature vectors with a length of D3, which are a first temperature feature vector, a first light intensity feature vector and a first PH value feature vector; The Uni-Mol model is used to perform atomic and atomic pair feature extraction processing according to the input first atomic encoding tensor and the first atomic pair encoding tensor to obtain corresponding first atomic feature tensors and first atomic pair feature tensors; and the first atomic feature tensors and the first atomic pair feature tensors are fused to obtain a corresponding first molecular feature tensor, which is sent to the feature fusion module; the tensor shape of the first atomic feature tensor is D4xA X , and D4 is a preset atomic feature dimension; the The first atomic feature tensor is specifically composed of A X first sub-feature vector groups each having a vector length of D4; the tensor shape of the first atomic pair feature tensor is D5xA X xA X , D5 is a preset atomic pair feature dimension; the first atomic pair feature tensor is specifically composed of A X first sub-feature tensors each having a shape of D5xA X , each of the first sub-feature tensors is specifically composed of A X second sub-feature vectors each having a vector length of D5; the tensor shape of the first molecular feature tensor is (D4+D5xA X )xA X ; the first molecular feature tensor is specifically composed of A X first atomic fusion vectors each having a vector length of (D4+D5xA X ); each of the first atomic fusion vectors is sequentially spliced from a corresponding first sub-feature vector and A X corresponding second sub-feature vectors. The feature fusion module is used for sequentially splicing each first atomic fusion vector of the first molecular feature tensor and three environment feature vectors of the first environment feature tensor to form a corresponding second atomic fusion vector; and all obtained second atomic fusion vectors are used to form a corresponding first fusion feature tensor, which is sent to the second and third MLP models; the tensor shape of the first fusion feature tensor is (D3×3+D4+D5×A X )×A X ; The second MLP model is sequentially connected by multiple full connection layers; the second MLP model is used for performing maximum absorbance prediction processing on the input first fusion feature tensor to obtain a corresponding predicted absorbance, and sending the predicted absorbance to the prediction output module; The third MLP model is sequentially connected by multiple full connection layers; the third MLP model is used for performing maximum absorption wavelength prediction processing on the input first fusion feature tensor to obtain a corresponding predicted wavelength, and sending the predicted wavelength to the prediction output module; The prediction output module is composed of the predicted absorbance and the predicted wavelength to obtain a corresponding prediction vector Y and output the prediction vector Y.

4. The processing method of claim 3, wherein the processing method is characterized by: The maximum absorption peak prediction model is trained based on the first data set, and specifically includes: Step 41, based on a preset first segmentation ratio, the first data set is segmented to obtain two sub-data sets, which are denoted as a corresponding first training set and a first evaluation set; Wherein, the first training set and the first evaluation set are both composed of a plurality of first data records; the total number ratio of the first training set and the first evaluation set meets the first segmentation ratio; Step 42, taking the first first data record of the first training set as a corresponding current training record; Step 43, taking the first training molecular structure and the first temperature, the first light intensity and the first PH value of the first training environment parameter of the current training record as the molecular structure X, the environment temperature, the environment light intensity and the environment PH value, inputting the maximum absorption peak prediction model for prediction processing to obtain a corresponding prediction vector Y; Step 44, taking the first evaluation set as a corresponding current evaluation set; Step 44, the first maximum absorption peak label of the current training record is extracted as a corresponding current label vector, and the current prediction vector Y is taken as a corresponding current prediction vector; and the current prediction vector and the current label vector are taken into a preset first model loss function for calculation to obtain a corresponding first loss value; The first model loss function is implemented based on an L1 loss function or an L2 loss function; Step 45, whether the first loss value meets a preset first loss value range is identified; if the first loss value meets the first loss value range, whether the current training record is the last first data record of the first training set is identified; if yes, step 46 is turned to; if no, the next first data record of the first training set is taken as a new current training record, and step 43 is returned to continue training; if the first loss value does not meet the first loss value range, the model parameters of the first, second and third MLP models and the Uni-Mol model of the maximum absorption peak prediction model are modulated in a round based on a preset first model optimizer in a direction of making the first model loss function reach a minimum value, and the round of modulation is ended and step 43 is returned to continue training; The first model optimizer at least includes an Adam optimizer and an SGD optimizer; Step 46, all the first data records of the first evaluation set are traversed in a round; and in the round of traversal, the first data record currently traversed is taken as a corresponding current evaluation record; the first training molecular structure of the current evaluation record and the first temperature, the first light intensity and the first PH value of the first training environment parameter are taken as the molecular structure X, the environment temperature, the environment light intensity and the environment PH value respectively, and the maximum absorption peak prediction model is input for prediction processing to obtain a corresponding prediction vector Y; the first maximum absorption peak label of the current evaluation record is extracted as a corresponding current label vector, and the current prediction vector Y is taken as a corresponding current prediction vector, and a corresponding prediction-label pair is formed by the current prediction vector and the current label vector; and when the round of traversal is ended, all the prediction-label pairs obtained are taken into a preset first model evaluation function for calculation to obtain a corresponding first evaluation value; The first model evaluation function is implemented based on an MAE function, an MSE function or an RMSE function; Step 47, whether the first evaluation value meets a preset first evaluation value range is identified; if no, step 42 is returned to continue training; if yes, training is stopped and model training is confirmed to be ended.

5. The processing method for the organic molecule maximum absorption peak prediction model according to claim 1, characterized in that, The maximum absorption peak information of the first molecular structure under all the first environment parameters in the first environment parameter set is predicted by the maximum absorption peak prediction model to obtain a corresponding first prediction vector set, and the first prediction vector set is fed back to a current user, and the method specifically comprises the following steps: perform a round of traversal on all the first environment parameters of the first environment parameter set; and during the current round of traversal, take the first environment parameter currently being traversed as a corresponding current environment parameter; take the first molecular structure and the temperature, the light intensity, and the PH value of the current environment parameter as the molecular structure X, the environment temperature, the environment light intensity, and the environment PH value respectively, input the maximum absorption peak prediction model for prediction processing to obtain a corresponding prediction vector Y; take the prediction absorbance and the prediction wavelength of the current prediction vector Y as the first absorbance and the first wavelength respectively to form a corresponding first prediction vector; and when the current round of traversal ends, form a corresponding first prediction vector set by all the first prediction vectors obtained, and feed back to the current user.

6. An apparatus for performing a processing method of the prediction model of the maximum absorption peak of an organic molecule according to any one of claims 1 to 5, characterized in that, The device comprises a model construction module, a data acquisition module, a model training module, and a model application module. The model construction module is configured to construct a deep learning model taking the MLP model and the Uni-Mol model as core components, denoted as a corresponding maximum absorption peak prediction model; the maximum absorption peak prediction model is configured to predict the maximum absorbance and the maximum absorption wavelength of the maximum absorption peak according to the molecular structure X, the environment temperature, the environment light intensity, and the environment PH value input by the model, and output a corresponding prediction vector Y; the prediction vector Y comprises a prediction absorbance and a prediction wavelength, which correspond to the maximum absorbance and the maximum absorption wavelength respectively; The data acquisition module is configured to collect the molecular structure data of multiple types of organic molecules of organic photovoltaic materials and the maximum absorption peak information of each organic molecule under different environment conditions through experiments, simulation calculations, and big data acquisition methods, and form a corresponding first data set based on the collected data; the environment conditions include temperature, light intensity, and PH value; The model training module is configured to train the maximum absorption peak prediction model based on the first data set; The model application module is configured to, after the model training ends, receive a first molecular structure and a first environment parameter set input by a user; and use the maximum absorption peak prediction model to predict the maximum absorption peak information of the first molecular structure under all first environment parameters in the first environment parameter set to obtain a corresponding first prediction vector set, and feed back to the current user; the first environment parameter set comprises multiple first environment parameters; the first environment parameters include temperature, light intensity, and PH value; the first prediction vector set comprises multiple first prediction vectors; and each first prediction vector is composed of a first absorbance and a first wavelength, and corresponds to a first environment parameter.

7. An electronic device, comprising: comprise: a memory, a processor, and a transceiver; the processor is configured to read and execute instructions in the memory to implement the method of any one of claims 1-5; the transceiver is coupled with the processor, and is controlled by the processor to perform message transmission and reception.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, when the computer instructions are executed by a computer, the computer instructions cause the computer to execute the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Push-pull type organic molecule two-photon absorption cross section prediction method and system

    CN115472237A

  • Treatment method and device for structural optimization of electrolyte formula components

    CN119560057A