Processing method and device for molecular affinity analysis

By constructing a molecular affinity prediction model based on Uni-Mol model and MLP model, the problem of high computational complexity in molecular affinity analysis of organic photovoltaic materials is solved, and efficient molecular affinity prediction and atomic contribution analysis are achieved.

CN120183560AActive Publication Date: 2025-06-20BEIJING DP TECH CO LTD

Patent Information

Application Number
CN202510283221.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-20
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The prior art has problems such as high computational complexity, long computational time and low computational efficiency in the molecular affinity analysis of organic photovoltaic materials.

Method used

The Uni-Mol model is used as an encoder and the molecular affinity prediction model is constructed in combination with the MLP model. The training data set is constructed through big data acquisition, the model is trained to achieve molecular affinity prediction, and the atomic-level affinity contribution is estimated through scrambling.

Benefits of technology

The overall affinity prediction and atomic affinity contribution analysis for organic photovoltaic materials molecules are realized, reducing the analysis complexity, shortening the analysis time, and improving the analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183560A_ABST
    Figure CN120183560A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a processing method and device for molecular affinity analysis. The method comprises the steps that a molecular affinity prediction model used for conducting molecular affinity prediction according to a molecular structure input by a model is constructed in the mode that a Uni-Mo < l > model serves as an encoder and an MLP model serves as an encoder downstream regression calculation task model; building a first data set through big data collection, and training a molecular affinity prediction model based on the first data set; after training is finished, prediction is carried out based on a molecular affinity prediction model according to a molecular structure input by a user, and contribution scores of atoms in the structure input by the user to molecular affinity are analyzed; and feeding back the obtained predicted molecular affinity and contribution scores of all atoms to the current user. According to the invention, the analysis complexity can be reduced, the analysis time is shortened, and the analysis efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a processing method and device for molecular affinity analysis. Background Art

[0002] Organic Photovoltaics (OPV) materials refer to organic materials that directly convert solar energy or other light energy into electrical energy through the photovoltaic effect and are composed of organic molecules. The molecular affinity of such organic molecules has a positive impact on improving the stability, light absorption efficiency, and photovoltaic conversion efficiency of the materials. Currently, there are two types of tasks for analyzing the molecular affinity of such organic molecules: one is to predict the overall molecular affinity, and the other is to predict the affinity contribution of each atom. Conventionally, both of these analysis tasks are implemented based on simulation calculation methods, that is, through a series of quantum chemical calculation methods (such as density functional theory calculations, etc.) to predict molecular affinity and atomic contribution fractions. However, from practical experience, such conventional analysis means often have problems such as high computational complexity, long computational time, and low computational efficiency. Summary of the Invention

[0003] The purpose of the present invention is to provide a processing method, device, electronic device, and computer-readable storage medium for molecular affinity analysis in view of the defects of the prior art. The present invention constructs a molecular affinity prediction model by using the Uni-Mol model as the encoder and the MLP model as the regression calculation task model downstream of the encoder. The molecular affinity prediction model is used to perform molecular affinity prediction processing based on the molecular structure input to the model and output the corresponding predicted affinity; and a first data set for model training is constructed through big data collection, and the molecular affinity prediction model is trained based on the first data set; and after the model training is completed, the molecular affinity prediction model is used to predict the molecular affinity of the molecular structure specified by the user, and the affinity contribution fraction of each atom in the current molecular structure is estimated and analyzed by scrambling the process features generated during the model processing. Through the present invention, not only can two affinity analysis tasks for organic photovoltaic material molecules (predicting the overall molecular affinity and analyzing the affinity contribution at the atomic level) be realized, but also the analysis complexity can be reduced, the analysis duration can be shortened, and the analysis efficiency can be improved.

[0004] To achieve the above object, in the first aspect of the embodiments of the present invention, a processing method for molecular affinity analysis is provided, and the method includes:

[0005] Construct a molecular affinity prediction model by using the Uni-Mol model as the encoder and the MLP model as the regression calculation task model downstream of the encoder; the molecular affinity prediction model is used to perform molecular affinity prediction processing on the input molecular structure X of the model and output the corresponding predicted affinity Y; the molecular structure X includes multiple atoms x i , where 1 ≤ atomic index i ≤ M, and M is the total number of atoms in the molecular structure X; each atom x i has atomic parameters including atomic type s i and atomic three-dimensional coordinates c i ; the predicted affinity Y is a real number;

[0006] Through a variety of preset public data collection channels, collect big data on the molecular structures and corresponding molecular affinity values of organic molecules used in organic photovoltaic materials; and use each collected molecular structure and the corresponding molecular affinity value as the corresponding first training molecular structure and the first labeled affinity to form a corresponding first data record; and form a corresponding first data set from all the obtained first data records; the variety of public data collection channels at least include multiple public molecular information databases and multiple types of public technical literatures in the field of organic photovoltaic materials; the first data set includes multiple first data records; the first data record includes the first training molecular structure and the first labeled affinity; the first training molecular structure includes multiple first atoms; the atomic parameters of each first atom include the first atomic type and the first atomic three-dimensional coordinates.

[0007] Based on the first data set, train the molecular affinity prediction model;

[0008] After the model training is completed, receive the first molecular structure input by the user; and use the first molecular structure as the corresponding molecular structure X to input into the molecular affinity prediction model for prediction processing, and use the predicted affinity Y output by the model in this processing as the corresponding first molecular affinity; and analyze the contribution scores of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atomic contribution scores; and feedback the obtained first molecular affinity and all the first atomic contribution scores to the current user; the first molecular structure includes multiple second atoms; the atomic parameters of each second atom include the second atomic type and the second atomic three-dimensional coordinates.

[0009] Preferably, the molecular affinity prediction model includes an embedding encoding module, the Uni-Mol model, and the MLP model;

[0010] The input end of the embedding encoding module is connected to the model input end of the molecular affinity prediction model, and the output end is connected to the input end of the Uni-Mol model; the output end of the Uni-Mol model is connected to the input end of the MLP model; the output end of the MLP model is connected to the output end of the molecular affinity prediction model;

[0011] The embedding encoding module is used to encode each atom x of the molecular structure X according to the one-hot encoding rule of atom types of the Uni-Mol model i to obtain the corresponding atom one-hot encoding vector es i , and all the obtained atom one-hot encoding vectors es i are used to form the corresponding atom type encoding tensor ES; and according to the pairwise feature initialization encoding rule of the Uni-Mol model, pairwise feature initialization is performed based on all the three-dimensional coordinates c of the atoms of the molecular structure X i to obtain the corresponding pairwise encoding tensor EP; and the atom type encoding tensor ES and the pairwise encoding tensor EP are sent to the Uni-Mol model;

[0012] The atom one-hot encoding vector es i is a one-dimensional one-hot encoding vector with a vector length of D1, where D1 is the total number of atom types of the molecular structure X; in the atom one-hot encoding vector es i , the one-hot encoding corresponding to the current atom type s of the atom x i is set to 1, and the remaining D1 - 1 one-hot encodings are set to 0; the tensor shape of the atom type encoding tensor ES is D1×M; i The pairwise encoding tensor EP is a pairwise encoding matrix with a shape of M×M; the pairwise encoding matrix is composed of M×M encoding matrix units; each encoding matrix unit is a pairwise encoding vector with a vector length of D2, where D2 is a preset encoding vector length; the matrix row or matrix column of the pairwise encoding matrix corresponds one-to-one with the atom x

[0013] i i one by one;

[0014] The Uni-Mol model is used to perform atom feature and pairwise feature extraction processing based on the input atom type encoding tensor ES and pairwise encoding tensor EP to obtain the corresponding atom feature tensor Q and pairwise feature tensor P; and perform feature fusion on the atom feature tensor Q and the pairwise feature tensor P to obtain the corresponding molecular feature tensor H and send it to the MLP model;

[0015] The tensor shape of the atomic feature tensor Q is D3×M, where D3 is the preset atomic feature dimension; the atomic feature tensor Q is specifically composed of M sub-feature vectors q with a vector length of D3 each i which i correspond one-to-one with the atom x i ;

[0016] The pairwise feature tensor P is a pairwise feature matrix with a shape of M×M; the pairwise feature matrix is composed of M×M feature matrix units; each feature matrix unit is a pairwise feature vector with a vector length of D4, where D4 is the preset pairwise feature dimension; the matrix rows or columns of the pairwise feature matrix correspond one-to-one with the atom x i ; the feature vectors corresponding to each matrix row of the pairwise feature matrix are represented as row feature vectors pl i whose i vector length is D4×M and is formed by sequentially concatenating the M pairwise feature vectors in the current row;

[0017] The molecular feature tensor H has a tensor shape of (D3 + D4×M)×M; the molecular feature tensor H is specifically composed of M sub-feature vectors h with a vector length of (D3 + D4×M) each i which i correspond one-to-one with the atom x i ; the sub-feature vector h i is formed by sequentially concatenating the corresponding sub-feature vector q with a length of D3 i and the row feature vector pl with a length of D4×M i ;

[0018] The MLP model is used to calculate the molecular affinity based on the input molecular feature tensor H and output the corresponding predicted affinity Y;

[0019] The MLP model is specifically composed of fully connected layers FC connected in sequence j where 1 ≤ layer index j ≤ N, N is the preset total number of fully connected layers, and N ≥ 3;

[0020] The function expression of the fully connected layer FC j is:

[0021]

[0022] When j = 1, y j-1 = y0 = H,

[0023] When j = N, Y = y j=N ;

[0024] Among them, y j-1 and y j are the output features of the (j - 1)-th and j-th fully connected layers; w j and b j are the weight parameters and bias parameters of each of the fully connected layers FC j ; σ is a preset activation function, which is only used in the first to (N - 1)-th fully connected layers; for the first fully connected layer FC j=1 , the output feature y j-1 of its corresponding upper layer is the molecular feature tensor H output by the Uni-Mol model; for the N-th fully connected layer FC j=N , the output feature of the current fully connected layer is a real number, i.e., the predicted affinity Y.

[0025] Preferably, the model training of the molecular affinity prediction model based on the first data set specifically includes:

[0026] Step 31, splitting the first data set into two sub-data sets denoted as the corresponding first training set and first evaluation set based on a preset first splitting ratio;

[0027] Among them, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first splitting ratio;

[0028] Step 32, taking the first first data record in the first training set as the corresponding current training record;

[0029] Step 33, inputting the first training molecular structure of the current training record as the corresponding molecular structure X into the molecular affinity prediction model for prediction processing, and taking the predicted affinity Y output by the model in this processing as the corresponding first predicted affinity;

[0030] Step 34, bringing the first predicted affinity and the first labeled affinity of the current training record into a preset first model loss function for calculation to obtain the corresponding first loss value;

[0031] Among them, the first model loss function is implemented based on the L1 loss function, smooth L1 loss function or L2 loss function;

[0032] Step 35: Identify whether the first loss value meets a preset first loss value range; if the first loss value meets the first loss value range, identify whether the current training record is the last first data record of the first training set. If so, go to Step 36; if not, use the next first data record of the first training set as the new current training record and return to Step 33 to continue training; if the first loss value does not meet the first loss value range, based on a preset first model optimizer, modulate the model parameters of the molecular affinity prediction model in the direction of minimizing the first model loss function for one round, and return to Step 33 to continue training at the end of this round of modulation;

[0033] Among them, the first model optimizer at least includes the Adam optimizer and the SGD optimizer;

[0034] Step 36: Conduct one round of traversal of all the first data records in the first evaluation set; during this round of traversal, use the currently traversed first data record as the corresponding current evaluation record; use the first training molecular structure of the current evaluation record as the corresponding molecular structure X and input it into the molecular affinity prediction model for prediction processing, and use the predicted affinity Y output by the model in this processing as the corresponding second predicted affinity; form a corresponding prediction-label pair from the second predicted affinity and the first label affinity of the current training record; at the end of this round of traversal, input all the obtained prediction-label pairs into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value;

[0035] Among them, the first model evaluation function is implemented based on the MAE function, the MSE function or the RMSE function;

[0036] Step 37: Identify whether the first evaluation value meets a preset first evaluation value range; if it does not meet, return to Step 32 to continue training; if it meets, stop training and confirm that the model training is completed.

[0037] Preferably, the analysis of the contribution scores of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atom contribution score specifically includes:

[0038] Step 41: Set a first counter initialized to 1; and set the corresponding first counting threshold to a preset number threshold K;

[0039] Step 42: Take the first molecular structure as the corresponding molecular structure X and input it into the embedding encoding module of the molecular affinity prediction model for processing to obtain the corresponding atomic type encoding tensor ES and the pairwise encoding tensor EP; input the atomic type encoding tensor ES and the pairwise encoding tensor EP obtained this time into the Uni-Mol model of the molecular affinity prediction model for processing to obtain the corresponding molecular feature tensor H; and take the molecular feature tensor H obtained this time as the corresponding current feature tensor H now ; and denote each sub-feature vector h of the molecular feature tensor H obtained this time i as the corresponding sub-feature vector h now,I ;

[0040] Step 43: Input the current feature tensor H now into the MLP model of the molecular affinity prediction model for processing and take the predicted affinity output by this processing as a corresponding affinity truth value z true,k ; 1 ≤ index k ≤ K;

[0041] Step 44: Create a random noise vector ε that follows a Gaussian distribution with a mean of 0, a standard deviation of a preset first standard deviation, and a vector shape that is consistent with the vector shape of the sub-feature vector h of the molecular feature tensor H i ; k ;

[0042] Step 45: Perform a round of traversal on all sub-feature vectors h of the current feature tensor H now ; and during this round of traversal, take the currently traversed sub-feature vector h now,I as the corresponding current vector h; and perform noise perturbation on the current vector h in the manner of h now,I = h + ε * to obtain the corresponding current perturbed vector h k ; and based on the current perturbed vector h * replace the corresponding sub-feature vector h in the current feature tensor H * to obtain a new feature tensor H now ; input the current feature tensor H now,I into the MLP model of the molecular affinity prediction model for processing and take the predicted affinity output by this processing as a corresponding perturbed affinity z * ; and at the end of this round of traversal, increment the first counter by 1; * ; i,k ; and at the end of this round of traversal, increment the first counter by 1;

[0043] Step 46: Identify whether the first counter is greater than the first counting threshold. If so, go to Step 47; if not, return to Step 42;

[0044] Step 47: Based on all the obtained true affinity values z true,k and all the perturbed affinity values z i,k estimate the contribution scores of each of the second atoms in the first molecular structure to obtain the corresponding first atomic contribution scores;

[0045] The estimation method of the first atomic contribution score is as follows:

[0046]

[0047] where score i is the first atomic contribution score corresponding to the i-th second atom in the first molecular structure.

[0048] In the second aspect of the embodiments of the present invention, there is provided an apparatus for implementing the processing method for molecular affinity analysis described in the first aspect above. The apparatus includes: a model construction module, a data acquisition module, a model training module, and a model application module;

[0049] The model construction module is used to construct a molecular affinity prediction model in a manner that takes the Uni-Mol model as the encoder and the MLP model as the regression calculation task model downstream of the encoder. The molecular affinity prediction model is used to perform molecular affinity prediction processing on the input molecular structure X of the model and output the corresponding predicted affinity Y. The molecular structure X includes multiple atoms x i , where 1 ≤ atomic index i ≤ M, and M is the total number of atoms in the molecular structure X; the atomic parameters of each atom x i include the atomic type s i and the atomic three-dimensional coordinates c i ; the predicted affinity Y is a real number;

[0050] The data acquisition module is used to collect big data on the molecular structures of organic molecules for organic photovoltaic materials and the corresponding molecular affinity values through a variety of preset public data acquisition channels; and take each collected molecular structure and the corresponding molecular affinity value as a corresponding first training molecular structure and a first labeled affinity to form a corresponding first data record; and form a corresponding first data set from all the obtained first data records; the variety of public data acquisition channels at least include multiple public molecular information databases and multiple types of public technical literatures in the field of organic photovoltaic materials; the first data set includes multiple first data records; the first data record includes the first training molecular structure and the first labeled affinity; the first training molecular structure includes multiple first atoms; the atomic parameters of each first atom include the first atom type and the first atomic three-dimensional coordinates.

[0051] The model training module is used to perform model training on the molecular affinity prediction model based on the first data set.

[0052] The model application module is used to, after the model training is completed, receive the first molecular structure input by the user; and take the first molecular structure as the corresponding molecular structure X and input it into the molecular affinity prediction model for prediction processing, and take the predicted affinity Y output by the model in this processing as the corresponding first molecular affinity; and analyze the contribution scores of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atomic contribution scores; and feedback the obtained first molecular affinity and all the first atomic contribution scores to the current user; the first molecular structure includes multiple second atoms; the atomic parameters of each second atom include the second atom type and the second atomic three-dimensional coordinates.

[0053] A third aspect of the embodiments of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;

[0054] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;

[0055] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

[0056] A fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the computer is enabled to execute the instructions of the method described in the first aspect above.

[0057] An embodiment of the present invention provides a processing method, device, electronic device, and computer-readable storage medium for molecular affinity analysis. As can be seen from the above, an embodiment of the present invention constructs a molecular affinity prediction model by using the Uni-Mol model as the encoder and the MLP model as the regression calculation task model downstream of the encoder. The molecular affinity prediction model is used to perform molecular affinity prediction processing based on the molecular structure input to the model and output the corresponding predicted affinity; and a first data set for model training is constructed through big data collection, and the molecular affinity prediction model is trained based on the first data set; and after the model training is completed, the molecular affinity prediction model is used to predict the molecular affinity of the molecular structure specified by the user, and the affinity contribution scores of each atom in the current molecular structure are estimated and analyzed by scrambling the process features generated during the model processing. Through the embodiment of the present invention, on the one hand, two affinity analysis tasks for organic photovoltaic material molecules are realized (predicting the overall molecular affinity and analyzing the atomic-level affinity contribution), and on the other hand, the analysis complexity is reduced, the analysis time is shortened, and the analysis efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 FIG. is a schematic diagram of a processing method for molecular affinity analysis provided in Embodiment 1 of the present invention;

[0059] Figure 2 FIG. is a module structure diagram of the molecular affinity prediction model provided in Embodiment 1 of the present invention;

[0060] Figure 3 FIG. is a module structure diagram of a processing device for molecular affinity analysis provided in Embodiment 2 of the present invention;

[0061] Figure 4 FIG. is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] Embodiment 1 of the present invention provides a processing method for molecular affinity analysis, as Figure 1 shown in the schematic diagram of a processing method for molecular affinity analysis provided in Embodiment 1 of the present invention, the method mainly includes the following steps:

[0064] Step 1, construct a molecular affinity prediction model by using the Uni-Mol model as the encoder and the MLP model as the regression calculation task model downstream of the encoder.

[0065] Here, the molecular affinity prediction model of the embodiment of the present invention is used to perform molecular affinity prediction processing according to the molecular structure X input to the model and output the corresponding predicted affinity Y; among them, the molecular structure X includes multiple atoms x i , 1 ≤ atomic index i ≤ M, where M is the total number of atoms in the molecular structure X; each atom x i has atomic parameters including atomic type s i and atomic three-dimensional coordinates c i ; the predicted affinity Y is a real number.

[0066] The model structure of the molecular affinity prediction model is as Figure 2 shown in the module structure diagram of the molecular affinity prediction model provided in Embodiment 1 of the present invention, and includes: an embedding encoding module, a Uni-Mol model, and an MLP model.

[0067] The connection relationship of each model component of the molecular affinity prediction model is: the input end of the embedding encoding module is connected to the model input end of the molecular affinity prediction model, and the output end is connected to the input end of the Uni-Mol model; the output end of the Uni-Mol model is connected to the input end of the MLP model; the output end of the MLP model is connected to the output end of the molecular affinity prediction model.

[0068] The component functions of each model component of the molecular affinity prediction model are as follows.

[0069] 1) Embedding encoding module:

[0070] The embedding encoding module of the embodiment of the present invention is used to encode each atom x of the molecular structure X according to the one-hot encoding rule of the atomic type of the Uni-Mol model to obtain the corresponding atomic one-hot encoding vector es i , and form the corresponding atomic type encoding tensor ES from all the obtained atomic one-hot encoding vectors es i ; and according to the pairwise feature initialization encoding rule of the Uni-Mol model, perform pairwise feature initialization according to all the atomic three-dimensional coordinates c of the molecular structure X i to obtain the corresponding pairwise encoding tensor EP; and send the atomic type encoding tensor ES and the pairwise encoding tensor EP to the Uni-Mol model. i Here, the atomic one-hot encoding vector es of the embodiment of the present invention

[0071] is a one-dimensional one-hot encoding vector with a vector length of D1, where D1 is the total number of atomic types of the molecular structure X; the atomic one-hot encoding vector es i is a one-dimensional one-hot encoding vector with a vector length of D1, and D1 is the total number of atomic types of the molecular structure X; this atomic one-hot encoding vector esi The current atom x i Atom type s i The corresponding one-hot encoding is set to 1, and the remaining D1-1 one-hot encodings are set to 0. The tensor shape of the atomic type encoding tensor ES in the embodiment of the present invention is D1×M.

[0072] The paired coding tensor EP of the embodiment of the present invention is a paired coding matrix with a shape of M×M; the paired coding matrix is ​​composed of M×M coding matrix units; wherein each coding matrix unit is a paired coding vector with a vector length of D2, D2 is a preset coding vector length; the matrix rows or matrix columns of the paired coding matrix are consistent with the atomic x i One to one correspondence.

[0073] It should be noted that the Uni-Mol model used in the embodiment of the present invention is an encoder model implemented based on the Encoder sub-model of the Transformer model and used to encode atomic-level features of molecular structures. The public technical document A "Uni-Mol: A Universal 3D Molecular Representation Learning Framework" provides a clear description of the detailed model structure, model reasoning principle, atom type one-hot encoding rules, paired feature initialization encoding rules, and model pre-training scheme of the Uni-Mol model. Therefore, the atom type one-hot encoding rules and paired feature initialization encoding rules used in the embedded encoding module will not be further elaborated here.

[0074] 2) Uni-Mol model:

[0075] The Uni-Mol model of an embodiment of the present invention is used to perform atomic feature and paired feature extraction processing according to the input atomic type encoding tensor ES and paired encoding tensor EP to obtain the corresponding atomic feature tensor Q and paired feature tensor P; and perform feature fusion on the atomic feature tensor Q and the paired feature tensor P to obtain the corresponding molecular feature tensor H and send it to the MLP model.

[0076] Here, the tensor shape of the atomic feature tensor Q in the embodiment of the present invention is D3×M, where D3 is a preset atomic feature dimension; the atomic feature tensor Q can be specifically regarded as a sub-feature vector q with M vector lengths of D3. i Composition, sub-eigenvector q i With atom x i One to one correspondence.

[0077] The pairwise feature tensor P in the embodiments of the present invention is a pairwise feature matrix with a shape of M×M; this pairwise feature matrix is composed of M×M feature matrix units; each feature matrix unit is a pairwise feature vector with a vector length of D4, where D4 is a preset pairwise feature dimension; the matrix rows or matrix columns of this pairwise feature matrix correspond to the atom x i in a one-to-one manner; the feature vector corresponding to each matrix row of this pairwise feature matrix can be expressed as a corresponding row feature vector pl i , and each row feature vector pl i has a vector length of D4×M and is sequentially concatenated by M feature matrix units, that is, M pairwise feature vectors, within the current row of the pairwise feature matrix.

[0078] The tensor shape of the molecular feature tensor H in the embodiments of the present invention is (D3 + D4×M)×M; this molecular feature tensor H can be specifically regarded as being composed of M sub-feature vectors h i each with a vector length of (D3 + D4×M); the sub-feature vectors h i correspond to the atom x i in a one-to-one manner; each sub-feature vector h i is sequentially concatenated by a corresponding sub-feature vector q i (with a length of D3) and a corresponding row feature vector pl i (with a length of D4×M).

[0079] As can be seen from the foregoing, the detailed model structure, model inference principle, and model pre-training scheme of the Uni-Mol model used in the embodiments of the present invention have been published in the public technical literature A, so the encoding process of the Uni-Mol model will not be further elaborated here. However, it should be noted that the Uni-Mol model used in the embodiments of the present invention has been pre-trained in advance according to the model pre-training scheme given in the public technical literature A.

[0080] 3) MLP model:

[0081] The MLP model in the embodiments of the present invention is used to calculate the molecular affinity based on the input molecular feature tensor H and output the corresponding predicted affinity Y.

[0082] As Figure 2 shown, this MLP model is specifically composed of fully connected layers FC j connected in sequence, where 1 ≤ layer index j ≤ N, N is the total number of preset fully connected layers, and N≥3.

[0083] The function expression of the fully connected layer FC j of this MLP model is as follows:

[0084]

[0085] When j = 1, y j-1 = y0 = H,

[0086] When j = N, Y = y j=N ;

[0087] where y j-1 and y j are the output features of the (j - 1)-th and j-th fully connected layers; w j and b j are the weight parameters and bias parameters of each fully connected layer FC j ; σ is a preset activation function, which is only used in the first to (N - 1)-th fully connected layers. This activation function σ can be a linear activation function or a non-linear activation function. In general, a non-linear activation function is often used; for the first fully connected layer FC j=1 , the corresponding output feature y j-1 of its previous layer is the molecular feature tensor H output by the Uni-Mol model; for the N-th fully connected layer FC j=N , the output feature of the current fully connected layer is a real number, that is, the predicted affinity Y.

[0088] Step 2: Through a variety of preset public data collection channels, collect big data on the molecular structures of organic molecules used in organic photovoltaic materials and their corresponding molecular affinity values; and take each collected molecular structure and its corresponding molecular affinity value as the corresponding first training molecular structure and first label affinity to form a corresponding first data record; and form a corresponding first data set from all the obtained first data records.

[0089] Here, the variety of public data collection channels in the embodiments of the present invention at least include multiple public molecular information databases and multiple types of public technical literatures in the field of organic photovoltaic materials. The first data set used for model training in the embodiments of the present invention includes multiple first data records; the first data record includes a first training molecular structure and a first label affinity; the first training molecular structure includes multiple first atoms; the atomic parameters of each first atom include a first atom type and a first atomic three-dimensional coordinate.

[0090] Step 3: Perform model training on the molecular affinity prediction model based on the first data set;

[0091] Specifically, it includes: Step 31: Divide the first data set into two sub-data sets according to a preset first splitting ratio, denoted as the corresponding first training set and first evaluation set;

[0092] Here, the first splitting ratio of the embodiments of the present invention is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first splitting ratio.

[0093] Step 32: Take the first first data record in the first training set as the corresponding current training record.

[0094] Step 33: Take the first training molecular structure of the current training record as the corresponding molecular structure X and input it into the molecular affinity prediction model for prediction processing, and take the predicted affinity Y output by the model in this processing as the corresponding first predicted affinity.

[0095] Step 34: Substitute the first predicted affinity and the first labeled affinity of the current training record into the preset first model loss function for calculation to obtain the corresponding first loss value.

[0096] Here, the first model loss function of the embodiments of the present invention is implemented based on the L1 loss function, the smooth L1 loss function or the L2 loss function.

[0097] Step 35: Identify whether the first loss value meets the preset first loss value range; if the first loss value meets the first loss value range, identify whether the current training record is the last first data record in the first training set. If so, go to step 36; if not, take the next first data record in the first training set as the new current training record and return to step 33 to continue training; if the first loss value does not meet the first loss value range, based on the preset first model optimizer, perform a round of modulation on the model parameters of the molecular affinity prediction model in the direction of minimizing the first model loss function, and return to step 33 to continue training at the end of this round of modulation.

[0098] Here, the first loss value range of the embodiments of the present invention is a preset loss value range; the first model optimizer includes at least the Adam optimizer and the SGD optimizer.

[0099] Step 36: Perform a round of traversal on all the first data records in the first evaluation set; during this round of traversal, take the currently traversed first data record as the corresponding current evaluation record; take the first training molecular structure of the current evaluation record as the corresponding molecular structure X and input it into the molecular affinity prediction model for prediction processing, and take the predicted affinity Y output by the model in this processing as the corresponding second predicted affinity; form a corresponding prediction-label pair from the second predicted affinity and the first labeled affinity of the current training record; at the end of this round of traversal, substitute all the obtained prediction-label pairs into the preset first model evaluation function for calculation to obtain the corresponding first evaluation value.

[0100] Here, the first model evaluation function of the embodiment of the present invention is implemented based on the MAE function, the MSE function or the RMSE function;

[0101] Step 37, identify whether the first evaluation value meets a preset first evaluation value range; if not, return to step 32 to continue training; if so, stop training and confirm that the model training is completed.

[0102] Here, the first evaluation value range of the embodiment of the present invention is a preset evaluation value range.

[0103] Step 4, after the model training is completed, receive the first molecular structure input by the user; and use the first molecular structure as the corresponding molecular structure X to input into the molecular affinity prediction model for prediction processing, and use the predicted affinity Y output by the model in this processing as the corresponding first molecular affinity; and analyze the contribution scores of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atomic contribution scores; and feedback the obtained first molecular affinity and all first atomic contribution scores to the current user;

[0104] Specifically, it includes: Step 41, after the model training is completed, receive the first molecular structure input by the user;

[0105] Here, the first molecular structure of the embodiment of the present invention includes a plurality of second atoms; the atomic parameters of each second atom include the second atom type and the three-dimensional coordinates of the second atom;

[0106] Step 2, and use the first molecular structure as the corresponding molecular structure X to input into the molecular affinity prediction model for prediction processing, and use the predicted affinity Y output by the model in this processing as the corresponding first molecular affinity;

[0107] Step 43, and analyze the contribution scores of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atomic contribution scores;

[0108] Specifically, it includes: Step 431, set a first counter initialized to 1; and set the corresponding first counting threshold to a preset number threshold K;

[0109] Here, the number threshold K of the embodiment of the present invention is a preset positive integer greater than 1;

[0110] Step 432: Take the first molecular structure as the corresponding molecular structure X and input it into the embedding encoding module of the molecular affinity prediction model to obtain the corresponding atomic type encoding tensor ES and pairwise encoding tensor EP; then input the obtained atomic type encoding tensor ES and pairwise encoding tensor EP into the Uni-Mol model of the molecular affinity prediction model for processing to obtain the corresponding molecular feature tensor H; and take the obtained molecular feature tensor H as the corresponding current feature tensor H now ; and denote each sub-feature vector h i of the obtained molecular feature tensor H as the corresponding sub-feature vector h now,I ;

[0111] Step 433: Input the current feature tensor H now into the MLP model of the molecular affinity prediction model for processing, and take the predicted affinity output by this processing as a corresponding affinity true value z true,k ; 1 ≤ index k ≤ K;

[0112] Step 434: Create a random noise vector ε that follows a Gaussian distribution with a mean of 0, a standard deviation of a preset first standard deviation, and a vector shape consistent with the vector shape of the sub-feature vector h of the molecular feature tensor H i ; k ;

[0113] Here, the first standard deviation is a preset standard deviation parameter;

[0114] Step 435: Conduct a round of traversal on all sub-feature vectors h now of the current feature tensor H now,I ; and during this round of traversal, take the currently traversed sub-feature vector h now,I as the corresponding current vector h; and based on h * = h + ε k perturb the current vector h with noise to obtain the corresponding current perturbed vector h * ; and based on the current perturbed vector h * replace the corresponding sub-feature vector h now in the current feature tensor H now,I to obtain a new feature tensor H * ; then input the current feature tensor H * into the MLP model of the molecular affinity prediction model for processing, and take the predicted affinity output by this processing as a corresponding perturbed affinity z i,k ; and at the end of this round of traversal, increment the first counter by 1;

[0115] Step 436: Identify whether the first counter is greater than the first counting threshold; if so, go to Step 437; if not, return to Step 432;

[0116] Step 437: Based on all the obtained true affinity values z true,k and all the perturbed affinity values z i,k estimate the contribution scores of each second atom in the first molecular structure to obtain the corresponding first atom contribution scores;

[0117] Here, the estimation method of the first atom contribution score in the embodiment of the present invention is:

[0118]

[0119] where score i is the first atom contribution score corresponding to the i-th second atom in the first molecular structure;

[0120] As can be seen from the above Steps 431 - 437, the embodiment of the present invention provides an analysis method for estimating the affinity contribution scores of each atom in the current molecular structure by scrambling the atomic-level features (i.e., the sub-feature vectors h i ), and performs K rounds of estimation during the processing to reduce the estimation error, and finally takes the average of the K estimation scores of each atom as the final estimation score (i.e., the first atom contribution score);

[0121] Step 44: Feed back the obtained first molecular affinity and all the first atom contribution scores to the current user.

[0122] Here, after the user obtains the feedback information, on the one hand, the user can intuitively understand the molecular affinity level of the current molecular structure based on the first molecular affinity; on the other hand, the user can intuitively understand the affinity contribution levels of each atom and each molecular fragment region in the current molecular structure based on all the first atom contribution scores, which also provides reference information for the user to further optimize the current molecular structure.

[0123] Figure 3 FIG. is a module structure diagram of a processing device for molecular affinity analysis provided by the second embodiment of the present invention. This device is a terminal device or a server for implementing the foregoing method embodiment, or can also be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, this device can be a device or a chip system of the foregoing terminal device or server. As Figure 3 shown, this device includes: a model construction module 201, a data acquisition module 202, a model training module 203, and a model application module 204.

[0124] The model construction module 201 is used to construct a molecular affinity prediction model by using the Uni-Mol model as the encoder and the MLP model as the regression calculation task model downstream of the encoder; the molecular affinity prediction model is used to perform molecular affinity prediction processing on the input molecular structure X of the model and output the corresponding predicted affinity Y; the molecular structure X includes multiple atoms x i , where 1 ≤ atomic index i ≤ M, and M is the total number of atoms in the molecular structure X; each atom x i has atomic parameters including atomic type s i and atomic three-dimensional coordinates c i ; the predicted affinity Y is a real number.

[0125] The data acquisition module 202 is used to collect big data on the molecular structures and corresponding molecular affinity values of organic molecules for organic photovoltaic materials through a variety of preset public data acquisition channels; and use the collected molecular structures and corresponding molecular affinity values as the corresponding first training molecular structures and first labeled affinities to form a corresponding first data record; and form a corresponding first data set from all the obtained first data records; the variety of public data acquisition channels at least include multiple public molecular information databases and multiple types of public technical literatures in the field of organic photovoltaic materials; the first data set includes multiple first data records; the first data record includes the first training molecular structure and the first labeled affinity; the first training molecular structure includes multiple first atoms; the atomic parameters of each first atom include the first atomic type and the first atomic three-dimensional coordinates.

[0126] The model training module 203 is used to train the molecular affinity prediction model based on the first data set.

[0127] The model application module 204 is used to, after the model training is completed, receive the first molecular structure input by the user; and use the first molecular structure as the corresponding molecular structure X to input into the molecular affinity prediction model for prediction processing and use the predicted affinity Y output by the model in this processing as the corresponding first molecular affinity; and analyze the contribution scores of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atomic contribution scores; and feedback the obtained first molecular affinity and all the first atomic contribution scores to the current user; the first molecular structure includes multiple second atoms; the atomic parameters of each second atom include the second atomic type and the second atomic three-dimensional coordinates.

[0128] The processing device for molecular affinity analysis provided by the embodiments of the present invention can execute the method steps in the above method embodiments, and its implementation principle and technical effects are similar, and will not be elaborated here.

[0129] It should be noted that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the model construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determined module can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through the integrated logic circuit in the processor element or the instruction in the form of software.

[0130] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).

[0131] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the above computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0132] Figure 4 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device can be a terminal device or a server for implementing the method of the foregoing embodiments, or a terminal device or a server for implementing the method of the foregoing embodiments connected to the foregoing terminal device or server. As Figure 4 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 to complete various processing functions and implement the processing steps described in the foregoing method embodiments. Preferably, the electronic device related to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0133] In Figure 4The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0134] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0135] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods and processing procedures provided in the above embodiments.

[0136] An embodiment of the present invention provides a processing method, device, electronic device, and computer-readable storage medium for molecular affinity analysis. As can be seen from the above content, in the embodiment of the present invention, a molecular affinity prediction model is constructed by using the Uni-Mol model as the encoder and the MLP model as the regression calculation task model downstream of the encoder. The molecular affinity prediction model is used to perform molecular affinity prediction processing based on the molecular structure input to the model and output the corresponding predicted affinity; and a first data set for model training is constructed through big data collection, and the molecular affinity prediction model is trained based on the first data set; and after the model training is completed, the molecular affinity prediction model is used to predict the molecular affinity of the molecular structure specified by the user, and the affinity contribution scores of each atom in the current molecular structure are estimated and analyzed by scrambling the process features generated during the model processing. Through the embodiment of the present invention, on the one hand, two affinity analysis tasks for organic photovoltaic material molecules are realized (predicting the overall molecular affinity and analyzing the atomic-level affinity contribution), and on the other hand, the analysis complexity is reduced, the analysis duration is shortened, and the analysis efficiency is improved.

[0137] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the art.

[0138] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A processing method for molecular affinity analysis, characterized in that The method comprises: A molecular affinity prediction model is constructed by using the Uni-Mol model as an encoder and the MLP model as a downstream regression calculation task model of the encoder; the molecular affinity prediction model is used to perform molecular affinity prediction processing according to the molecular structure X input by the model and output the corresponding predicted affinity Y; the molecular structure X includes multiple atoms x i , 1≤atom index i≤M, M is the total number of atoms in the molecular structure X; each atom x i The atomic parameters include the atom type s i and the atomic three-dimensional coordinates c i ; The predicted affinity Y is a real number; Through a plurality of preset public data collection channels, the molecular structures of organic molecules used for organic photovoltaic materials and the corresponding molecular affinity values ​​are collected for big data; and each collected molecular structure and the corresponding molecular affinity value are used as a corresponding first training molecular structure and a first label affinity to form a corresponding first data record; and all the first data records obtained form a corresponding first data set; the plurality of public data collection channels include at least a plurality of public molecular information repositories and a plurality of types of public technical literature in the field of organic photovoltaic materials; the first data set includes a plurality of the first data records; the first data record includes the first training molecular structure and the first label affinity; the first training molecular structure includes a plurality of first atoms; the atomic parameters of each of the first atoms include a first atom type and a first atom three-dimensional coordinate; Performing model training on the molecular affinity prediction model based on the first data set; After the model training is completed, the first molecular structure input by the user is received; the first molecular structure is input as the corresponding molecular structure X into the molecular affinity prediction model for prediction processing, and the predicted affinity Y output by the model in this processing is used as the corresponding first molecular affinity; the contribution score of each atom in the first molecular structure to the molecular affinity is analyzed to obtain the corresponding first atomic contribution score; and the obtained first molecular affinity and all the first atomic contribution scores are fed back to the current user; the first molecular structure includes multiple second atoms; the atomic parameters of each second atom include a second atom type and a second atom three-dimensional coordinate.

2. The method for molecular affinity analysis according to claim 1, characterized in that: The molecular affinity prediction model includes an embedded coding module, the Uni-Mol model and an MLP model; The input end of the embedded coding module is connected to the model input end of the molecular affinity prediction model, and the output end is connected to the input end of the Uni-Mol model; the output end of the Uni-Mol model is connected to the input end of the MLP model; the output end of the MLP model is connected to the output end of the molecular affinity prediction model; The embedded coding module is used to encode each atom x of the molecular structure X according to the one-hot encoding rule of the atom type of the Uni-Mol model. i Encode to get the corresponding atomic one-hot encoding vector es i , and the obtained one-hot encoding vectors es of all the atoms are i The atomic type encoding tensor ES corresponding to the composition; According to the paired feature initialization encoding rule of the Uni-Mol model, all the three-dimensional coordinates c of the atoms in the molecular structure X are i Performing paired feature initialization to obtain a corresponding paired encoding tensor EP; and sending the atom type encoding tensor ES and the paired encoding tensor EP to the Uni-Mol model; The atom one-hot encoding vector es i is a one-dimensional one-hot encoding vector with a vector length of D1, where D1 is the total number of atomic types in the molecular structure X; The atom one-hot encoding vector es i The atom x in the i The atom type s i The corresponding one-hot encoding is set to 1, and the remaining D1-1 one-hot encodings are set to 0; the tensor shape of the atom type encoding tensor ES is D1×M; The pairwise encoding tensor EP is a pairwise encoding matrix with a shape of M×M; the pairwise encoding matrix is ​​composed of M×M encoding matrix units; each encoding matrix unit is a pairwise encoding vector with a vector length of D2, where D2 is a preset encoding vector length; the matrix rows or matrix columns of the pairwise encoding matrix are consistent with the atomic x i One to one correspondence; The Uni-Mol model is used to perform atomic feature and paired feature extraction processing according to the input atomic type encoding tensor ES and the paired encoding tensor EP to obtain the corresponding atomic feature tensor Q and paired feature tensor P; and perform feature fusion on the atomic feature tensor Q and the paired feature tensor P to obtain the corresponding molecular feature tensor H and send it to the MLP model; The tensor shape of the atomic feature tensor Q is D3×M, where D3 is the preset atomic feature dimension; the atomic feature tensor Q is specifically composed of M sub-feature vectors q whose vector lengths are all D3. i The sub-eigenvector q i With the atom x i One to one correspondence; The paired feature tensor P is a paired feature matrix with a shape of M×M; the paired feature matrix is ​​composed of M×M feature matrix units; each of the feature matrix units is a paired feature vector with a vector length of D4, where D4 is a preset paired feature dimension; the matrix rows or matrix columns of the paired feature matrix are the same as the atomic x i One-to-one correspondence; the eigenvector corresponding to each matrix row of the paired feature matrix is ​​represented as the row eigenvector pl i , the row feature vector pl i The length of the vector is D4×M, which is formed by sequentially concatenating the M pairs of feature vectors in the current row; The tensor shape of the molecular feature tensor H is (D3+D4×M)×M; the molecular feature tensor H is specifically composed of M sub-feature vectors h whose vector lengths are all (D3+D4×M) i The sub-feature vector h i With the atom x i One-to-one correspondence; the sub-feature vector h i The corresponding sub-feature vector q of length D3 i and the row feature vector pl of length D4×M i Sequential splicing; The MLP model is used to calculate the molecular affinity according to the input molecular feature tensor H and output the corresponding predicted affinity Y; The MLP model specifically consists of multiple layers of sequentially connected fully connected layers FC j Composition, 1≤layer index j≤N, N is the total number of preset fully connected layers, N≥3; The fully connected layer FC j The function expression is: When j=1, y j-1 =y0=H, When j = N, Y = y j=N ; Among them, y j-1 ,y j is the output feature of the j-1th and jth fully connected layers; w j , b j For each of the fully connected layers FC j The weight parameters and bias parameters of ; σ is a type of preset activation function, which is only used in the 1st to N-1th fully connected layers; for the 1st fully connected layer FC j=1 For example, the corresponding previous layer output feature y j-1 It is the molecular feature tensor H output by the Uni-Mol model; for the Nth fully connected layer FC j=N In terms of, the output feature of the current fully connected layer is a real number, namely the predicted affinity Y.

3. The processing method for molecular affinity analysis according to claim 1, characterized in that: The performing model training on the molecular affinity prediction model based on the first data set specifically includes: Step 31, based on a preset first segmentation ratio, the first data set is divided into two sub-data sets which are recorded as a corresponding first training set and a first evaluation set; Wherein, both the first training set and the first evaluation set are composed of a plurality of the first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 32, taking the first first data record of the first training set as the corresponding current training record; Step 33, inputting the first training molecular structure of the current training record as the corresponding molecular structure X into the molecular affinity prediction model for prediction processing and using the predicted affinity Y output by the model in this processing as the corresponding first predicted affinity; Step 34, bringing the first predicted affinity and the first label affinity of the current training record into a preset first model loss function to calculate and obtain a corresponding first loss value; Wherein, the first model loss function is implemented based on an L1 loss function, a smoothed L1 loss function or an L2 loss function; Step 35, identifying whether the first loss value satisfies a preset first loss value range; if the first loss value satisfies the first loss value range, identifying whether the current training record is the last first data record of the first training set, if so, turning to step 36, if not, taking the next first data record of the first training set as the new current training record and returning to step 33 to continue training; if the first loss value does not satisfy the first loss value range, performing a round of modulation on the model parameters of the molecular affinity prediction model in a direction to minimize the first model loss function based on a preset first model optimizer, and returning to step 33 to continue training at the end of this round of modulation; Wherein, the first model optimizer includes at least an Adam optimizer and an SGD optimizer; Step 36, perform a round of traversal on all the first data records of the first evaluation set; and in this round of traversal, use the first data record currently traversed as the corresponding current evaluation record; and input the first training molecular structure of the current evaluation record as the corresponding molecular structure X into the molecular affinity prediction model for prediction processing, and use the predicted affinity Y output by the model this time as the corresponding second predicted affinity; and form a corresponding prediction-label pair by the second predicted affinity and the first label affinity of the current training record; and at the end of this round of traversal, bring all the obtained prediction-label pairs into the preset first model evaluation function to calculate and obtain the corresponding first evaluation value; Wherein, the first model evaluation function is implemented based on the MAE function, the MSE function or the RMSE function; Step 37, identifying whether the first evaluation value meets the preset first evaluation value range; if not, returning to step 32 to continue training; if satisfied, stopping training and confirming the end of model training.

4. The method for molecular affinity analysis according to claim 2, characterized in that: The step of analyzing the contribution score of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atom contribution score specifically includes: Step 41, setting a first counter initialized to 1; and setting the corresponding first counting threshold to a preset number threshold K; Step 42: input the first molecular structure as the corresponding molecular structure X into the embedded coding module of the molecular affinity prediction model for processing to obtain the corresponding atom type encoding tensor ES and the paired encoding tensor EP; and input the atom type encoding tensor ES and the paired encoding tensor EP obtained this time into the Uni-Mol model of the molecular affinity prediction model for processing to obtain the corresponding molecular feature tensor H; and use the molecular feature tensor H obtained this time as the corresponding current feature tensor H now ; And each of the sub-feature vectors h of the molecular feature tensor H obtained this time i Denoted as the corresponding sub-eigenvector h now,I ; Step 43: convert the current feature tensor H now The MLP model of the molecular affinity prediction model is input for processing and the predicted affinity output by this processing is used as a corresponding affinity true value z true,k ; 1≤indexk≤K; Step 44, create a sub-feature vector h that satisfies Gaussian distribution and has a mean of 0, a standard deviation of a preset first standard deviation, and a vector shape that is the same as the molecular feature tensor H i The vector shape remains consistent with the random noise vector ε k ; Step 45: the current feature tensor H now All the sub-eigenvectors h now,I Perform a round of traversal; and in this round of traversal, the sub-feature vector h currently traversed now,I as the corresponding current vector h; and based on h * =h+ε k The current vector h is perturbed by noise in a manner to obtain the corresponding current perturbation vector h * ; and based on the current disturbance vector h * For the current feature tensor H now The corresponding sub-feature vector h now,I Replace it to get a new feature tensor H * ; And the current feature tensor H * The MLP model of the molecular affinity prediction model is input for processing and the predicted affinity output by this processing is used as a corresponding perturbation affinity z i,k ; and at the end of this round of traversal, add 1 to the first counter; Step 46, identifying whether the first counter is greater than the first counting threshold; if so, go to step 47; if not, return to step 42; Step 47, based on all the obtained affinity true values ​​z true,k and all the perturbation affinities z i,k estimating the contribution fraction of each of the second atoms in the first molecular structure to obtain the corresponding contribution fraction of the first atom; The first atomic contribution fraction is estimated as follows: Among them, score i is the first atom contribution score corresponding to the i-th second atom in the first molecular structure.

5. A device for executing the processing method for molecular affinity analysis according to any one of claims 1 to 4, characterized in that: The device comprises: a model building module, a data acquisition module, a model training module and a model application module; The model building module is used to build a molecular affinity prediction model by using the Uni-Mol model as an encoder and the MLP model as a downstream regression calculation task model of the encoder; the molecular affinity prediction model is used to perform molecular affinity prediction processing according to the molecular structure X input by the model and output the corresponding predicted affinity Y; the molecular structure X includes multiple atoms x i , 1≤atom index i≤M, M is the total number of atoms in the molecular structure X; each atom x i The atomic parameters include the atom type s i and the atomic three-dimensional coordinates c i ; The predicted affinity Y is a real number; The data acquisition module is used to collect big data on the molecular structures of organic molecules used for organic photovoltaic materials and the corresponding molecular affinity values ​​through a plurality of preset public data acquisition channels; and each collected molecular structure and the corresponding molecular affinity value are used as a corresponding first training molecular structure and a first label affinity to form a corresponding first data record; and all the first data records obtained form a corresponding first data set; the plurality of public data acquisition channels at least include a plurality of public molecular information libraries and a plurality of types of public technical literature in the field of organic photovoltaic materials; the first data set includes a plurality of the first data records; the first data record includes the first training molecular structure and the first label affinity; the first training molecular structure includes a plurality of first atoms; the atomic parameters of each of the first atoms include a first atom type and a first atom three-dimensional coordinate; The model training module is used to perform model training on the molecular affinity prediction model based on the first data set; The model application module is used to receive a first molecular structure input by a user after model training is completed; and input the first molecular structure as the corresponding molecular structure X into the molecular affinity prediction model for prediction processing and use the predicted affinity Y output by the model this time as the corresponding first molecular affinity; and analyze the contribution score of each atom in the first molecular structure to the molecular affinity to obtain the corresponding first atomic contribution score; and feed back the obtained first molecular affinity and all the first atomic contribution scores to the current user; the first molecular structure includes multiple second atoms; the atomic parameters of each second atom include a second atom type and a second atom three-dimensional coordinate.

6. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is used to couple with the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 4; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method and device for processing affinity prediction data of MHC molecules and polypeptides

    CN116386761A

  • Antibody-antigen affinity prediction method, device and system and storage medium

    CN118629501A

  • Training method and device for molecular generation model

    US20240330657A1

Cited By

  • Method and device for processing bimolecular transfer integral prediction model

    CN120783879A