Processing method and device for photoelectric molecule generation model
By designing the photoelectric molecular generation model and training framework, using the underdamped Langzhiwan MCMC algorithm to generate rich molecular structures, the problems of long experimental periods and limited structure types in the existing photoelectric molecular generation methods are solved, and the effect of rapid generation of multiple structures is achieved.
Patent Information
- Application Number
- CN202510283219.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing photoelectric molecule generation methods have problems such as long experimental periods and limited structural types, which are difficult to meet the needs of rapid generation of multiple structures.
A photoelectric molecular generation model is designed, and a rich molecular structure is generated by constructing voxel tensor initialization module and photoelectric molecular generation model, and training based on the underdamped Langzhiwan MCMC algorithm.
It realizes the generation of multiple molecular structures at one time without being restricted by artificial experience, shortens the generation cycle and enriches the types of structures.
Smart Images

Figure CN120183559A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a processing method and device for an optoelectronic molecule generation model. Background Art
[0002] Optoelectronic molecules have unique electronic structures and optoelectronic properties, and play important roles in aspects such as optoelectronic conversion, light detection, and optical storage. At present, most of the conventional generation methods of optoelectronic molecules are realized based on experimental methods, such as synthesis methods, gas-phase methods, solution methods, etc., and all experimental methods have problems such as long experimental periods and limited types of generated structures.
[0003] With the popularization and application of artificial intelligence (AI) technology in the field of new drug discovery, we found that a series of customized models for new drug molecule generation based on AI technology can achieve the purpose of shortening the generation period and enriching the types of structures. In view of this, if AI technology is introduced into the field of optoelectronic molecule generation and a corresponding customized model is designed, a similar improvement effect can also be achieved. Therefore, the technical problem to be solved by the present invention is: designing an optoelectronic molecule generation model capable of generating optoelectronic molecules with rich structures, and training and applying it. Summary of the Invention
[0004] The purpose of the present invention is to provide a processing method, device, electronic device, and computer-readable storage medium for an optoelectronic molecule generation model in view of the deficiencies of the prior art. The present invention pre-constructs a voxel tensor initialization module capable of creating a voxel tensor V according to the original molecular structure M0 input by the module, and an optoelectronic molecule generation model capable of performing molecular structure generation processing on the voxel tensor V input by the model and outputting a predicted molecular structure M, and a model training framework is composed of the voxel tensor initialization module and the optoelectronic molecule generation model; then, a corresponding model data set is constructed by collecting data on known optoelectronic molecule structures, and the optoelectronic molecule generation model is trained based on the model data set and the model training framework; then, after the model training is completed, the step threshold s max and the sampling interval △s input by the user are received, and an initial voxel tensor is initialized based on the random noise addition method, and based on the underdamped Langevin Markov Chain Monte Carlo (MCMC) algorithm, with the initial voxel tensor as the starting quantity for s maxSample step by step, and during the sampling process, extract the sampling voxel tensor corresponding to the current step every △s steps as a corresponding first voxel tensor. Then, use each first voxel tensor as a corresponding voxel tensor V and input it into the optoelectronic molecule generation model for processing to obtain a corresponding predicted molecular structure M. Finally, form a new structure set with all the obtained predicted molecular structures M and feedback it to the current user. Based on the optoelectronic molecule generation model of the present invention, multiple molecular structures can be generated at one time without being restricted by artificial experience. That is to say, through the present invention, both the generation cycle can be shortened and the structure types can be enriched.
[0005] To achieve the above object, a first aspect of an embodiment of the present invention provides a processing method for an optoelectronic molecule generation model, and the method includes:
[0006] Construct a voxel tensor initialization module and an optoelectronic molecule generation model; and form a model training framework with the voxel tensor initialization module and the optoelectronic molecule generation model. The voxel tensor initialization module is used to create a corresponding voxel tensor V according to the input original molecular structure M0 of the module. The optoelectronic molecule generation model is used to perform molecular structure generation processing on the input voxel tensor V of the model and output a corresponding predicted molecular structure M.
[0007] Collect data on known optoelectronic molecular structures to construct a corresponding model data set; and train the optoelectronic molecule generation model based on the model data set and the model training framework.
[0008] After the model training is completed, receive the step threshold s max and the sampling interval △s input by the user; and perform voxel tensor initialization processing to obtain a corresponding initial voxel tensor; and based on the underdamped Langevin MCMC algorithm, use the initial voxel tensor as the starting quantity to perform s max step-by-step continuous sampling to obtain corresponding s max sampling voxel tensors, and during the sampling process, extract the sampling voxel tensor corresponding to the current step every △s steps as a corresponding first voxel tensor; then, use each first voxel tensor as a corresponding voxel tensor V and input it into the optoelectronic molecule generation model for processing to obtain a corresponding predicted molecular structure M; and form a corresponding new structure set with all the obtained predicted molecular structures M in this time and feedback it to the current user.
[0009] Preferably, the voxel tensor V is a probability distribution tensor with a shape of N×L 3 and is composed of L 3 voxel vectors v with a length of N i,j,k ; the voxel vector v i,j,kComposed of N probability values ranging from 0 to 1; L is the side length of a preset three-dimensional voxel grid space and L is an even number, and the volume of the three-dimensional voxel grid space is L 3 is greater than the maximum molecular volume among all known optoelectronic molecules; the three-dimensional voxel grid space is composed of L 3 voxel grids, and the voxel grids correspond one-to-one with the voxel vector v i,j,k , where i, j, and k are the three-dimensional coordinates of the voxel grid, 0 ≤ i ≤ L, 0 ≤ j ≤ L, 0 ≤ k ≤ L; the central voxel grid coordinates of the three-dimensional voxel grid space are (i = L / 2, j = L / 2, k = L / 2); N is the total number of atomic types in a preset first atomic type set, and the first atomic type set is the total set of atomic types available for optoelectronic molecules; the N probability values in the voxel vector v i,j,k correspond one-to-one with the N types of atomic types in the first atomic type set, and each probability value in the voxel vector v i,j,k is used to represent the probability that the current voxel grid is occupied by an atom of a corresponding atomic type;
[0010] The predicted molecular structure M includes multiple atoms a p , 1 ≤ atomic index p ≤ A M , and A M is the total number of atoms in the predicted molecular structure M; the atomic parameters of each atom a p include the atomic type t p and the three-dimensional atomic coordinate c p (x, y, z); all the atomic types t p satisfy the first atomic type set;
[0011] The original molecular structure M0 includes multiple atoms a q , 1 ≤ atomic index q ≤ A0, and A0 is the total number of atoms in the original molecular structure M0; the atomic parameters of each atom a q include the atomic type t q and the three-dimensional atomic coordinate c q (x, y, z); all the atomic types t q satisfy the first atomic type set;
[0012] The model data set includes multiple collected molecular structures M S ; each collected molecular structure M S is composed of multiple atoms a w , and the atomic parameters of each atom a w include the atomic type t w and the three-dimensional atomic coordinate c w (x, y, z), 1 ≤ atomic index w ≤ AS , A S is the total number of atoms of the currently described collected molecular structure M S ; all the atoms of all the collected molecular structures M S of all the atom types t w all satisfy the first set of atom types.
[0013] Preferably, the voxel tensor initialization module is used to create a corresponding voxel tensor V according to the original molecular structure M0 input by the module, specifically including:
[0014] The voxel tensor initialization module sets a zero tensor with a shape of N×L 3 as the initial tensor V init ; the initial tensor V init consists of L 3 voxel vectors v of length N init,(i,j,k) ; the voxel vector v init,(i,j,k) consists of N preset initialization probability values, 0 < initialization probability value < 1;
[0015] and performs a round of traversal on all the atoms a q of the original molecular structure M0; and during this round of traversal, the currently traversed atom a q is used as the corresponding current atom; and the Euclidean distance between the three-dimensional atomic coordinates c q (x, y, z) of the current atom and each voxel grid coordinate (i, j, k) of the three-dimensional voxel grid space is calculated to obtain the corresponding first distance, and the voxel grid corresponding to the smallest of the first distances is used as the corresponding current matching grid; and the voxel vector v init,(i,j,k) corresponding to the current matching grid is used as the corresponding current matching vector; and the probability value corresponding to the atom type t q of the current atom among the N probability values of the current matching vector is reset to 1;
[0016] and after this round of traversal of all the atoms a q of the original molecular structure M0, a random noise tensor NS that satisfies an isotropic Gaussian distribution with a mean of 0, a standard deviation σ of a preset first standard deviation, and a shape of N×L 3 is created; and a corresponding noise-added tensor V noise is obtained in the manner of V init = V noise + NS; and it is identified whether all the probability values in the currently described noise-added tensor V noise are non-negative. If not, a new random noise tensor NS is created again for the initial tensor Vinit Perform noise addition until all probability values in the latest noise-added tensor V noise are non-negative. If so, use the current noise-added tensor V noise as the corresponding voxel tensor V.
[0017] Preferably, the optoelectronic molecule generation model includes a high-dimensional tensor mapper, a 3D U-Net model, a molecular structure conversion module, and a pre-trained Uni-Mol model;
[0018] The input end of the high-dimensional tensor mapper is connected to the input end of the optoelectronic molecule generation model, and the output end is connected to the input end of the 3D U-Net model; the output end of the 3D U-Net model is connected to the input end of the molecular structure conversion module; the output end of the molecular structure conversion module is connected to the input end of the pre-trained Uni-Mol model; the output end of the pre-trained Uni-Mol model is connected to the output end of the optoelectronic molecule generation model;
[0019] The high-dimensional tensor mapper is used to map the default input tensor shape N ’ ×L 3 in N ’ of the voxel tensor V input to the model as the target dimension, and map the N in the tensor shape N×L of the voxel tensor V input to the model as the initial dimension; and perform high-dimensional tensor mapping processing on the voxel tensor V from the initial dimension to the target dimension to obtain a high-dimensional encoded tensor E with a shape of N 3 ×L ’ and send it to the 3D U-Net model; the default input tensor shape N 3 ×L ’ in the 3D U-Net model 3 where N ’ is a power of 2 and greater than N;
[0020] The structure of the high-dimensional tensor mapper is sequentially connected by the first and second residual blocks; each of the first and second residual blocks is sequentially connected by a 1×1×1 convolutional layer, two 3×3×3 convolutional layers, and a connection unit ⊕; in the first or second residual block, the 1×1×1 convolutional layer is used to increase the dimension of the input tensor of this convolutional layer, each of the 3×3×3 convolutional layers is used to encode the input tensor of this convolutional layer to obtain a corresponding encoded tensor, and process the current encoded tensor based on a SiLU activation function, and the connection unit ⊕ is used to add the output tensor of the 1×1×1 convolutional layer of this residual module to the output tensor of the second 3×3×3 convolutional layer;
[0021] The 3D U-Net model is used to treat the high-dimensional encoded tensor E as a high-dimensional voxel tensor with noise and perform denoising voxel tensor segmentation to obtain a denoised voxel tensor V whose shape is consistent with the shape of the voxel tensor V input to the model. clean Send it to the molecular structure conversion module; the denoised voxel tensor V clean has a shape of N×L 3 , and is composed of L 3 voxel vectors v of length N clean,(i,j,k) ; the voxel vector v clean,(i,j,k) is composed of N probability values ranging from 0 to 1;
[0022] The molecular structure conversion module is used to perform a round of traversal on all the voxel vectors v clean of the denoised voxel tensor V clean,(i,j,k) ; and during this round of traversal, the currently traversed voxel vector v clean,(i,j,k) is used as the corresponding current vector; and it is identified whether the N probability values of the current vector are all lower than a preset first probability threshold; if the N probability values of the current vector are all lower than the first probability threshold, a corresponding first voxel feature is set to 0; if at least one of the N probability values of the current vector is not lower than the first probability threshold, the maximum value among the N probability values of the current vector is used as the corresponding current maximum probability value, and it is identified whether the current maximum probability value is higher than a preset second probability threshold, if so, the corresponding first voxel feature is set to 1, if not, the corresponding first voxel feature is set to 0; and when the obtained first voxel feature is 1, with the center voxel grid coordinate of the three-dimensional voxel grid space as the coordinate origin (x = 0, y = 0, z = 0) of the atomic coordinate system, the voxel grid coordinate (i, j, k) corresponding to the current vector is subjected to coordinate conversion from the voxel grid coordinate to the atomic coordinate system to obtain a corresponding first atomic coordinate, and the atomic type corresponding to the current maximum probability value corresponding to the current first voxel feature is used as the corresponding first atomic type, and then a corresponding first atom is formed by the first atomic type and the first atomic coordinate obtained this time; and at the end of this round of traversal, all the first atoms obtained are used to form a corresponding first molecular structure and sent to the pre-trained Uni-Mol model;
[0023] The pre-trained Uni-Mol model is implemented based on the pre-training framework structure of the Uni-Mol model and consists of the Uni-Mol model, an atomic type prediction head, and an atomic coordinate prediction head; the Uni-Mol model is an atomic-level encoder implemented based on the Encoder module of the Transformer model, the atomic type prediction head is a prediction network formed by connecting a linear layer and a Softmax layer, and the atomic coordinate prediction head is implemented based on an SE(3) equivariant prediction head and the position update process of the EGNN model; the input end of the Uni-Mol model is connected to the input end of the pre-trained Uni-Mol model, the first output end is connected to the input end of the atomic type prediction head, and the second output end is connected to the input end of the atomic coordinate prediction head; the pre-trained Uni-Mol model has been pre-trained according to the pre-training method of the pre-training framework of the Uni-Mol model.
[0024] The pre-trained Uni-Mol model is used to perform structural optimization on the first molecular structure and output the corresponding predicted molecular structure M. Specifically: the Uni-Mol model performs atomic-level feature encoding on the input first molecular structure to obtain the corresponding atomic type encoding vector and atomic pair feature matrix, and sends the atomic type encoding vector to the atomic type prediction head and the atomic pair feature matrix to the atomic coordinate prediction head; and the atomic type prediction head performs atomic type prediction based on the atomic type encoding vector to obtain multiple predicted atomic types; and the atomic coordinate prediction head performs atomic coordinate prediction based on the atomic pair feature matrix to obtain multiple predicted atomic coordinates, and the predicted atomic coordinates correspond one-to-one to the predicted atomic types; and each of the predicted atomic types is used as a corresponding atomic type t p and each of the predicted atomic coordinates corresponding to the predicted atomic types is used as a corresponding three-dimensional atomic coordinate c p (x, y, z), and from each atomic type t p and its corresponding three-dimensional atomic coordinate c p (x, y, z) to form a corresponding atom a p ; and all the obtained atoms a p are used to form the corresponding predicted molecular structure M.
[0025] Preferably, training the optoelectronic molecule generation model based on the model dataset and the model training framework specifically includes:
[0026] Step 501, initialize the voxel tensor initialization module of the model training framework and the optoelectronic molecule generation model as the corresponding current initialization module and current generation model.
[0027] Step 502: Split the model dataset into two sub-datasets according to a preset first splitting ratio, denoted as the corresponding first training set and first evaluation set;
[0028] Among them, both the first training set and the first evaluation set consist of multiple collected molecular structures M S ; The ratio of the total number of molecular structures in the first training set and the first evaluation set satisfies the first splitting ratio;
[0029] Step 503: Take the first collected molecular structure M in the first training set S as the corresponding current training structure;
[0030] Step 504: Input the current training structure as a corresponding original molecular structure M0 into the current initialization module for processing to obtain the corresponding voxel tensor V; and use the initial tensor V init that generates the current voxel tensor V through noise addition during the processing as the corresponding first label tensor;
[0031] Step 505: Input the current voxel tensor V into the high-dimensional tensor mapper of the current generation model for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model of the current generation model for processing to obtain the corresponding denoised voxel tensor V clean ; and use the current denoised voxel tensor V clean as the corresponding first prediction tensor;
[0032] Step 506: Input the first prediction tensor and the first label tensor corresponding to the current training structure into a preset first model loss function L M1 for calculation to obtain the corresponding first loss value;
[0033] Among them, the first model loss function L M1 is implemented based on the L1 loss function, L2 loss function or cross-entropy loss function;
[0034] Step 507: Identify whether the first loss value meets a preset first loss value range; if the first loss value meets the first loss value range, then identify whether the current training structure is the last collected molecular structure M in the first training set S , if so, go to Step 508, if not, take the next collected molecular structure M in the first training set SAs the new current training structure and return to step 504; if the first loss value does not satisfy the first loss value range, then based on a preset first model optimizer, towards making the first model loss function L M1 reach the minimum value direction to modulate the model parameters of the high-dimensional tensor mapper and the 3D U-Net model for one round, and return to step 505 at the end of this round of modulation;
[0035] Among them, the first model optimizer at least includes Adam optimizer and SGD optimizer;
[0036] Step 508, perform a round of traversal on all the collected molecular structures M in the first evaluation set S ; and during this round of traversal, use the currently traversed collected molecular structure M S as the corresponding current evaluation structure; and use the current evaluation structure as a corresponding original molecular structure M0 to input into the current initialization module for processing to obtain the corresponding voxel tensor V, and use the initial tensor V init that generates the current voxel tensor V by adding noise during the processing as the corresponding second label tensor; and input the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor V clean ; and use the current denoised voxel tensor V clean as the corresponding second prediction tensor; and form a corresponding first prediction-label pair from the second prediction tensor and the second label tensor corresponding to the current evaluation structure; and at the end of this round of traversal, input all the obtained first prediction-label pairs into a preset first model evaluation function for calculation to obtain the corresponding first evaluation value;
[0037] Among them, the first model evaluation function is implemented based on the MAE function, MSE function or RMSE function;
[0038] Step 509, identify whether the first evaluation value satisfies a preset first evaluation value range; if it satisfies, go to step 510; if it does not satisfy, return to step 503 to continue training;
[0039] Step 510, based on a preset second segmentation ratio, split the model data set into two sub-data sets denoted as the corresponding second training set and second evaluation set;
[0040] Among them, both the second training set and the second evaluation set consist of multiple collected molecular structures M SComposition; the ratio of the total number of molecular structures of the second training set to that of the second evaluation set satisfies the first segmentation ratio;
[0041] Step 511, take the first collected molecular structure M of the second training set S as the corresponding current training structure;
[0042] Step 512, take the current training structure as a corresponding first label structure; and take the current training structure as a corresponding original molecular structure M0 and input it into the current initialization module for processing to obtain the corresponding voxel tensor V; and input the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor V clean ; and input the current denoised voxel tensor V clean into the molecular structure conversion module of the current generation model for processing to obtain the corresponding first molecular structure;
[0043] Step 513, input the current first molecular structure into the pre-trained Uni-Mol model of the current generation model for processing to obtain the corresponding predicted molecular structure M; and take the current predicted molecular structure M as a corresponding first prediction structure;
[0044] Step 514, bring the first prediction structure corresponding to the current training structure and the first label structure into the preset second model loss function L M2 for calculation to obtain the corresponding second loss value;
[0045] wherein, the second model loss function L M2 is implemented based on the L1 loss function, L2 loss function or cross-entropy loss function;
[0046] Step 515, identify whether the second loss value satisfies the preset second loss value range; if the second loss value satisfies the second loss value range, then identify whether the current training structure is the last collected molecular structure M of the second training set S for identification, if so, go to Step 516, if not, take the next collected molecular structure M of the second training set S as the new current training structure and return to Step 512; if the second loss value does not satisfy the second loss value range, then based on the preset second model optimizer, fine-tune the model parameters of the pre-trained Uni-Mol model in the direction of minimizing the second model loss function L M2 for one round and return to Step 513 at the end of this round of fine-tuning;
[0047] Among them, the second model optimizer includes at least an Adam optimizer and an SGD optimizer;
[0048] Step 516, perform a round of traversal on all the collected molecular structures M in the second evaluation set; and during this round of traversal, use the currently traversed collected molecular structure M S as the corresponding current evaluation structure; and use the current evaluation structure as a corresponding second label structure; and use the current evaluation structure as a corresponding original molecular structure M0 and input it into the current initialization module for processing to obtain the corresponding voxel tensor V; and input the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor V S ; and input the current denoised voxel tensor V clean into the molecular structure conversion module for processing to obtain the corresponding first molecular structure; and input the current first molecular structure into the pre-trained Uni-Mol model for processing to obtain the corresponding predicted molecular structure M; and use the current predicted molecular structure M as a corresponding second predicted structure; and form a corresponding second prediction-label pair from the second predicted structure and the second label structure corresponding to the current evaluation structure; and at the end of this round of traversal, input all the obtained second prediction-label pairs into a preset second model evaluation function for calculation to obtain the corresponding second evaluation value; clean
[0049] Among them, the second model evaluation function is implemented based on the MAE function, the MSE function, or the RMSE function;
[0050] Step 517, identify whether the second evaluation value meets a preset second evaluation value range; if not, return to step 511 to continue training; if so, stop training and confirm that the model training is completed.
[0051] Preferably, the voxel tensor initialization process to obtain a corresponding initial voxel tensor specifically includes:
[0052] Set a zero tensor with a shape of N×L 3 as the corresponding first tensor V1; the first tensor V1 is composed of L 3 voxel vectors v with a length of N 1,(i,j,k) ; the voxel vector v 1,(i,j,k) is composed of N preset initialization probability values, 0 < initialization probability value < 1;
[0053] and create a random noise tensor NA that satisfies an isotropic Gaussian distribution with a mean of 0, a standard deviation σ of a preset first standard deviation, and a shape of N×L 3 ; and obtain a corresponding second tensor V2 in the manner of V2 = V1 + NA; and identify whether all probability values in the current second tensor V2 are non-negative. If not, create a new random noise tensor NA again to add noise to the first tensor V1 until all probability values in the latest second tensor V2 are non-negative. If so, use the current second tensor V2 as the corresponding initial voxel tensor.
[0054] Preferably, the underdamped Langevin MCMC algorithm starts with the initial voxel tensor and performs s max steps of continuous sampling to obtain corresponding s max sampled voxel tensors, and during the sampling process, extract the sampled voxel tensor corresponding to the current step number every △s steps as a corresponding first voxel tensor, specifically including:
[0055] Set the first variable y s=0 to the corresponding initial voxel tensor, and set the second variable v s=0 to 0; and substitute y s=0 , v s=0 into the underdamped Langevin MCMC algorithm formula for s max steps of continuous derivation to obtain s max first variables y 1≤s and second variables v 1≤s ; and use the obtained s max first variables y 1≤s as the corresponding s max sampled voxel tensors; and during the sampling process, use the sampled voxel tensor corresponding to the first variable y 1≤s where the sampling step number s is a non-zero integer multiple of △s as a corresponding first voxel tensor;
[0056] where 0 ≤ sampling step number s ≤ s max ;
[0057] The underdamped Langevin MCMC algorithm formula is:
[0058]
[0059] δ is a preset step size parameter, u is a preset inverse mass parameter, γ is a preset friction parameter; ε is a random noise tensor that satisfies an isotropic Gaussian distribution with a mean of 0, a standard deviation σ of a preset first standard deviation, and a shape of N×L 3 ; y s , y s+1They are the first variables in the s-th and (s + 1)-th steps respectively; v s and v s+1 They are the second variables in the s-th and (s + 1)-th steps respectively; g θ () is a scoring function, and θ is a built-in parameter of this scoring function.
[0060] In a second aspect of the embodiments of the present invention, there is provided an apparatus for implementing the processing method of the optoelectronic molecule generation model described in the first aspect above. The apparatus includes: a model construction module, a model training module, and a model application module;
[0061] The model construction module is used to construct a voxel tensor initialization module and an optoelectronic molecule generation model; and a model training framework is composed of the voxel tensor initialization module and the optoelectronic molecule generation model. The voxel tensor initialization module is used to create a corresponding voxel tensor V according to the input original molecular structure M0 of the module. The optoelectronic molecule generation model is used to perform molecular structure generation processing on the input voxel tensor V of the model and output a corresponding predicted molecular structure M;
[0062] The model training module is used to construct a corresponding model data set by collecting data on known optoelectronic molecular structures; and train the optoelectronic molecule generation model based on the model data set and the model training framework;
[0063] The model application module is used to, after the model training is completed, receive the step number threshold s max input by the user and the sampling interval △s; and perform voxel tensor initialization processing to obtain a corresponding initial voxel tensor; and based on the underdamped Langevin MCMC algorithm, use the initial voxel tensor as the starting quantity to perform s max steps of continuous sampling to obtain corresponding s max sampled voxel tensors, and extract the sampled voxel tensor corresponding to the current step number as a corresponding first voxel tensor every △s steps during the sampling process; and use each of the first voxel tensors as a corresponding voxel tensor V to input into the optoelectronic molecule generation model for processing to obtain a corresponding predicted molecular structure M; and form a corresponding new structure set from all the predicted molecular structures M obtained this time and feedback it to the current user.
[0064] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: a memory, a processor, and a transceiver;
[0065] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;
[0066] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0067] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.
[0068] The embodiments of the present invention provide a processing method, apparatus, electronic device, and computer-readable storage medium for an optoelectronic molecule generation model. As can be seen from the above, the embodiments of the present invention pre-construct a voxel tensor initialization module capable of creating a voxel tensor V according to the original molecular structure M0 input by the module, and an optoelectronic molecule generation model capable of performing molecular structure generation processing on the voxel tensor V input by the model and outputting a predicted molecular structure M, and a model training framework is composed of the voxel tensor initialization module and the optoelectronic molecule generation model; then, data collection is performed on known optoelectronic molecular structures to construct a corresponding model data set, and the optoelectronic molecule generation model is trained based on the model data set and the model training framework; then, after the model training is completed, the step threshold s max and the sampling interval △s are received, and an initial voxel tensor is initialized based on the random noise addition method, and s max continuous samplings are performed with the initial voxel tensor as the starting quantity based on the underdamped Langevin MCMC algorithm, and the sampling voxel tensor corresponding to the current step is extracted every △s steps during the sampling process as a corresponding first voxel tensor, and each first voxel tensor is used as a corresponding voxel tensor V to be input into the optoelectronic molecule generation model for processing to obtain a corresponding predicted molecular structure M, and all the obtained predicted molecular structures M are used to form a new structure set and fed back to the current user. The optoelectronic molecule generation model based on the embodiments of the present invention can generate multiple molecular structures at one time without being restricted by manual experience. Through the embodiments of the present invention, both the generation cycle is shortened and the types of molecular structures are enriched. Description of the Drawings
[0069] Figure 1 It is a schematic diagram of a processing method for an optoelectronic molecule generation model provided in Embodiment 1 of the present invention;
[0070] Figure 2 It is a module structure diagram of the optoelectronic molecule generation model and the model training framework provided in Embodiment 1 of the present invention;
[0071] Figure 3 It is a module structure diagram of a processing apparatus for an optoelectronic molecule generation model provided in Embodiment 2 of the present invention;
[0072] Figure 4A schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. Detailed implementation manners
[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0074] Embodiment 1 of the present invention provides a processing method for an optoelectronic molecule generation model, as Figure 1 shown in the schematic diagram of the processing method for an optoelectronic molecule generation model provided in Embodiment 1 of the present invention, the method mainly includes the following steps:
[0075] Step 1, construct a voxel tensor initialization module and an optoelectronic molecule generation model; and form a model training framework with the voxel tensor initialization module and the optoelectronic molecule generation model.
[0076] Here, the voxel tensor initialization module of the embodiment of the present invention is used to create a voxel tensor V according to the input original molecular structure M0 of the module; the optoelectronic molecule generation model of the embodiment of the present invention is used to perform molecular structure generation processing on the input voxel tensor V of the model and output the corresponding predicted molecular structure M.
[0077] The voxel tensor V of the embodiment of the present invention is a probability distribution tensor with a shape of N×L 3 and is composed of L 3 voxel vectors v with a length of N i,j,k ; each voxel vector v i,j,k is composed of N probability values ranging from 0 to 1. Among them:
[0078] 1) L is the side length of a preset three-dimensional voxel grid space and L is an even number, and the volume L of this three-dimensional voxel grid space 3 is greater than the maximum molecular volume among all known optoelectronic molecules; this three-dimensional voxel grid space is composed of L 3 voxel grids, and the voxel grids correspond one-to-one with the voxel vectors v i,j,k ; the subscripts i, j, k are the three-dimensional coordinates of the voxel grid, 0≤i≤L, 0≤j≤L, 0≤k≤L; the central voxel grid coordinates of this three-dimensional voxel grid space are (i = L / 2, j = L / 2, k = L / 2);
[0079] 2) N is the total number of atomic types in a preset first atomic type set, and this first atomic type set is the total set of atomic types available for optoelectronic molecules;
[0080] 3) The voxel vector v i,j,k The N probability values in correspond one-to-one with N types of atomic types in the first atomic type set. Each probability value in the voxel vector v i,j,k is used to represent the probability that the current voxel grid is occupied by an atom of a corresponding atomic type.
[0081] The predicted molecular structure M of the embodiment of the present invention includes multiple atoms a p ; where 1 ≤ atomic index p ≤ A M , and A M is the total number of atoms in the predicted molecular structure M; the atomic parameters of each atom a p include the atomic type t p and the three-dimensional atomic coordinate c p (x, y, z); all atomic types t p satisfy the first atomic type set.
[0082] The original molecular structure M0 of the embodiment of the present invention includes multiple atoms a q ; where 1 ≤ atomic index q ≤ A0, and A0 is the total number of atoms in the original molecular structure M0; the atomic parameters of each atom a q include the atomic type t q and the three-dimensional atomic coordinate c q (x, y, z); all atomic types t q satisfy the first atomic type set.
[0083] It should be noted that the voxel tensor initialization module is used to create a corresponding voxel tensor V according to the input original molecular structure M0 of the module, specifically including:
[0084] Step A1, the voxel tensor initialization module sets a zero tensor with a shape of N × L 3 as the initial tensor V init ;
[0085] Here, the initial tensor V init is composed of L 3 voxel vectors v init,(i,j,k) with a length of N; the voxel vector v init,(i,j,k) is composed of N preset initialization probability values, where 0 < initialization probability value < 1;
[0086] Step A2, perform a round of traversal on all atoms a q of the original molecular structure M0; and during this round of traversal, regard the currently traversed atom a q as the corresponding current atom; and for the three-dimensional atomic coordinate c q(x, y, z) calculates the Euclidean distance from each voxel grid coordinate (i, j, k) in the three-dimensional voxel grid space to obtain the corresponding first distance, and takes the voxel grid corresponding to the smallest first distance as the corresponding current matching grid; and takes the voxel vector v corresponding to the current matching grid init,(i,j,k) as the corresponding current matching vector; and resets the probability value corresponding to the atomic type t of the current atom among the N probability values of the current matching vector q to 1;
[0087] Step A3, and after the current round of traversal of all atoms a of the original molecular structure M0 q ends, create a random noise tensor NS that satisfies an isotropic Gaussian distribution, has a mean of 0, a standard deviation σ of a preset first standard deviation, and a shape of N×L 3 ; and obtain a corresponding noisy tensor V by the method of V noise = V init + NS; and identify whether all probability values in the current noisy tensor V noise are non-negative. If not, create a new random noise tensor NS again to add noise to the initial tensor V noise until all probability values in the latest noisy tensor V init are non-negative. If so, take the current noisy tensor V noise as the corresponding voxel tensor V. noise Here, the first standard deviation is a preset standard deviation value.
[0088] It should also be noted that, as
[0089] shown in the module structure diagram of the optoelectronic molecule generation model and the model training framework provided in Embodiment 1 of the present invention, the optoelectronic molecule generation model of the present invention includes a high-dimensional tensor mapper, a 3D U-Net model, a molecular structure conversion module, and a pre-trained Uni-Mol model. Figure 2 The connection relationships of the components of the optoelectronic molecule generation model are as follows: the input end of the high-dimensional tensor mapper is connected to the input end of the optoelectronic molecule generation model, and the output end is connected to the input end of the 3D U-Net model; the output end of the 3D U-Net model is connected to the input end of the molecular structure conversion module; the output end of the molecular structure conversion module is connected to the input end of the pre-trained Uni-Mol model; the output end of the pre-trained Uni-Mol model is connected to the output end of the optoelectronic molecule generation model.
[0090] The functions of the components of the optoelectronic molecule generation model are as follows.
[0091] 1) High-dimensional tensor mapper:
[0092] 1) High-dimensional tensor mapper:
[0093] The high-dimensional tensor mapper according to an embodiment of the present invention is used to map the N in the default input tensor shape N ’ ×L 3 in ’ to the target dimension, and the N in the tensor shape N×L of the voxel tensor V input to the model is denoted as the initial dimension; and the voxel tensor V is subjected to high-dimensional tensor mapping processing from the initial dimension to the target dimension to obtain a high-dimensional encoded tensor E with a shape of N 3 ×L ’ and sent to the 3D U-Net model. Here, the N in the default input tensor shape N 3 ×L ’ in 3 is a power of 2 and greater than N. ’
[0094] As Figure 2 shown, the structure of the high-dimensional tensor mapper is sequentially connected by first and second residual blocks; wherein, the first and second residual blocks are each sequentially connected by a 1×1×1 convolutional layer, two 3×3×3 convolutional layers, and a connection unit ⊕; and in the first or second residual block: a) the 1×1×1 convolutional layer is used to increase the dimension of the input tensor of this convolutional layer, b) each 3×3×3 convolutional layer is used to encode the input tensor of this convolutional layer to obtain a corresponding encoded tensor, and process the current encoded tensor based on a SiLU activation function, c) the connection unit ⊕ is used to add the output tensor of the 1×1×1 convolutional layer of this residual module to the output tensor of the second 3×3×3 convolutional layer.
[0095] 2) 3D U-Net model:
[0096] The 3D U-Net model according to an embodiment of the present invention is used to regard the high-dimensional encoded tensor E as a high-dimensional voxel tensor with noise and perform denoising voxel tensor segmentation on it to obtain a denoised voxel tensor V whose shape is the same as the shape of the voxel tensor V input to the model clean and send it to the molecular structure conversion module. Here, the shape of the denoised voxel tensor V clean is N×L 3 , and it consists of L 3 voxel vectors v clean,(i,j,k) with a length of N; the voxel vector v clean,(i,j,k) consists of N probability values ranging from 0 to 1.
[0097] It should be noted that the detailed model structure of the 3D U-Net model used in the embodiments of the present invention can be understood through the public technical literature A, "3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation", and will not be further elaborated here.
[0098] 3) Molecular structure conversion module:
[0099] The molecular structure conversion module of the embodiments of the present invention is used to perform a round of traversal on all voxel vectors v clean of the denoised voxel tensor V clean,(i,j,k) ; and during this round of traversal, the currently traversed voxel vector v clean,(i,j,k) is used as the corresponding current vector; and it is identified whether the N probability values of the current vector are all lower than a preset first probability threshold; if the N probability values of the current vector are all lower than the first probability threshold, a corresponding first voxel feature is set to 0; if at least one of the N probability values of the current vector is not lower than the first probability threshold, the maximum value among the N probability values of the current vector is used as the corresponding current maximum probability value, and it is identified whether the current maximum probability value is higher than a preset second probability threshold, if so, the corresponding first voxel feature is set to 1, if not, the corresponding first voxel feature is set to 0; and when the obtained first voxel feature is 1, the coordinate transformation from the voxel grid coordinate (i, j, k) corresponding to the current vector to the atomic coordinate system is performed with the central voxel grid coordinate of the three-dimensional voxel grid space as the coordinate origin (x = 0, y = 0, z = 0) of the atomic coordinate system to obtain a corresponding first atomic coordinate, and the atomic type corresponding to the current maximum probability value corresponding to the current first voxel feature is used as the corresponding first atomic type, and then a corresponding first atom is formed by the first atomic type and the first atomic coordinate obtained this time; and at the end of this round of traversal, a corresponding first molecular structure composed of all the first atoms obtained is sent to the pre-trained Uni-Mol model. Here, the first and second probability thresholds are two pre-set probability parameters.
[0100] 4) Pre-trained Uni-Mol model:
[0101] The pre-trained Uni-Mol model of the embodiment of the present invention is implemented based on the pre-training framework structure of the Uni-Mol model, and is composed of a Uni-Mol model, an atomic type prediction head, and an atomic coordinate prediction head; wherein, the Uni-Mol model is an atomic-level encoder implemented based on the Encoder module of the Transformer model, the atomic type prediction head is a prediction network connected by a linear layer and a Softmax layer, and the atomic coordinate prediction head is implemented by an SE(3) equivariant prediction head and a position update process based on the EGNN model. The input end of the Uni-Mol model is connected to the input end of the pre-trained Uni-Mol model, the first output end is connected to the input end of the atomic type prediction head, and the second output end is connected to the input end of the atomic coordinate prediction head. It should be noted that the pre-training framework of the Uni-Mol model mentioned above, as well as the specific model structure implementation methods of the Uni-Mol model, the atomic type prediction head, and the atomic coordinate prediction head can be further understood through the public technical literature B 《Uni-Mol: A Universal 3D Molecular Representation Learning Framework》, and will not be elaborated further here. It should also be noted that the pre-trained Uni-Mol model of the embodiment of the present invention has been pre-trained according to the pre-training method of the pre-training framework of the Uni-Mol model, specifically, it has been pre-trained according to the various data sets and training methods mentioned in the public technical literature B.
[0102] The pre-trained Uni-Mol model of the embodiment of the present invention is used to perform structural optimization processing on the first molecular structure and output the corresponding predicted molecular structure M. Specifically: the Uni-Mol model performs atomic-level feature encoding on the input first molecular structure to obtain the corresponding atomic type encoding vector and atomic pair feature matrix, and sends the atomic type encoding vector to the atomic type prediction head and the atomic pair feature matrix to the atomic coordinate prediction head; and the atomic type prediction head performs atomic type prediction based on the atomic type encoding vector to obtain multiple predicted atomic types; and the atomic coordinate prediction head performs atomic coordinate prediction based on the atomic pair feature matrix to obtain multiple predicted atomic coordinates, and the predicted atomic coordinates correspond one-to-one to the predicted atomic types; and each predicted atomic type is used as a corresponding atomic type t p and each predicted atomic coordinate corresponding to the predicted atomic type is used as a corresponding three-dimensional atomic coordinate c p (x, y, z), and each atomic type t p and its corresponding three-dimensional atomic coordinate c p (x, y, z) form a corresponding atom a p ; and all the obtained atoms a p form the corresponding predicted molecular structure M.
[0103] It should also be noted that, as Figure 2 shown, the model training framework of the embodiment of the present invention is composed of a voxel tensor initialization module and an optoelectronic molecule generation model connected in sequence.
[0104] Step 2: Construct a corresponding model data set by collecting data on known optoelectronic molecule structures; and train the optoelectronic molecule generation model based on the model data set and the model training framework;
[0105] Specifically, it includes: Step 21: Construct a corresponding model data set by collecting data on known optoelectronic molecule structures;
[0106] Here, the model data set of the embodiment of the present invention includes multiple collected molecular structures M S ; each collected molecular structure M S is composed of multiple atoms a w ; each atom a w has atomic parameters including atomic type t w and three-dimensional atomic coordinates c w (x, y, z), 1 ≤ atomic index w ≤ A S , A S is the total number of atoms of the current collected molecular structure M S ; all atomic types t S of all collected molecular structures M w satisfy the first atomic type set;
[0107] Step 22: Train the optoelectronic molecule generation model based on the model data set and the model training framework;
[0108] Specifically, it includes: Step 22-1: Use the voxel tensor initialization module and the optoelectronic molecule generation model of the model training framework as the corresponding current initialization module and current generation model;
[0109] Step 22-2: Divide the model data set into two sub-data sets according to a preset first splitting ratio, denoted as the corresponding first training set and first evaluation set;
[0110] Here, the first splitting ratio is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set are composed of multiple collected molecular structures M S ; the ratio of the total number of molecular structures of the first training set and the first evaluation set satisfies the first splitting ratio;
[0111] Step 22-3: Use the first collected molecular structure M S of the first training set as the corresponding current training structure;
[0112] Step 22-4: Input the current training structure as a corresponding original molecular structure M0 into the current initialization module for processing to obtain the corresponding voxel tensor V; and use the initial tensor V generated by adding noise during the processing as the corresponding first label tensor. init as the corresponding first label tensor;
[0113] Step 22-5: Input the current voxel tensor V into the high-dimensional tensor mapper of the current generation model for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model of the current generation model for processing to obtain the corresponding denoised voxel tensor V; and use the current denoised voxel tensor V clean as the corresponding first prediction tensor; clean as the corresponding first prediction tensor;
[0114] Step 22-6: Input the first prediction tensor and the first label tensor corresponding to the current training structure into the preset first model loss function L M1 for calculation to obtain the corresponding first loss value;
[0115] Here, the first model loss function L of the embodiment of the present invention M1 is implemented based on the L1 loss function, the L2 loss function or the cross-entropy loss function;
[0116] Step 22-7: Identify whether the first loss value satisfies the preset first loss value range; if the first loss value satisfies the first loss value range, then identify whether the current training structure is the last collected molecular structure M of the first training set S for identification. If so, go to Step 22-8. If not, use the next collected molecular structure M of the first training set S as the new current training structure and return to Step 22-4; if the first loss value does not satisfy the first loss value range, then based on the preset first model optimizer, modulate the model parameters of the high-dimensional tensor mapper and the 3D U-Net model in the direction of minimizing the first model loss function L M1 for one round, and return to Step 22-5 at the end of this round of modulation;
[0117] Here, the first loss value range of the embodiment of the present invention is a preset numerical range; the first model optimizer includes at least the Adam optimizer and the SGD optimizer;
[0118] Step 22-8: Perform one round of traversal on all the collected molecular structures M of the first evaluation set; and during this round of traversal, use the currently traversed collected molecular structure M S for one round of traversal; and during this round of traversal, use the currently traversed collected molecular structure M Sas the corresponding current evaluation structure; and use the current evaluation structure as a corresponding original molecular structure M0 to input it into the current initialization module for processing to obtain the corresponding voxel tensor V, and use the initial tensor V generated by adding noise during the processing as the current voxel tensor V init as the corresponding second label tensor; and input the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor V clean ; and use the current denoised voxel tensor V clean as the corresponding second prediction tensor; and form a corresponding first prediction-label pair from the second prediction tensor and the second label tensor corresponding to the current evaluation structure; and at the end of this round of traversal, input all the obtained first prediction-label pairs into the preset first model evaluation function for calculation to obtain the corresponding first evaluation value;
[0119] Here, the first model evaluation function of the embodiments of the present invention is implemented based on the MAE function, the MSE function or the RMSE function;
[0120] Step 22-9, identify whether the first evaluation value meets the preset first evaluation value range; if it meets, go to Step 22-10; if it does not meet, return to Step 22-3 to continue training;
[0121] Here, the first evaluation value range of the embodiments of the present invention is a preset numerical range;
[0122] Step 22-10, divide the model data set into two sub-data sets according to the preset second segmentation ratio, denoted as the corresponding second training set and second evaluation set;
[0123] Here, the second segmentation ratio is a preset ratio parameter, such as 7:3; both the second training set and the second evaluation set are composed of multiple collected molecular structures M S ; the ratio of the total number of molecular structures in the second training set and the second evaluation set meets the first segmentation ratio;
[0124] Step 22-11, use the first collected molecular structure M in the second training set S as the corresponding current training structure;
[0125] Step 22-12, use the current training structure as a corresponding first label structure; and use the current training structure as a corresponding original molecular structure M0 to input it into the current initialization module for processing to obtain the corresponding voxel tensor V; and input the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor Vclean ; and input the current voxel tensor V after denoising clean into the molecular structure conversion module of the current generation model for processing to obtain the corresponding first molecular structure;
[0126] Step 22-13, and input the current first molecular structure into the pre-trained Uni-Mol model of the current generation model for processing to obtain the corresponding predicted molecular structure M; and use the current predicted molecular structure M as a corresponding first predicted structure;
[0127] Step 22-14, bring the first predicted structure corresponding to the current training structure and the first label structure into the preset second model loss function L M2 for calculation to obtain the corresponding second loss value;
[0128] Here, the second model loss function L of the embodiment of the present invention M2 is implemented based on the L1 loss function, the L2 loss function or the cross-entropy loss function;
[0129] Step 22-15, identify whether the second loss value satisfies the preset second loss value range; if the second loss value satisfies the second loss value range, then identify whether the current training structure is the last collected molecular structure M of the second training set S for identification, if so, go to Step 22-16, if not, use the next collected molecular structure M of the second training set S as the new current training structure and return to Step 22-12; if the second loss value does not satisfy the second loss value range, then based on the preset second model optimizer, adjust the model parameters of the pre-trained Uni-Mol model in the direction of making the second model loss function L M2 reach the minimum value for a round of fine-tuning, and return to Step 22-13 at the end of this round of fine-tuning;
[0130] Here, the second loss value range of the embodiment of the present invention is a preset numerical range; the second model optimizer includes at least the Adam optimizer and the SGD optimizer;
[0131] Step 22-16, perform a round of traversal on all the collected molecular structures M of the second evaluation set S ; and during this round of traversal, use the currently traversed collected molecular structure M S as the corresponding current evaluation structure; and use the current evaluation structure as a corresponding second label structure; and use the current evaluation structure as a corresponding original molecular structure M0 to input into the current initialization module for processing to obtain the corresponding voxel tensor V; and input the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoded tensor E; and input the current high-dimensional encoded tensor E into the 3D U-Net model for processing to obtain the corresponding voxel tensor V after denoisingclean ; and input the current denoised voxel tensor V clean into the molecular structure conversion module for processing to obtain the corresponding first molecular structure; and input the current first molecular structure into the pre-trained Uni-Mol model for processing to obtain the corresponding predicted molecular structure M; and use the current predicted molecular structure M as a corresponding second predicted structure; and form a corresponding second prediction-label pair from the second predicted structure and the second label structure corresponding to the current evaluation structure; and at the end of this round of traversal, input all the obtained second prediction-label pairs into the preset second model evaluation function for calculation to obtain the corresponding second evaluation value;
[0132] Here, the second model evaluation function of the embodiments of the present invention is implemented based on the MAE function, the MSE function or the RMSE function;
[0133] Step 22-17, identify whether the second evaluation value meets the preset second evaluation value range; if not, return to step 22-11 to continue training; if so, stop training and confirm that the model training is completed.
[0134] Here, the second evaluation value range of the embodiments of the present invention is a preset numerical range.
[0135] Step 3, after the model training is completed, receive the step threshold s max input by the user and the sampling interval △s; and perform voxel tensor initialization processing to obtain a corresponding initial voxel tensor; and based on the underdamped Langevin MCMC algorithm, use the initial voxel tensor as the starting quantity to perform max s steps of continuous sampling to obtain the corresponding s max sampled voxel tensors, and extract the sampled voxel tensor corresponding to the current step as a corresponding first voxel tensor every △s steps during the sampling process; and input each first voxel tensor as a corresponding voxel tensor V into the optoelectronic molecule generation model for processing to obtain a corresponding predicted molecular structure M; and form a corresponding new structure set from all the obtained predicted molecular structures M and feedback it to the current user;
[0136] Specifically, it includes: Step 31, after the model training is completed, receive the step threshold s max input by the user and the sampling interval △s;
[0137] Step 32, and perform voxel tensor initialization processing to obtain a corresponding initial voxel tensor;
[0138] Specifically, it includes: Step 321, set a all-zero tensor with a shape of N×L 3 as the corresponding first tensor V1;
[0139] Here, the first tensor V1 is composed of L3 a voxel vector v of length N 1,(i,j,k) consisting of; the voxel vector v 1,(i,j,k) is composed of N preset initial probability values, where 0 < initial probability value < 1;
[0140] Step 322, and create a random noise tensor NA that satisfies an isotropic Gaussian distribution with a mean of 0, a standard deviation σ of a preset first standard deviation, and a shape of N×L 3 ; and obtain a corresponding second tensor V2 in the way of V2 = V1 + NA; and identify whether all the probability values in the current second tensor V2 are non-negative. If not, create a new random noise tensor NA again to add noise to the first tensor V1 until all the probability values in the latest second tensor V2 are non-negative. If so, use the current second tensor V2 as the corresponding initial voxel tensor;
[0141] Step 33, and perform s max steps of continuous sampling based on the underdamped Langevin MCMC algorithm with the initial voxel tensor as the starting quantity to obtain corresponding s max sampled voxel tensors, and extract the sampled voxel tensor corresponding to the current step number as a corresponding first voxel tensor every △s steps during the sampling process;
[0142] Specifically include: set the first variable y s=0 as the corresponding initial voxel tensor, and set the second variable v s=0 to 0; and substitute y s=0 , v s=0 into the underdamped Langevin MCMC algorithm formula for s max steps of continuous derivation to obtain s max first variables y 1≤s and second variables v 1≤s ; and use the obtained s max first variables y 1≤s as the corresponding s max sampled voxel tensors; and during the sampling process, use the sampled voxel tensor corresponding to the first variable y 1≤s where the sampling step number s is a non-zero integer multiple of △s as a corresponding first voxel tensor;
[0143] where 0 ≤ sampling step number s ≤ s max ;
[0144] The underdamped Langevin MCMC algorithm formula is:
[0145]
[0146] In the above formula: δ is a preset step parameter, u is a preset anti-mass parameter, γ is a preset friction parameter; ε is a random noise tensor that satisfies an isotropic Gaussian distribution with a mean of 0, a standard deviation σ of a preset first standard deviation, and a shape of N×L 3 ; y s and y s+1 are the first variables at the s-th and (s + 1)-th steps respectively; v s and v s+1 are the second variables at the s-th and (s + 1)-th steps respectively; g θ () is a scoring function, and θ is a built-in parameter of this scoring function;
[0147] Step 34, and take each first voxel tensor as a corresponding voxel tensor V and input it into the optoelectronic molecule generation model for processing to obtain a corresponding predicted molecular structure M; and form a corresponding new structure set from all the predicted molecular structures M obtained this time and feedback it to the current user.
[0148] It should be noted that the predicted molecular structure M obtained by the optoelectronic molecule generation model of the embodiment of the present invention does not contain the chemical bond information between atom pairs. The embodiment of the present invention can also call some conventional chemical information tools, such as OpenBabel, etc., when knowing all the atoms a p of each predicted molecular structure M p atomic type t p and three-dimensional atomic coordinates c p (x, y, z), set corresponding chemical bonds for the atom pairs that can have connection relationships in the current predicted molecular structure M, and thus obtain a corresponding chemical bond set, and form a corresponding atom set from all the atoms a
[0149] Figure 3 is a module structure diagram of a processing device of an optoelectronic molecule generation model provided in the second embodiment of the present invention. This device is a terminal device or a server that implements the foregoing method embodiment, or can also be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, this device can be a device or a chip system of the foregoing terminal device or server. As Figure 3 shown, this device includes: a model construction module 201, a model training module 202, and a model application module 203.
[0150] The model construction module 201 is used to construct a voxel tensor initialization module and a photoelectric molecule generation model; and a model training framework is composed of the voxel tensor initialization module and the photoelectric molecule generation model; the voxel tensor initialization module is used to create a voxel tensor V corresponding to the original molecular structure M0 input by the module; the photoelectric molecule generation model is used to perform molecular structure generation processing on the voxel tensor V input by the model and output the corresponding predicted molecular structure M.
[0151] The model training module 202 is used to construct a corresponding model data set by collecting data on known photoelectric molecular structures; and train the photoelectric molecule generation model based on the model data set and the model training framework.
[0152] The model application module 203 is used to receive the step threshold s max and the sampling interval △s input by the user after the model training is completed; and perform voxel tensor initialization processing to obtain a corresponding initial voxel tensor; and based on the underdamped Langevin MCMC algorithm, use the initial voxel tensor as the starting quantity to perform s max steps of continuous sampling to obtain the corresponding s max sampled voxel tensors, and extract the sampled voxel tensor corresponding to the current step as a corresponding first voxel tensor every △s steps during the sampling process; and use each first voxel tensor as a corresponding voxel tensor V to input into the photoelectric molecule generation model for processing to obtain a corresponding predicted molecular structure M; and form a corresponding new structure set from all the predicted molecular structures M obtained this time and feedback it to the current user.
[0153] The processing device for a photoelectric molecule generation model provided by the embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, and will not be described in detail here.
[0154] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the model construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determined module can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0155] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0156] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0157] Figure 4 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device can be a terminal device or a server for implementing the method of the foregoing embodiments, or a terminal device or a server for implementing the method of the foregoing embodiments connected to the foregoing terminal device or server. As Figure 4 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions can be stored in the memory 302 to complete various processing functions and implement the processing steps described in the foregoing method embodiments. Preferably, the electronic device related to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.
[0158] In Figure 4The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.
[0159] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0160] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored. When it runs on a computer, the computer is made to execute the methods and processing procedures provided in the above embodiments.
[0161] The embodiments of the present invention provide a processing method, device, electronic device, and computer-readable storage medium for an optoelectronic molecule generation model. As can be seen from the above, the embodiments of the present invention pre-construct a voxel tensor initialization module that can create a voxel tensor V according to the original molecular structure M0 input by the module, and an optoelectronic molecule generation model that can perform molecular structure generation processing on the voxel tensor V input by the model and output a predicted molecular structure M, and form a model training framework from the voxel tensor initialization module and the optoelectronic molecule generation model; then collect data on known optoelectronic molecule structures to construct a corresponding model data set, and train the optoelectronic molecule generation model based on the model data set and the model training framework; then after the model training is completed, receive the step threshold s max and the sampling interval △s input by the user, initialize an initial voxel tensor based on the random noise addition method, and perform s starting from the initial voxel tensor based on the underdamped Langevin MCMC algorithm maxStep - by - step continuous sampling is performed. During the sampling process, every △s steps, the sampling voxel tensor corresponding to the current step number is extracted as a corresponding first voxel tensor. Each first voxel tensor is used as a corresponding voxel tensor V and input into the optoelectronic molecule generation model for processing to obtain a corresponding predicted molecular structure M. All the obtained predicted molecular structures M form a new structure set and are fed back to the current user. The optoelectronic molecule generation model based on the embodiments of the present invention can generate multiple molecular structures at one time without being restricted by artificial experience. Through the embodiments of the present invention, both the generation cycle is shortened and the types of molecular structures are enriched.
[0162] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination thereof. The software modules may be placed in a random access memory (RAM), memory, read - only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD - ROM, or any other form of storage medium known in the art.
[0163] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above - mentioned are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for processing a photoelectric molecular generation model, characterized in that: The method comprises: Construct a voxel tensor initialization module and a photoelectric molecule generation model; and the voxel tensor initialization module and the photoelectric molecule generation model form a model training framework; the voxel tensor initialization module is used to create a voxel tensor according to the original molecular structure M0 input by the module to obtain a corresponding voxel tensor V; the photoelectric molecule generation model is used to perform molecular structure generation processing according to the voxel tensor V input by the model and output the corresponding predicted molecular structure M; Constructing a corresponding model data set by collecting data on known optoelectronic molecular structures; and training the optoelectronic molecule generation model based on the model data set and the model training framework; After the model training is completed, the step threshold s input by the user is received max and sampling interval △s; and perform voxel tensor initialization processing to obtain a corresponding initial voxel tensor; and perform s based on the underdamped Langevin MCMC algorithm with the initial voxel tensor as the starting quantity max Step continuous sampling to get the corresponding s max A sampling voxel tensor is obtained, and during the sampling process, the sampling voxel tensor corresponding to the current step is extracted every △s steps as a corresponding first voxel tensor; each of the first voxel tensors is input into the photoelectric molecular generation model as a corresponding voxel tensor V for processing to obtain a corresponding predicted molecular structure M; and all the predicted molecular structures M obtained this time form a corresponding new structure set to feedback to the current user.
2. The method for processing a photoelectric molecule generation model according to claim 1, characterized in that: The voxel tensor V is a shape of N×L 3 The probability distribution tensor of L 3 A voxel vector v of length N i,j,k Composition; the voxel vector v i,j,k It is composed of N probability values between 0 and 1; L is the side length of a preset three-dimensional voxel grid space and L is an even number, and the volume of the three-dimensional voxel grid space is L 3 is larger than the maximum molecular volume of all known photoelectric molecules; the three-dimensional voxel grid space is composed of L 3 The voxel grid is composed of a voxel vector v i,j,k One-to-one correspondence, i, j, k are the three-dimensional coordinates of the voxel grid, 0≤i≤L, 0≤j≤L, 0≤k≤L; the central voxel grid coordinates of the three-dimensional voxel grid space are (i=L / 2, j=L / 2, k=L / 2); N is the total number of atom types of a preset first atom type set, and the first atom type set is a total set of atom types that can be used for photoelectric molecules; the voxel vector v i,j,k The N probability values in the first atom type set correspond one-to-one to the N types of atoms in the first atom type set, and the voxel vector v i,j,k Each probability value in is used to represent the probability that the current voxel grid is occupied by an atom of the corresponding atomic type; The predicted molecular structure M includes multiple atoms a p , 1≤atomic index p≤A M , A M is the total number of atoms in the predicted molecular structure M; each of the atoms a p The atomic parameters include the atomic type t p and the three-dimensional atomic coordinates c p (x,y,z); all the atoms of type t p All satisfy the first set of atom types; The original molecular structure M0 includes a plurality of atoms a q , 1≤atom index q≤A0, A0 is the total number of atoms in the original molecular structure M0; each a q The atomic parameters include the atomic type t q and the three-dimensional atomic coordinates c q (x,y,z); all the atoms of type t q All satisfy the first set of atom types; The model data set includes a plurality of collected molecular structures M S Each of the collected molecular structures M S are composed of multiple atoms a w Composition, each of the atoms a w The atomic parameters include the atomic type t w and the three-dimensional atomic coordinates c w (x,y,z), 1≤atomic index w≤A S , A S The molecular structure M is collected for the current description S The total number of atoms; all the collected molecular structures M S All the atomic types t w All satisfy the first set of atom types.
3. The method for processing a photoelectric molecule generation model according to claim 2, characterized in that: The voxel tensor initialization module is used to create a voxel tensor according to the original molecular structure M0 input by the module to obtain a corresponding voxel tensor V, specifically including: The voxel tensor initialization module sets a shape of N×L 3 The all-zero tensor V is used as the initial tensor init ; The initial tensor V init By L 3 A voxel vector v of length N init,(i,j,k) Composition; the voxel vector v init,(i,j,k) It consists of N preset initialization probability values, 0<initialization probability value<1; And for all the atoms a of the original molecular structure M0 q Perform a round of traversal; and in this round of traversal, the atom a currently traversed q as the corresponding current atom; and the three-dimensional atomic coordinate c of the current atom q The Euclidean distance between (x, y, z) and each voxel grid coordinate (i, j, k) in the three-dimensional voxel grid space is calculated to obtain a corresponding first distance, and the voxel grid corresponding to the smallest first distance is used as the corresponding current matching grid; and the voxel vector v corresponding to the current matching grid is calculated. init,(i,j,k) As the corresponding current matching vector; and the N probability values of the current matching vector and the atom type t of the current atom q The corresponding probability value is reset to 1; And in all the atoms a of the original molecular structure M0 q After this round of traversal, create an isotropic Gaussian distribution with a mean of 0, a standard deviation σ of the preset first standard deviation, and a shape of N×L 3 The random noise tensor NS is noise =V init +NS to get a corresponding noise tensor V noise ; and add noise to the current tensor V noise Whether all probability values in are not negative is identified, if not, a new random noise tensor NS is created again for the initial tensor V init Noise is added until the latest noisy tensor V noise Until all probability values in are not negative, if so, the current noise tensor V noise As the corresponding voxel tensor V.
4. The method for processing a photoelectric molecule generation model according to claim 2, characterized in that: The optoelectronic molecule generation model includes a high-dimensional tensor mapper, a 3D U-Net model, a molecular structure conversion module and a pre-trained Uni-Mol model; The input end of the high-dimensional tensor mapper is connected to the input end of the photoelectric molecule generation model, and the output end is connected to the input end of the 3D U-Net model; the output end of the 3D U-Net model is connected to the input end of the molecular structure conversion module; the output end of the molecular structure conversion module is connected to the input end of the pre-trained Uni-Mol model; the output end of the pre-trained Uni-Mol model is connected to the output end of the photoelectric molecule generation model; The high-dimensional tensor mapper is used to convert the default input tensor shape N of the 3D U-Net model ’ ×L 3 N ’ Denoted as the target dimension, the tensor shape of the voxel tensor V input to the model is N×L 3 The N in the figure is recorded as the initial dimension; and the voxel tensor V is mapped from the initial dimension to the target dimension to obtain a voxel tensor with a shape of N ’ ×L 3 The high-dimensional encoded tensor E is sent to the 3D U-Net model; the default input tensor shape N of the 3D U-Net model ’ ×L 3 N ’ It is 2 to the power of n and is greater than N; The structure of the high-dimensional tensor mapper is formed by sequentially connecting the first and second residual blocks; the first and second residual blocks are each formed by sequentially connecting a 1×1×1 convolutional layer, two 3×3×3 convolutional layers and a connection unit ⊕; in the first or second residual block, the 1×1×1 convolutional layer is used to increase the dimension of the input tensor of the convolutional layer, each of the 3×3×3 convolutional layers is used to encode the input tensor of the convolutional layer to obtain a corresponding encoding tensor, and process the current encoding tensor based on a Si LU activation function, and the connection unit ⊕ is used to add the output tensor of the 1×1×1 convolutional layer of the residual module to the output tensor of the second 3×3×3 convolutional layer; The 3D U-Net model is used to treat the high-dimensional encoding tensor E as a high-dimensional voxel tensor with noise and perform denoising voxel tensor segmentation on it to obtain a denoised voxel tensor V with a shape consistent with the shape of the voxel tensor V input by the model. clean Send to the molecular structure conversion module; the denoised voxel tensor V clean The shape is N×L 3 , by L 3 A voxel vector v of length N clean,(i,j,k) The voxel vector v clean,(i,j,k) It consists of N probability values between 0 and 1; The molecular structure conversion module is used to convert the denoised voxel tensor V clean All the voxel vectors v clean,(i,j,k) Perform a round of traversal; and in this round of traversal, the voxel vector v currently traversed is clean,(i,j,k) as the corresponding current vector; and identifying whether the N probability values of the current vector are all lower than a preset first probability threshold; If all N probability values of the current vector are lower than the first probability threshold, setting a corresponding first voxel feature to 0; If at least one of the N probability values of the current vector is not lower than the first probability threshold, the maximum value of the N probability values of the current vector is used as the corresponding current maximum probability value, and whether the current maximum probability value is higher than the preset second probability threshold is identified. If so, the corresponding first voxel feature is set to 1, otherwise the corresponding first voxel feature is set to 0; and when the first voxel feature obtained is 1, the voxel grid coordinates (i, j, k) corresponding to the current vector are transformed from voxel grid coordinates to atomic coordinate system with the central voxel grid coordinates of the three-dimensional voxel grid space as the coordinate origin of the atomic coordinate system (x=0, y=0, z=0) to obtain a corresponding first atomic coordinate, and the atomic type corresponding to the current maximum probability value corresponding to the current first voxel feature is used as the corresponding first atomic type, and then the first atomic type obtained this time and the first atomic coordinates form a corresponding first atom; and at the end of this round of traversal, all the first atoms obtained form a corresponding first molecular structure and send it to the pre-trained Uni-Mol model; The pre-trained Uni-Mol model is implemented based on the pre-training framework structure of the Uni-Mol model, and is composed of the Uni-Mol model, an atom type prediction head and an atom coordinate prediction head; the Uni-Mol model is an atomic-level encoder implemented by an Encoder module based on a Transformer model, the atom type prediction head is a prediction network formed by connecting a linear layer and a Softmax layer, and the atom coordinate prediction head is an SE (3) equivariant prediction head and implemented based on a position update process of an EGNN model; the input end of the Uni-Mol model is connected to the input end of the pre-trained Uni-Mol model, the first output end is connected to the input end of the atom type prediction head, and the second output end is connected to the input end of the atom coordinate prediction head; the pre-trained Uni-Mol model has been pre-trained according to the pre-training method of the pre-training framework of the Uni-Mol model; The pre-trained Uni-Mol model is used to perform structural optimization processing on the first molecular structure and output the corresponding predicted molecular structure M, specifically: the Uni-Mol model performs atomic-level feature encoding according to the input first molecular structure to obtain the corresponding atomic type encoding vector and atom pair feature matrix, and sends the atom type encoding vector to the atom type prediction head, and sends the atom pair feature matrix to the atomic coordinate prediction head; the atom type prediction head performs atomic type prediction according to the atom type encoding vector to obtain multiple predicted atomic types; the atomic coordinate prediction head performs atomic coordinate prediction according to the atom pair feature matrix to obtain multiple predicted atomic coordinates, and the predicted atomic coordinates correspond to the predicted atomic type one by one; and each of the predicted atomic types is used as a corresponding atomic type t p , taking the predicted atomic coordinates corresponding to each predicted atomic type as a corresponding three-dimensional atomic coordinate c p (x, y, z), and each of the atom types t p and its corresponding three-dimensional atomic coordinates c p (x, y, z) forms a corresponding atom a p ; and by all the atoms a obtained p The predicted molecular structure M corresponds to the composition.
5. The method for processing a photoelectric molecular generation model according to claim 4, characterized in that: The training of the photoelectric molecule generation model based on the model data set and the model training framework specifically includes: Step 501, using the voxel tensor initialization module and the photoelectric molecule generation model of the model training framework as the corresponding current initialization module and current generation model; Step 502, based on a preset first segmentation ratio, the model data set is divided into two sub-data sets recorded as a corresponding first training set and a first evaluation set; Wherein, the first training set and the first evaluation set are both composed of a plurality of the collected molecular structures M S composition; the ratio of the total number of molecular structures in the first training set to the first evaluation set satisfies the first segmentation ratio; Step 503: The first collected molecular structure M of the first training set S As the corresponding current training structure; Step 504: input the current training structure as a corresponding original molecular structure M0 into the current initialization module for processing to obtain the corresponding voxel tensor V; and generate the initial tensor V of the current voxel tensor V by adding noise during the processing. init as the corresponding first label tensor; Step 505: input the current voxel tensor V into the high-dimensional tensor mapper of the current generation model for processing to obtain the corresponding high-dimensional encoding tensor E; and input the current high-dimensional encoding tensor E into the 3D U-Net model of the current generation model for processing to obtain the corresponding denoised voxel tensor V. clean ; and the current denoised voxel tensor V clean as the corresponding first prediction tensor; Step 506: Substitute the first prediction tensor and the first label tensor corresponding to the current training structure into the preset first model loss function L M1 Calculate and obtain the corresponding first loss value; Among them, the first model loss function L M1 Implemented based on L1 loss function, L2 loss function or cross entropy loss function; Step 507, identifying whether the first loss value meets the preset first loss value range; if the first loss value meets the first loss value range, determining whether the current training structure is the last collected molecular structure M of the first training set. S If yes, go to step 508, if no, the next collected molecular structure M of the first training set is S as the new current training structure and return to step 504; if the first loss value does not satisfy the first loss value range, based on the preset first model optimizer, move toward making the first model loss function L M1 Perform a round of modulation on the model parameters of the high-dimensional tensor mapper and the 3D U-Net model in the direction in which the minimum value is reached, and return to step 505 when this round of modulation ends; Wherein, the first model optimizer includes at least an Adam optimizer and an SGD optimizer; Step 508: all the collected molecular structures M in the first evaluation set S Perform a round of traversal; and in this round of traversal, the collected molecular structure M currently traversed S as the corresponding current evaluation structure; and input the current evaluation structure as a corresponding original molecular structure M0 into the current initialization module for processing to obtain the corresponding voxel tensor V, and generate the initial tensor V of the current voxel tensor V by adding noise during the processing init As the corresponding second label tensor; and input the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoding tensor E; and input the current high-dimensional encoding tensor E into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor V clean ; and the current denoised voxel tensor V clean as the corresponding second prediction tensor; and the second prediction tensor corresponding to the current evaluation structure and the second label tensor form a corresponding first prediction-label pair; and at the end of this round of traversal, all the obtained first prediction-label pairs are brought into the preset first model evaluation function to calculate and obtain the corresponding first evaluation value; Wherein, the first model evaluation function is implemented based on the MAE function, the MSE function or the RMSE function; Step 509, identifying whether the first evaluation value meets the preset first evaluation value range; if so, go to step 510; if not, return to step 503 to continue training; Step 510, dividing the model data set into two sub-data sets based on a preset second division ratio and recording them as a corresponding second training set and a second evaluation set; Wherein, the second training set and the second evaluation set are both composed of a plurality of the collected molecular structures M S composition; the ratio of the total number of molecular structures of the second training set and the second evaluation set satisfies the first segmentation ratio; Step 511: The first collected molecular structure M of the second training set S As the corresponding current training structure; Step 512, taking the current training structure as a corresponding first label structure; and inputting the current training structure as a corresponding original molecular structure M0 into the current initialization module for processing to obtain the corresponding voxel tensor V; and inputting the current voxel tensor V into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoding tensor E; and inputting the current high-dimensional encoding tensor E into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor V clean ; and the current denoised voxel tensor V clean The molecular structure conversion module of the current generation model is input for processing to obtain the corresponding first molecular structure; Step 513, input the current first molecular structure into the pre-trained Uni-Mo l model of the current generation model for processing to obtain the corresponding predicted molecular structure M; and use the current predicted molecular structure M as a corresponding first predicted structure; Step 514: Substitute the first prediction structure and the first label structure corresponding to the current training structure into the preset second model loss function L M2 Calculate and obtain a corresponding second loss value; Among them, the second model loss function L M2 Implemented based on L1 loss function, L2 loss function or cross entropy loss function; Step 515, identifying whether the second loss value meets the preset second loss value range; if the second loss value meets the second loss value range, determining whether the current training structure is the last collected molecular structure M of the second training set. S If yes, go to step 516, if no, the next collected molecular structure M of the second training set is S as the new current training structure and return to step 512; if the second loss value does not satisfy the second loss value range, then based on the preset second model optimizer, move toward making the second model loss function L M2 Perform a round of fine-tuning on the model parameters of the pre-trained Uni-Mo l model in the direction of reaching the minimum value, and return to step 513 when this round of fine-tuning is completed; Wherein, the second model optimizer includes at least an Adam optimizer and an SGD optimizer; Step 516: all the collected molecular structures M in the second evaluation set S Perform a round of traversal; and in this round of traversal, the collected molecular structure M currently traversed S as the corresponding current evaluation structure; and the current evaluation structure is used as a corresponding second label structure; and the current evaluation structure is used as a corresponding original molecular structure M0 and input into the current initialization module for processing to obtain the corresponding voxel tensor V; and the current voxel tensor V is input into the high-dimensional tensor mapper for processing to obtain the corresponding high-dimensional encoding tensor E; and the current high-dimensional encoding tensor E is input into the 3D U-Net model for processing to obtain the corresponding denoised voxel tensor V clean ; and the current denoised voxel tensor V clean Input the molecular structure conversion module for processing to obtain the corresponding first molecular structure; input the current first molecular structure into the pre-trained Uni-Mo l model for processing to obtain the corresponding predicted molecular structure M; and use the current predicted molecular structure M as a corresponding second predicted structure; and the second predicted structure corresponding to the current evaluation structure and the second label structure form a corresponding second prediction-label pair; and at the end of this round of traversal, all the obtained second prediction-label pairs are brought into the preset second model evaluation function for calculation to obtain the corresponding second evaluation value; Wherein, the second model evaluation function is implemented based on the MAE function, the MSE function or the RMSE function; Step 517, identifying whether the second evaluation value meets the preset second evaluation value range; if not, returning to step 511 to continue training; if satisfied, stopping training and confirming the end of model training.
6. The method for processing a photoelectric molecular generation model according to claim 2, characterized in that: The voxel tensor initialization process is performed to obtain a corresponding initial voxel tensor, specifically including: Set a shape of N×L 3 The all-zero tensor of L is used as the corresponding first tensor V1; the first tensor V1 is composed of L 3 A voxel vector v of length N 1,(i,j,k) The voxel vector v 1,(i,j,k) It consists of N preset initialization probability values, 0<initialization probability value<1; And create an isotropic Gaussian distribution with a mean of 0, a standard deviation σ of the preset first standard deviation, and a shape of N×L 3 A random noise tensor NA is obtained; and a corresponding second tensor V2 is obtained by V2=V1+NA; and it is identified whether all probability values in the current second tensor V2 are not negative. If not, a new random noise tensor NA is created again to add noise to the first tensor V1 until all probability values in the latest second tensor V2 are not negative. If so, the current second tensor V2 is used as the corresponding initial voxel tensor.
7. The method for processing a photoelectric molecule generation model according to claim 2, characterized in that: The underdamped Langevin MCMC algorithm is performed using the initial voxel tensor as a starting quantity. max Step continuous sampling to get the corresponding s max A sampling voxel tensor is obtained, and the sampling voxel tensor corresponding to the current step number is extracted as a corresponding first voxel tensor every △s steps during the sampling process, specifically including: The first variable y s=0 Set to the corresponding initial voxel tensor, and set the second variable v s=0 Set to 0; and set y s=0 、v s=0 Substitute the underdamped Langevin MCMC algorithm formula for s max Step by step, we can get s max The first variable y 1≤s and the second variable v 1≤s ; and the obtained s max The first variable y 1≤s As the corresponding s max The sampled voxel tensor; and in the sampling process, each sampling step s is a first variable y that is a non-zero integer multiple of △s 1≤s The corresponding sampled voxel tensor is used as a corresponding first voxel tensor; Among them, 0≤sampling steps s≤s max ; The underdamped Langevin MCMC algorithm formula is: δ is the preset step size parameter, u is the preset inverse mass parameter, γ is the preset friction parameter; ε is an isotropic Gaussian distribution with a mean of 0, a standard deviation σ is the preset first standard deviation, and a shape of N×L 3 A random noise tensor of y s ,y s+1 are the first variables of the sth and s+1th steps respectively; v s 、v s+1 are the second variables of the sth and s+1th steps respectively; g θ () is a scoring function and θ is the built-in parameter of the scoring function.
8. A device for executing the method for processing a photoelectric molecular generation model according to any one of claims 1 to 7, characterized in that: The device comprises: a model building module, a model training module and a model application module; The model building module is used to construct a voxel tensor initialization module and an optoelectronic molecule generation model; and the voxel tensor initialization module and the optoelectronic molecule generation model form a model training framework; the voxel tensor initialization module is used to create a voxel tensor according to the original molecular structure M0 input by the module to obtain a corresponding voxel tensor V; the optoelectronic molecule generation model is used to perform molecular structure generation processing according to the voxel tensor V input by the model and output the corresponding predicted molecular structure M; The model training module is used to construct a corresponding model data set by collecting data on known optoelectronic molecular structures; and to train the optoelectronic molecule generation model based on the model data set and the model training framework; The model application module is used to receive the step count threshold s input by the user after the model training is completed. max and sampling interval △s; and perform voxel tensor initialization processing to obtain a corresponding initial voxel tensor; and perform s based on the underdamped Langevin MCMC algorithm with the initial voxel tensor as the starting quantity max Step continuous sampling to get the corresponding s max A sampling voxel tensor is obtained, and during the sampling process, the sampling voxel tensor corresponding to the current step is extracted every △s steps as a corresponding first voxel tensor; each of the first voxel tensors is input into the photoelectric molecular generation model as a corresponding voxel tensor V for processing to obtain a corresponding predicted molecular structure M; and all the predicted molecular structures M obtained this time form a corresponding new structure set to feedback to the current user.
9. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is used to couple with the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 7; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Molecular design using local exploration
CN117529782A
System and method for generating a novel molecular structure using a protein structure
US20220406403A1