A processing method and device for a molecular conformation generator with conditional constraints

By constructing an end-to-end molecular conformation generator based on the Uni-Mol model, the problems of inefficient design efficiency and insufficient exploration capabilities in the existing molecular design process are solved, and a new molecular structure that meets the target properties are efficiently generated.

CN119181441BActive Publication Date: 2025-06-17BEIJING DP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411310784.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-06-17
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

The existing molecular design process has problems such as inefficient design efficiency and insufficient ability to explore unknown chemical spaces.

Method used

An end-to-end molecular conformation generator is constructed based on multiple Uni-Mol models and a property prediction model that has been trained in the model. The molecular sequence is transformed by chemical informatics tools, and batch molecules are generated based on the molecular property vectors and the number of generated molecules entered by the user.

Benefits of technology

Without experiments and molecular motion simulations, a large number of new molecular structures are generated and directional conditions can be constrained according to the set target molecular properties, improving design efficiency and exploring ability of unknown chemical spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181441B_ABST
    Figure CN119181441B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention relates to a processing method and device for a molecular conformation generator with conditional constraints. The method includes: constructing a molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training; initializing the model parameters of all Uni-Mol models in the molecular conformation generator based on preset base model parameters; receiving a first molecular sequence, a first molecular property vector, and a first number of generated molecules input by a user; performing three-dimensional conformation conversion processing on the first molecular sequence based on a preset chemoinformatics tool to obtain a corresponding first converted conformation; performing batch molecular generation task processing according to the first converted conformation, the first molecular property vector, the first number of generated molecules, and the molecular conformation generator to obtain a corresponding first batch molecular task report and feedback it to the user. Through the present invention, the molecular design efficiency can be improved, and the exploration ability of unknown molecular structures can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a processing method and device for a molecular conformation generator with conditional constraints. Background Art

[0002] In the fields of molecular design and drug research and development, generating novel molecules with specific properties is an important research direction. Currently, the conventional molecular design process generally includes the following steps: a) First, design multiple new molecular structures based on expert knowledge and manual structure adjustment methods, and pre-give a set of molecular property indicators; b) Then, measure or calculate various properties of each new molecular structure (such as melting point, solubility, density, acid / base type, stability score, reactivity score, etc.) based on experimental means or molecular motion simulation means; c) Then, compare the obtained various molecular property data with the preset various molecular property indicators, and screen the new molecular structures based on the comparison results. If the new molecular structures do not meet the standards, the above steps a-c need to be repeated until the properties of the obtained new molecular structures meet the standards. It is not difficult to see that this conventional molecular design method has defects such as low design efficiency and insufficient exploration ability of unknown chemical space. Summary of the Invention

[0003] The purpose of the present invention is to provide a processing method, device, electronic device and computer-readable storage medium for a molecular conformation generator with conditional constraints in view of the defects of the prior art. The present invention first constructs an end-to-end molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training, and initializes the model parameters of all Uni-Mol models in the generator based on the model parameters of a base Uni-Mol model that has completed model pre-training; after receiving the molecular sequence, molecular property vector and the number of generated molecules input by the user, first use chemical informatics tools (such as OpenBabel, RDKit, etc.) to perform three-dimensional conformation conversion on the molecular sequence, and then perform batch molecular generation task processing according to the converted conformation, molecular property vector, the number of generated molecules and the molecular conformation generator. Through the molecular conformation generator provided by the present invention, a large number of molecular structures can be generated without experiments, without molecular motion simulation, and without being restricted by expert knowledge, and the generated molecular structures can be directionally conditionally constrained according to the set target molecular properties. Through the present invention, not only the design efficiency can be improved, but also the exploration ability of unknown chemical space can be improved.

[0004] To achieve the above object, in the first aspect of the embodiments of the present invention, a processing method for a molecular conformation generator with conditional constraints is provided, and the method includes:

[0005] Construct a molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training; wherein, the molecular conformation generator is used to generate a new molecular conformation based on the input molecular conformation X ini and the conditional constraint vector C T to perform new molecular conformation generation processing and output the corresponding generated molecular conformation X D ;

[0006] Initialize the model parameters of all the Uni-Mol models in the molecular conformation generator based on preset base model parameters; wherein, the base model parameters are the model parameters of a base Uni-Mol model that has completed model pre-training;

[0007] Receive the first molecular sequence, the first molecular property vector, and the first number of generated molecules input by the user; wherein, the first molecular sequence is a one-dimensional SMIELS sequence; the first molecular property vector includes multiple first property data; each of the first property data corresponds to a specified molecular property; the specified molecular properties at least include melting point, solubility, density, acid / alkaline type, stability score, and reactivity score; when each of the first property data is 0, it means that the specified molecular property corresponding to the current property data is not constrained when generating a new molecular conformation, and when the first property data is not 0, it means that the specified molecular property corresponding to the current property data is constrained by the corresponding conditional property data when generating a new molecular conformation;

[0008] Perform three-dimensional conformation conversion processing on the first molecular sequence based on a preset chemoinformatics tool to obtain the corresponding first converted conformation; wherein, the chemoinformatics tool at least includes OpenBabel and RDKit;

[0009] Perform batch molecular generation task processing according to the first converted conformation, the first molecular property vector, the first number of generated molecules, and the molecular conformation generator, and feedback the corresponding first batch molecular task report to the user.

[0010] Preferably, the Uni-Mol model is used to perform molecular conformation optimization processing on the input three-dimensional molecular conformation and output the corresponding optimized molecular conformation;

[0011] The property prediction model is composed of multiple parallel first prediction models; each of the first prediction models is implemented based on a type of non-linear regression model, and the non-linear regression model at least includes the XGBoost model, the GBDT model, the Random Forest model, and the MLP model; each of the first prediction models corresponds to one of the specified molecular properties;

[0012] The property prediction model is used to send the currently input three-dimensional molecular conformation to each of the first prediction models to obtain corresponding property prediction data; and all the obtained property prediction data are combined to form a corresponding property prediction vector and output; the property prediction vector is composed of multiple pieces of the property prediction data.

[0013] Among the first prediction models corresponding to melting point, solubility, density, acid / base type, stability score or reactivity score in the property prediction model, each is used to perform corresponding prediction processing on melting point, solubility, density, acid / base type, stability score or reactivity score according to the currently input three-dimensional molecular conformation to obtain corresponding predicted values of melting point, solubility, density, acid / base type, stability score or reactivity score as a corresponding piece of the property prediction data and output.

[0014] Preferably, the first model input end of the molecular conformation generator is used to receive the molecular conformation X input by the model ini and the second model input end is used to receive the conditional constraint vector C input by the model T , and the model output end is used to output the corresponding generated molecular conformation X D ;

[0015] The molecular conformation X ini and the generated molecular conformation X D are each a three-dimensional molecular conformation; the conditional constraint vector C T is composed of multiple pieces of the conditional property data, and each piece of the conditional property data corresponds to a specified molecular property.

[0016] The molecular conformation generator includes a conformation optimization model, an encoder, the property prediction model, a constraint condition comparison module, an encoding modulation module and a decoder.

[0017] The input end of the conformation optimization model is connected to the first model input end, and the output end is connected to the first input end of the encoder; the second input end of the encoder is connected to the output end of the encoding modulation module, the first output end is connected to the input end of the property prediction model, and the second output end is connected to the first input end of the constraint condition comparison module; the output end of the property prediction model is connected to the second input end of the constraint condition comparison module; the third input end of the constraint condition comparison module is connected to the second model input end, the first output end is connected to the input end of the encoding modulation module, and the second output end is connected to the input end of the decoder; the output end of the decoder is connected to the model output end.

[0018] The conformation optimization model is implemented based on a Uni-Mol model.

[0019] The conformational optimization model is used to optimize the input molecular conformation X ini and output the corresponding optimized molecular conformation X O ; and send the optimized molecular conformation X O to the encoder;

[0020] The encoder is composed of a preset first number N E of the Uni-Mol models connected in series;

[0021] The encoder is used to save the received optimized molecular conformation X O when receiving the optimized molecular conformation X sent by the conformational optimization model; and use N O Uni-Mol models to gradually optimize the received optimized molecular conformation X E and take the optimized molecular conformation output by the last Uni-Mol model as the corresponding encoded output conformation X O ; and send the obtained encoded output conformation X E in this time to the property prediction model and the constraint condition comparison module; E

[0022] The encoder is also used to reset its own encoder parameters with the received modulation model parameters when receiving the modulation model parameters sent by the encoding modulation module; and at the end of this parameter reset, use N E Uni-Mol models to gradually optimize the saved optimized molecular conformation X O and take the optimized molecular conformation output by the N E th Uni-Mol model as the latest encoded output conformation X E ; and send the obtained encoded output conformation X E in this time to the property prediction model and the constraint condition comparison module; E

[0023] The property prediction model is used to send the currently input encoded output conformation X E to each of the first prediction models to obtain the corresponding property prediction data; and form the corresponding property prediction vector C P from all the obtained property prediction data in this time; and send the obtained property prediction vector C P in this time to the constraint condition comparison module;

[0024] The constraint condition comparison module is used to compare the input property prediction vector C P with the conditional constraint vector C TPerform constraint condition comparison processing to obtain the corresponding condition satisfaction status; and identify the condition satisfaction status; if the condition satisfaction status is not meeting the constraint conditions, then use the latest property prediction vector C P and the condition constraint vector C T to send to the encoding modulation module; if the condition satisfaction status is meeting the constraint conditions, then use the latest encoded output conformation X E to send to the decoder; where the condition satisfaction status includes not meeting the constraint conditions and meeting the constraint conditions;

[0025] The encoding modulation module is used to perform a round of parameter modulation processing on the encoder according to the property prediction vector C P and the condition constraint vector C T received this time to obtain the corresponding modulation model parameters; and send the modulation model parameters obtained this time to the encoder;

[0026] The decoder is composed of a preset second number N D of the Uni-Mol models connected in series;

[0027] The decoder is used to use N D Uni-Mol models to perform step-by-step optimization on the input encoded output conformation X E and use the optimized molecular conformation output by the N D th Uni-Mol model as the corresponding generated molecular conformation X D and output it.

[0028] Furthermore, the performing constraint condition comparison processing according to the input property prediction vector C P and the condition constraint vector C T to obtain the corresponding condition satisfaction status specifically includes:

[0029] The constraint condition comparison module records each non-zero condition property data in the condition constraint vector C T as the corresponding first comparison property data;

[0030] And perform a round of traversal on all the first comparison property data; and during this round of traversal, use the currently traversed first comparison property data as the corresponding current threshold data; and use the specified molecular property corresponding to the current threshold data as the corresponding current property; and use the property prediction vector C PThe property prediction data corresponding to the current property is used as the corresponding current prediction data; and based on the current property, a preset first correspondence table is queried, and the first relative error range field of the first correspondence record in the first correspondence table whose first property field matches the current property is extracted as the corresponding current relative error range; and a corresponding prediction-threshold data pair is formed by the current prediction data and the current threshold data; and a relative error is calculated for the current prediction data and the current threshold data, and the corresponding current relative error = |current prediction data - current threshold data| / current threshold data; and it is identified whether the current relative error meets the corresponding current relative error range. If it meets, the corresponding single-item error comparison result is set to meet, and if it does not meet, the corresponding single-item error comparison result is set to not meet; and at the end of this round of traversal, all the obtained prediction-threshold data pairs are brought into a preset error evaluation function for calculation to obtain the corresponding first evaluation value; wherein, the first correspondence table is a correspondence table reflecting the correspondence between properties and relative error ranges, and is composed of multiple first correspondence records; each first correspondence record corresponds to a specified molecular property; the first correspondence record includes the first property field and the first relative error range field; the first property field is used to store a corresponding specified molecular property, and the first relative error range field is used to store a corresponding relative error range; the single-item error comparison result includes meet and not meet; the error evaluation function at least includes the RMSE function;

[0031] And it is identified whether the first evaluation value meets a preset first evaluation value range. If it meets, the corresponding full-item comparison result is set to meet, and if it does not meet, the corresponding full-item comparison result is set to not meet; wherein, the full-item comparison result includes meet and not meet;

[0032] And all the obtained single-item error comparison results and the full-item comparison result are identified; if any one of the single-item error comparison results is not meet or the full-item comparison result is not meet, the corresponding condition satisfaction status is set to not meet the constraint condition; if all the single-item error comparison results and the full-item comparison result are meet, the corresponding condition satisfaction status is set to meet the constraint condition.

[0033] Further, the encoder is subjected to a round of parameter modulation processing according to the property prediction vector C received this time P and the condition constraint vector C T to obtain the corresponding modulation model parameters, specifically including:

[0034] Step 51, the encoding modulation module uses the condition constraint vector CT The condition property data that are all 0 in it are recorded as the corresponding invalid property data;

[0035] Step 52, and use the property prediction vector C P Reset the property prediction data corresponding to each of the invalid property data in it to 0;

[0036] Step 53, and use the property prediction vector C after reset P and the condition constraint vector C T Input them into a preset model optimization objective function;

[0037] Among them, the model optimization objective function is:

[0038] θ is the model parameter of the encoder;

[0039] L M is a preset model loss function, and the model loss function L M at least includes L1 loss function, L2 loss function and cross-entropy loss function;

[0040] Step 54, and based on a preset model parameter optimizer, perform a round of parameter modulation processing on the model parameter θ of the encoder in the direction of minimizing the model optimization objective function, and use the model parameter θ obtained after this round of modulation as the corresponding modulated model parameter at the end of this round of parameter modulation processing;

[0041] Among them, the model parameter optimizer at least includes SGD optimizer, ADAM optimizer, RMSprop optimizer, AdamW optimizer, LBGFS optimizer.

[0042] Preferably, the three-dimensional conformation conversion processing of the first molecular sequence based on a preset cheminformatics tool to obtain the corresponding first conversion conformation specifically includes:

[0043] Input the first molecular sequence into the first processing interface provided by the cheminformatics tool to perform molecular sequence normalization processing to obtain the corresponding first standard sequence; and input the first standard sequence into the second processing interface provided by the cheminformatics tool to perform three-dimensional conformation conversion processing of the molecular sequence to obtain the corresponding first conversion conformation;

[0044] Among them, the first and second processing interfaces are two processing interfaces provided by the chemoinformatics tool; the first processing interface is used to perform a standardized sequence conversion on the molecular sequence input through the interface according to a preset standardized molecular sequence arrangement rule and output the obtained standardized sequence as the interface processing result; the first processing interface outputs the same standardized sequence for different SMIELS sequence expressions of the same molecule; the second processing interface is used to generate a corresponding three-dimensional molecular conformation based on the molecular sequence input through the interface and output the obtained three-dimensional molecular conformation as the interface processing result.

[0045] Preferably, feeding back the corresponding first batch of molecular task reports to the user by processing the batch molecular generation task according to the first converted conformation, the first molecular property vector, the first generated molecular quantity, and the molecular conformation generator specifically includes:

[0046] Step 71, initialize the first counter to 1;

[0047] Step 72, use the first converted conformation and the first molecular property vector as the corresponding molecular conformation X ini and the conditional constraint vector C T input into the molecular conformation generator to perform new molecular conformation generation processing to obtain the corresponding generated molecular conformation X D as the corresponding first generated molecular conformation;

[0048] Step 73, increment the first counter by 1;

[0049] Step 74, identify whether the first counter is greater than the first generated molecular quantity; if so, go to Step 75; if not, return to Step 72;

[0050] Step 75, identify the similarity between every two of the first generated molecular conformations to obtain the corresponding first similarity; and cluster all the obtained first generated molecular conformations based on the first similarity to obtain one or more corresponding first conformation sets;

[0051] Among them, each first conformation set is composed of one or more of the first generated molecular conformations; when the total number of the first generated molecular conformations in the first conformation set is unique, the unique first generated molecular conformation has a first similarity greater than a preset first similarity threshold with any other first generated molecular conformation; when the total number of the first generated molecular conformations in the first conformation set is not unique, the first similarity between every two of the first generated molecular conformations in the set is less than or equal to the first similarity threshold;

[0052] Step 76: Feed back to the user a corresponding first batch molecular task report composed of the first molecular sequence, the first molecular property vector, the first generated molecular quantity, and all the first conformation sets.

[0053] In the second aspect of the embodiments of the present invention, there is provided an apparatus for implementing the processing method of the molecular conformation generator with conditional constraints described in the first aspect above. The apparatus includes: a model construction module, a model initialization module, a user data reception module, a data preprocessing module, and a batch generation task processing module;

[0054] The model construction module is used to construct a molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training; wherein, the molecular conformation generator is used to perform new molecular conformation generation processing according to the input molecular conformation X ini and the conditional constraint vector C T and output the corresponding generated molecular conformation X D ;

[0055] The model initialization module is used to initialize the model parameters of all the Uni-Mol models in the molecular conformation generator based on preset base model parameters; wherein, the base model parameters are the model parameters of a base Uni-Mol model that has completed model pre-training;

[0056] The user data reception module is used to receive the first molecular sequence, the first molecular property vector, and the first generated molecular quantity input by the user; wherein, the first molecular sequence is a one-dimensional SMIELS sequence; the first molecular property vector includes multiple first property data; each first property data corresponds to a specified molecular property; the specified molecular properties at least include melting point, solubility, density, acid / alkaline type, stability score, and reactivity score; when each first property data is 0, it means that there is no constraint on the specified molecular property corresponding to the current property data when generating a new molecular conformation, and when the first property data is not 0, it means that the specified molecular property corresponding to the current property data is constrained by the current property data as the corresponding conditional property data when generating a new molecular conformation;

[0057] The data preprocessing module is used to perform three-dimensional conformation conversion processing on the first molecular sequence based on preset cheminformatics tools to obtain the corresponding first conversion conformation; wherein, the cheminformatics tools at least include OpenBabel and RDKit;

[0058] The batch generation task processing module is used to perform batch molecule generation task processing according to the first converted conformation, the first molecular property vector, the first number of generated molecules, and the molecular conformation generator, and feedback the corresponding first batch molecule task report to the user.

[0059] In the third aspect of the embodiments of the present invention, an electronic device is provided, including: a memory, a processor, and a transceiver;

[0060] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;

[0061] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.

[0062] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.

[0063] The embodiments of the present invention provide a processing method, device, electronic device, and computer-readable storage medium of a molecular conformation generator with conditional constraints. As can be seen from the above content, the embodiments of the present invention first construct an end-to-end molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training, and initialize the model parameters of all Uni-Mol models in the generator based on the model parameters of a base Uni-Mol model that has completed model pre-training; after receiving the molecular sequence, molecular property vector, and number of generated molecules input by the user, first use chemical informatics tools (such as OpenBabel, RDKit, etc.) to perform three-dimensional conformation conversion on the molecular sequence, and then perform batch molecule generation task processing according to the converted conformation, molecular property vector, number of generated molecules, and molecular conformation generator. The molecular conformation generator given by the embodiments of the present invention can generate a large number of molecular structures without experiments, molecular motion simulations, and expert knowledge limitations, and can perform directional conditional constraints on the generated molecular structures according to the set target molecular properties; through the embodiments of the present invention, not only can the design efficiency of molecular design be effectively improved, but also the exploration ability of unknown novel molecular structures can be effectively improved. Description of the Drawings

[0064] Figure 1 It is a schematic diagram of a processing method of a molecular conformation generator with conditional constraints provided in Embodiment 1 of the present invention;

[0065] Figure 2Schematic diagram of the module of the molecular conformation generator provided in the first embodiment of the present invention;

[0066] Figure 3 Structural diagram of the module of a processing device of a molecular conformation generator with conditional constraints provided in the second embodiment of the present invention;

[0067] Figure 4 Schematic diagram of the structure of an electronic device provided in the third embodiment of the present invention. Detailed implementation manners

[0068] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] The first embodiment of the present invention provides a processing method for a molecular conformation generator with conditional constraints. As Figure 1 shown in the schematic diagram of the processing method for a molecular conformation generator with conditional constraints provided in the first embodiment of the present invention, the method mainly includes the following steps:

[0070] Step 1: Construct a molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training.

[0071] Here, the molecular conformation generator in the embodiment of the present invention is used to perform new molecular conformation generation processing according to the input molecular conformation X ini and the conditional constraint vector C T and output the corresponding generated molecular conformation X D .

[0072] It should be noted that the Uni-Mol model in the embodiments of the present invention is used to optimize the input three-dimensional molecular conformation and output the corresponding optimized molecular conformation; this model is an end-to-end three-dimensional molecular conformation optimization model implemented based on the transformer model, and the detailed model design can be understood through the public technical literature "UNI-MOL: A UNIVERSAL 3D MOLECULAR REPRESENTATION LEARNING FRAMEWORK". It can be known from this technical literature that a new three-dimensional molecular conformation can be generated based on the input three-dimensional molecular conformation by the Uni-Mol model; the pre-training scheme of this model can also be completed with reference to the pre-training scheme given in the above technical literature; hereinafter, the Uni-Mol model that has completed pre-training in the embodiments of the present invention is referred to as the base Uni-Mol model, and the model parameters of the base Uni-Mol model are referred to as base model parameters.

[0073] It should be noted that the property prediction model in the embodiments of the present invention is composed of multiple parallel first prediction models; each first prediction model is respectively implemented based on a type of non-linear regression model, and the non-linear regression models mentioned here at least include the XGBoost model, the GBDT model, the Random Forest model, and the MLP model. Each first prediction model in the embodiments of the present invention corresponds to a specified molecular property; the specified molecular property at least includes property types such as melting point, solubility, density, acid / alkaline type, stability score, and reactivity score. The property prediction model in the embodiments of the present invention is used to send the currently input three-dimensional molecular conformation to each first prediction model to obtain the corresponding property prediction data; and all the obtained property prediction data form the corresponding property prediction vector and output; the property prediction vector here is composed of multiple property prediction data. The first prediction model corresponding to the melting point, solubility, density, acid / alkaline type, stability score, or reactivity score in the property prediction model of the embodiments of the present invention is respectively used to perform the corresponding melting point, solubility, density, acid / alkaline type, stability score, or reactivity score prediction processing according to the currently input three-dimensional molecular conformation to obtain the corresponding melting point, solubility, density, acid / alkaline type, stability score, or reactivity score prediction value as a corresponding property prediction data and output.

[0074] It should be noted that the property prediction model in the embodiments of the present invention is a mature model that has been pre-trained. Its training method is carried out by using the conventional supervised training method, and the technical solution for training this property prediction model does not belong to the component of the technical solution of the present invention. Therefore, the training process of the property prediction model will not be elaborated.

[0075] It should be noted that as Figure 2As shown in the module schematic diagram of the molecular conformation generator provided in the first embodiment of the present invention, the first model input end of the molecular conformation generator of the present invention is used to receive the molecular conformation X input by the model ini The second model input end is used to receive the conditional constraint vector C input by the model T , and the model output end is used to output the corresponding generated molecular conformation X D ; The molecular conformation X ini and the generated molecular conformation X D are each a three-dimensional molecular conformation; The conditional constraint vector C T is composed of multiple conditional property data, and each conditional property data corresponds to a specified molecular property.

[0076] As Figure 2 shown, the molecular conformation generator of the present invention is composed of a conformation optimization model, an encoder, a property prediction model, a constraint condition comparison module, an encoding modulation module, and a decoder. The connection relationship of the internal components of the molecular conformation generator is as follows: the input end of the conformation optimization model is connected to the first model input end, and the output end is connected to the first input end of the encoder; the second input end of the encoder is connected to the output end of the encoding modulation module, the first output end is connected to the input end of the property prediction model, and the second output end is connected to the first input end of the constraint condition comparison module; the output end of the property prediction model is connected to the second input end of the constraint condition comparison module; the third input end of the constraint condition comparison module is connected to the second model input end, the first output end is connected to the input end of the encoding modulation module, and the second output end is connected to the input end of the decoder; the output end of the decoder is connected to the model output end.

[0077] The function descriptions of the components of the molecular conformation generator are as follows:

[0078] 1) Conformation optimization model:

[0079] The conformation optimization model of the molecular conformation generator in the embodiment of the present invention is implemented based on a Uni-Mol model;

[0080] This conformation optimization model is used to perform molecular conformation optimization processing on the input molecular conformation X ini and output the corresponding optimized molecular conformation X O ; And send the optimized molecular conformation X O to the encoder;

[0081] 2) Encoder:

[0082] The encoder of the molecular conformation generator in the embodiment of the present invention is composed of a preset first number N E of Uni-Mol models connected in series; The first number N E here is a preset positive integer;

[0083] This encoder is used to save the optimized molecular conformation X received from the conformation optimization model when it is received. O When receiving the optimized molecular conformation X this time, O it is saved; and N E Uni-Mol models are used to perform step-by-step optimization on the optimized molecular conformation X received this time O and the optimized molecular conformation output by the last Uni-Mol model is used as the corresponding encoded output conformation X. E And the encoded output conformation X obtained this time E is sent to the property prediction model and the constraint condition comparison module.

[0084] This encoder is also used to reset its own encoder parameters using the received modulation model parameters when receiving the modulation model parameters sent by the encoding modulation module; and at the end of this parameter reset, N E Uni-Mol models are used again to perform step-by-step optimization on the saved optimized molecular conformation X O and the optimized molecular conformation output by the Nth E Uni-Mol model is used as the latest encoded output conformation X. E And the encoded output conformation X obtained this time E is sent to the property prediction model and the constraint condition comparison module.

[0085] 3) Property prediction model:

[0086] The property prediction model of the molecular conformation generator in the embodiment of the present invention is used to send the currently input encoded output conformation X E to each first prediction model to obtain corresponding property prediction data; and all the property prediction data obtained this time are used to form a corresponding property prediction vector C. P And the property prediction vector C obtained this time P is sent to the constraint condition comparison module.

[0087] 4) Constraint condition comparison module:

[0088] The constraint condition comparison module of the molecular conformation generator in the embodiment of the present invention is used to perform constraint condition comparison processing on the input property prediction vector C P and the condition constraint vector C T to obtain a corresponding condition satisfaction status; and identify the condition satisfaction status; if the condition satisfaction status is not meeting the constraints, the latest property prediction vector C P and the condition constraint vector C T are sent to the encoding modulation module; if the condition satisfaction status is meeting the constraints, the latest encoded output conformation X E is sent to the decoder.

[0089] Among them, the condition satisfaction status includes not satisfying the constraint condition and satisfying the constraint condition;

[0090] Furthermore, when the constraint condition comparison module predicts the vector C according to the nature of the input P and the condition constraint vector C T to perform the constraint condition comparison process to obtain the corresponding condition satisfaction status, the specific processing steps are composed of the following steps A1 - A4:

[0091] Step A1, the constraint condition comparison module records the non - zero condition property data in the condition constraint vector C T as the corresponding first comparison property data;

[0092] Step A2, and perform a round of traversal on all the first comparison property data; and during this round of traversal, take the currently traversed first comparison property data as the corresponding current threshold data; and take the specified molecular property corresponding to the current threshold data as the corresponding current property; and take the property prediction data corresponding to the current property in the property prediction vector C P as the corresponding current prediction data; and query the preset first correspondence table based on the current property, and extract the first relative error range field of the first correspondence record whose first property field matches the current property in the first correspondence table as the corresponding current relative error range; and form a corresponding prediction - threshold data pair from the current prediction data and the current threshold data; and calculate the relative error corresponding to the current prediction data and the current threshold data, the current relative error = |current prediction data - current threshold data| / current threshold data; and identify whether the current relative error satisfies the corresponding current relative error range, if it satisfies, set the corresponding single - item error comparison result to satisfy, if it does not satisfy, set the corresponding single - item error comparison result to not satisfy; and at the end of this round of traversal, bring all the obtained prediction - threshold data pairs into the preset error evaluation function for calculation to obtain the corresponding first evaluation value;

[0093] Among them, the first correspondence table is a correspondence table reflecting the correspondence between properties and relative error ranges, and is composed of multiple first correspondence records; each first correspondence record corresponds to a specified molecular property; the first correspondence record includes a first property field and a first relative error range field; the first property field is used to store a corresponding specified molecular property, and the first relative error range field is used to store a corresponding relative error range; the single - item error comparison result includes satisfy and not satisfy; the error evaluation function at least includes the RMSE function;

[0094] Step A3, identify whether the first evaluation value meets the preset first evaluation value range. If it meets, set the corresponding full-item comparison result to meet; if it does not meet, set the corresponding full-item comparison result to not meet;

[0095] Among them, the first evaluation value range is a preset evaluation value range; the full-item comparison result includes meet and not meet;

[0096] Step A4, identify all the single-item error comparison results and the full-item comparison results obtained; if any single-item error comparison result is not meet or the full-item comparison result is not meet, set the corresponding condition satisfaction status to not meet the constraint condition; if all single-item error comparison results and the full-item comparison result are meet, set the corresponding condition satisfaction status to meet the constraint condition;

[0097] 5) Encoding modulation module:

[0098] The encoding modulation module of the molecular conformation generator according to the embodiment of the present invention is used to perform a round of parameter modulation processing on the encoder according to the property prediction vector C received this time P and the conditional constraint vector C T to obtain the corresponding modulation model parameters; and send the modulation model parameters obtained this time to the encoder;

[0099] Further, when the encoding modulation module performs a round of parameter modulation processing on the encoder according to the property prediction vector C received this time P and the conditional constraint vector C T to obtain the corresponding modulation model parameters, the specific processing steps are composed of the following steps B1-B4:

[0100] Step B1, the encoding modulation module records the conditional property data with each value of 0 in the conditional constraint vector C T as the corresponding invalid property data;

[0101] Step B2, and reset the property prediction data corresponding to each invalid property data in the property prediction vector C P to 0;

[0102] Step B3, and substitute the reset property prediction vector C P and the conditional constraint vector C T into the preset model optimization objective function;

[0103] Here, the model optimization objective function of the embodiment of the present invention is:

[0104]

[0105] Among them, θ is the model parameter of the encoder; L Mis a preset model loss function, and the model loss function L M at least includes the L1 loss function, the L2 loss function, and the cross-entropy loss function;

[0106] Step B4, and based on a preset model parameter optimizer, perform a round of parameter modulation processing on the model parameters θ of the encoder in the direction of minimizing the model optimization objective function, and use the model parameters θ obtained after this round of modulation as the corresponding modulated model parameters at the end of this round of parameter modulation processing;

[0107] Among them, the model parameter optimizer at least includes the SGD optimizer, the ADAM optimizer, the RMSprop optimizer, the AdamW optimizer, and the LBGFS optimizer;

[0108] 6) Decoder:

[0109] The decoder of the molecular conformation generator in the embodiment of the present invention is composed of a preset second number N D of Uni-Mol models connected in series; the second number N here D is a preset positive integer, and N D <N E ;

[0110] This decoder is used to use N D Uni-Mol models to perform step-by-step optimization on the input encoded output conformation X E and use the optimized molecular conformation output by the N D th Uni-Mol model as the corresponding generated molecular conformation X D and output it.

[0111] Step 2, initialize the model parameters of all Uni-Mol models in the molecular conformation generator based on the preset base model parameters.

[0112] Here, the base model parameters of the embodiment of the present invention are the model parameters of a base Uni-Mol model that has completed model pre-training.

[0113] Step 3, receive the first molecular sequence, the first molecular property vector, and the first generated molecular quantity input by the user.

[0114] Here, the first molecular sequence is a one-dimensional SMIELS sequence; the first molecular property vector includes multiple first property data; each first property data corresponds to a specified molecular property; when each first property data is 0, it means that the specified molecular property corresponding to the current property data is not restricted when generating a new molecular conformation, and when the first property data is not 0, it means that the specified molecular property corresponding to the current property data is restricted by the corresponding conditional property data when generating a new molecular conformation; the first generated molecular quantity is a positive integer.

[0115] Step 4: Perform three-dimensional conformation conversion processing on the first molecular sequence based on a preset chemoinformatics tool to obtain the corresponding first converted conformation;

[0116] Among them, the chemoinformatics tool includes at least OpenBabel and RDKit;

[0117] Specifically, it includes: inputting the first molecular sequence into the first processing interface provided by the chemoinformatics tool to perform molecular sequence standardization processing to obtain the corresponding first standard sequence; and inputting the first standard sequence into the second processing interface provided by the chemoinformatics tool to perform three-dimensional conformation conversion processing of the molecular sequence to obtain the corresponding first converted conformation.

[0118] Here, the first and second processing interfaces of the embodiments of the present invention are two processing interfaces provided by the chemoinformatics tool; among them: 1) The first processing interface is used to perform standardized sequence conversion on the molecular sequence input into the interface according to a preset standardized molecular sequence arrangement rule and output the obtained standardized sequence as the interface processing result. It should be noted that the first processing interface will output the same standardized sequence for different SMIELS sequence expressions of the same molecule; 2) The second processing interface is used to generate the corresponding three-dimensional molecular conformation according to the molecular sequence input into the interface and output the obtained three-dimensional molecular conformation as the interface processing result.

[0119] Step 5: Perform batch molecular generation task processing according to the first converted conformation, the first molecular property vector, the first number of generated molecules, and the molecular conformation generator, and feedback the corresponding first batch molecular task report to the user;

[0120] Specifically, it includes: Step 51: Initialize the first counter to 1;

[0121] Step 52: Use the first converted conformation and the first molecular property vector as the corresponding molecular conformation X ini and the conditional constraint vector C T Input them into the molecular conformation generator to perform new molecular conformation generation processing to obtain the corresponding generated molecular conformation X D as the corresponding first generated molecular conformation;

[0122] Step 53: Increment the first counter by 1;

[0123] Step 54: Identify whether the first counter is greater than the first number of generated molecules; if so, go to Step 55; if not, return to Step 52;

[0124] Step 55: Identify the similarity between every two first generated molecular conformations to obtain the corresponding first similarity; and cluster all the obtained first generated molecular conformations based on the first similarity to obtain one or more corresponding first conformation sets;

[0125] Wherein, each first conformation set is composed of one or more first generated molecular conformations; when the total number of first generated molecular conformations in the first conformation set is unique, the first similarity between the unique first generated molecular conformation and any other first generated molecular conformation is greater than a preset first similarity threshold; when the total number of first generated molecular conformations in the first conformation set is not unique, the first similarity between every two first generated molecular conformations in the set is less than or equal to the first similarity threshold;

[0126] Here, in the embodiment of the present invention, when identifying the similarity between every two first generated molecular conformations, the respective first generated molecular conformations are pre-converted into corresponding first conformation vectors, and then the similarity between every two first conformation vectors is calculated based on a conventional vector similarity algorithm (such as the cosine similarity algorithm, etc.) to obtain the corresponding first similarity;

[0127] Step 56: Feed back to the user a corresponding first batch molecular task report composed of the first molecular sequence, the first molecular property vector, the first generated molecular quantity, and all the first conformation sets.

[0128] Figure 3 FIG. is a module structure diagram of a processing device of a molecular conformation generator with conditional constraints provided in the second embodiment of the present invention. The device is a terminal device or a server for implementing the foregoing method embodiment, or may be a device capable of enabling the foregoing terminal device or server to implement the foregoing method embodiment. For example, the device may be a device or a chip system of the foregoing terminal device or server. As Figure 3 shown, the device includes: a model construction module 201, a model initialization module 202, a user data receiving module 203, a data preprocessing module 204, and a batch generation task processing module 205.

[0129] The model construction module 201 is used to construct a molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training; wherein, the molecular conformation generator is used to perform new molecular conformation generation processing on the input molecular conformation X ini and the conditional constraint vector C T and output the corresponding generated molecular conformation X D .

[0130] The model initialization module 202 is used to initialize the model parameters of all Uni-Mol models in the molecular conformation generator based on the preset base model parameters; wherein, the base model parameters are the model parameters of a base Uni-Mol model that has completed model pre-training.

[0131] The user data receiving module 203 is used to receive the first molecular sequence, the first molecular property vector, and the first number of generated molecules input by the user; wherein, the first molecular sequence is a one-dimensional SMIELS sequence; the first molecular property vector includes multiple first property data; each first property data corresponds to a specified molecular property; the specified molecular properties at least include melting point, solubility, density, acid / alkaline type, stability score, and reactivity score; when each first property data is 0, it means that the specified molecular property corresponding to the current property data is not constrained when generating a new molecular conformation, and when the first property data is not 0, it means that the corresponding specified molecular property is constrained by the current property data as the corresponding conditional property data when generating a new molecular conformation.

[0132] The data preprocessing module 204 is used to perform three-dimensional conformation conversion processing on the first molecular sequence based on the preset cheminformatics tools to obtain the corresponding first converted conformation; wherein, the cheminformatics tools at least include OpenBabel and RDKit.

[0133] The batch generation task processing module 205 is used to perform batch molecular generation task processing according to the first converted conformation, the first molecular property vector, the first number of generated molecules, and the molecular conformation generator, and feedback the corresponding first batch molecular task report to the user.

[0134] The processing device of the molecular conformation generator with conditional constraints provided by the embodiments of the present invention can execute the method steps in the above method embodiments, and its implementation principle and technical effects are similar, and will not be described in detail here.

[0135] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; they can also be partially implemented in the form of software called by processing elements and partially implemented in the form of hardware. For example, the model construction module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determined module can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0136] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0137] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0138] Figure 4 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server for implementing the method of the foregoing embodiments, or may be a terminal device or a server for implementing the method of the foregoing embodiments and connected to the foregoing terminal device or server. As Figure 4 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 to complete various processing functions and implement the processing steps described in the foregoing method embodiments. Preferably, the electronic device according to the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.

[0139] In Figure 4The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.

[0140] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0141] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods and processing procedures provided in the above embodiments.

[0142] An embodiment of the present invention provides a processing method, apparatus, electronic device, and computer-readable storage medium for a molecular conformation generator with conditional constraints. As can be seen from the above, in the embodiment of the present invention, an end-to-end molecular conformation generator is first constructed based on multiple Uni-Mol models and a property prediction model that has completed model training, and the model parameters of all Uni-Mol models in the generator are initialized based on the model parameters of a base Uni-Mol model that has completed model pre-training; after receiving the molecular sequence, molecular property vector, and the number of generated molecules input by the user, first use chemical informatics tools (such as OpenBabel, RDKit, etc.) to perform three-dimensional conformation conversion on the molecular sequence, and then perform batch molecular generation task processing according to the converted conformation, molecular property vector, the number of generated molecules, and the molecular conformation generator. Through the molecular conformation generator given by the embodiment of the present invention, a large number of molecular structures can be generated without experiments, without molecular motion simulation, and without being restricted by expert knowledge, and the generated molecular structures can be directionally conditionally constrained according to the set target molecular properties; through the embodiment of the present invention, not only can the design efficiency of molecular design be effectively improved, but also the exploration ability of unknown novel molecular structures can be effectively improved.

[0143] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0144] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A processing method for a molecular conformation generator with conditional constraints, characterized in that: The method comprises: A molecular conformation generator is constructed based on multiple Uni-Mol models and a property prediction model that has completed model training; wherein the molecular conformation generator is used to generate a molecular conformation X according to the input molecular conformation X. ini and the conditional constraint vector C T Perform new molecular conformation generation processing and output the corresponding generated molecular conformation X D ; Initializing the model parameters of all the Uni-Mol models in the molecular conformation generator based on the preset base model parameters; wherein the base model parameters are the model parameters of a base Uni-Mol model that has completed model pre-training; Receive a first molecular sequence, a first molecular property vector and a first number of generated molecules input by a user; wherein the first molecular sequence is a one-dimensional SMIELS sequence; the first molecular property vector includes a plurality of first property data; each of the first property data corresponds to a specified molecular property; the specified molecular properties at least include melting point, solubility, density, acid / alkaline type, stability score and reactivity score; when each of the first property data is 0, it indicates that when generating a new molecular conformation, no constraint is imposed on the specified molecular property corresponding to the current property data; when the first property data is not 0, it indicates that when generating a new molecular conformation, the current property data is used as the corresponding conditional property data to constrain the corresponding specified molecular property; Based on a preset chemical informatics tool, a three-dimensional conformation conversion process is performed on the first molecular sequence to obtain a corresponding first conversion conformation; wherein the chemical informatics tool includes at least OpenBabel and RDKit; Performing batch molecule generation task processing according to the first conversion conformation, the first molecular property vector, the first number of generated molecules and the molecular conformation generator to obtain a corresponding first batch molecule task report and provide feedback to the user; The Uni-Mol model is used to optimize the input three-dimensional molecular conformation and output the corresponding optimized molecular conformation; The property prediction model is composed of a plurality of parallel first prediction models; each of the first prediction models is implemented based on a type of nonlinear regression model, and the nonlinear regression model includes at least an XGBoost model, a GBDT model, a RandomForest model and an MLP model; each of the first prediction models corresponds to one of the specified molecular properties; The property prediction model is used to send the currently input three-dimensional molecular conformation to each of the first prediction models to obtain corresponding property prediction data; and all the obtained property prediction data are used to form a corresponding property prediction vector and output it; the property prediction vector is composed of a plurality of the property prediction data; The first prediction model corresponding to the melting point, solubility, density, acid / alkalinity type, stability score or reactivity score in the property prediction model is used to perform corresponding melting point, solubility, density, acid / alkalinity type, stability score or reactivity score prediction processing according to the currently input three-dimensional molecular conformation to obtain the corresponding melting point, solubility, density, acid / alkalinity type, stability score or reactivity score prediction value as a corresponding property prediction data and output it; The first model input terminal of the molecular conformation generator is used to receive the molecular conformation X input by the model. ini The second model input terminal is used to receive the conditional constraint vector C of the model input T The model output terminal is used to output the corresponding generated molecular conformation X D ; The molecular conformation X ini and the generated molecular conformation X D Each is a three-dimensional molecular conformation; the conditional constraint vector C T It is composed of a plurality of the conditional property data, each of which corresponds to one of the specified molecular properties; The molecular conformation generator includes a conformation optimization model, an encoder, the property prediction model, a constraint comparison module, a coding modulation module and a decoder; The input end of the conformation optimization model is connected to the input end of the first model, and the output end is connected to the first input end of the encoder; the second input end of the encoder is connected to the output end of the coding and modulation module, the first output end is connected to the input end of the property prediction model, and the second output end is connected to the first input end of the constraint comparison module; the output end of the property prediction model is connected to the second input end of the constraint comparison module; the third input end of the constraint comparison module is connected to the second model input end, the first output end is connected to the input end of the coding and modulation module, and the second output end is connected to the input end of the decoder; the output end of the decoder is connected to the output end of the model; The conformation optimization model is implemented based on the Uni-Mol model; The conformation optimization model is used to input the molecular conformation X ini Perform molecular conformation optimization and output the corresponding optimized molecular conformation X O ; and the optimized molecular conformation X O Sending to the encoder; The encoder is composed of a preset first number N E The Uni-Mol model is connected in series; The encoder is used to receive the optimized molecular conformation X sent by the conformation optimization model O When the optimized molecular conformation X received this time is O to save; and use N E The Uni-Mol model is used to optimize the molecular conformation X received this time. O Perform step-by-step optimization and use the optimized molecular conformation output by the last Uni-Mol model as the corresponding coded output conformation X E ; And the encoding output conformation X obtained this time E Sending to the property prediction model and the constraint condition comparison module; The encoder is also used to reset its own encoder parameters using the modulation model parameters received this time when receiving the modulation model parameters sent by the coding modulation module; and to use N again when the parameter reset is completed. E The Uni-Mol model is used to store the optimized molecular conformation X O Perform step-by-step optimization and set the Nth E The optimized molecular conformation output by the Uni-Mol model is used as the latest encoded output conformation X E ; And the encoding output conformation X obtained this time E Sending to the property prediction model and the constraint condition comparison module; The property prediction model is used to convert the encoding output conformation X of the current input E Send the corresponding property prediction data to each of the first prediction models; and form a corresponding property prediction vector C from all the property prediction data obtained this time P ; and the property prediction vector C obtained this time P Sending to the constraint comparison module; The constraint condition comparison module is used to predict the vector C according to the input property P and the conditional constraint vector C T Perform constraint condition comparison processing to obtain the corresponding condition satisfaction state; and identify the condition satisfaction state; if the condition satisfaction state is that the constraint condition is not satisfied, then the latest property prediction vector C P and the conditional constraint vector C T Send to the coding and modulation module; if the condition is satisfied, the latest coding output constellation X E Sending to the decoder; wherein the condition satisfaction state includes not satisfying the constraint condition and satisfying the constraint condition; The coding and modulation module is used to predict the vector C according to the property received at that time P and the conditional constraint vector C T Performing a round of parameter modulation processing on the encoder to obtain the corresponding modulation model parameters; and sending the modulation model parameters obtained this time to the encoder; The decoder is composed of a preset second number N D The Uni-Mol model is connected in series; The decoder is used to use N D The Uni-Mol model encodes the output conformation X of the input E Perform step-by-step optimization and set the Nth D The optimized molecular conformation output by the Uni-Mol model is used as the corresponding generated molecular conformation X D And output; The performing batch molecule generation task processing according to the first conversion conformation, the first molecular property vector, the first number of generated molecules and the molecular conformation generator to obtain a corresponding first batch molecule task report and feeding back to the user specifically includes: Step 71, initializing the first counter to 1; Step 72: taking the first conversion conformation and the first molecular property vector as the corresponding molecular conformation X ini and the conditional constraint vector C T Input the molecular conformation generator to perform new molecular conformation generation processing to obtain the corresponding generated molecular conformation X D As the corresponding first generated molecular conformation; Step 73, adding 1 to the first counter; Step 74, identifying whether the first counter is greater than the first generated molecule quantity; if so, go to step 75; if not, return to step 72; Step 75, identifying the similarity between every two of the first generated molecular conformations to obtain a corresponding first similarity; and clustering all the first generated molecular conformations obtained based on the first similarity to obtain one or more corresponding first conformation sets; Wherein, each first conformation set is composed of one or more first generated molecular conformations; when the total number of the first generated molecular conformations in the first conformation set is unique, the first similarity between the unique first generated molecular conformation and any other first generated molecular conformations is greater than a preset first similarity threshold; when the total number of the first generated molecular conformations in the first conformation set is not unique, the first similarity between every two first generated molecular conformations in the set is less than or equal to the first similarity threshold; Step 76: Feedback to the user the first batch of molecular task report corresponding to the first molecular sequence, the first molecular property vector, the first generated molecular quantity and all the first conformation sets.

2. The processing method of the molecular conformation generator with conditional constraints according to claim 1, characterized in that: The prediction vector C according to the property of the input P and the conditional constraint vector C T The constraint condition comparison process is performed to obtain the corresponding condition satisfaction status, which specifically includes: The constraint condition comparison module converts the condition constraint vector C T Each conditional property data that is not 0 is recorded as the corresponding first comparison property data; And perform a round of traversal on all the first comparison property data; and in this round of traversal, use the first comparison property data currently traversed as the corresponding current threshold data; and use the specified molecular property corresponding to the current threshold data as the corresponding current property; and use the property prediction vector C P The property prediction data corresponding to the current property is used as the corresponding current prediction data; and based on the current property, a preset first correspondence table is queried, and the first relative error range field of the first correspondence record in which the first property field in the first correspondence table matches the current property is extracted as the corresponding current relative error range; and a corresponding prediction-threshold data pair is formed by the current prediction data and the current threshold data; and a relative error calculation is performed on the current prediction data and the current threshold data, and the corresponding current relative error = |current prediction data-current threshold data| / current threshold data; and whether the current relative error satisfies the corresponding current relative error range is identified, and if so, the corresponding single error comparison result is set to be satisfied, and if not, the single error comparison result is set to be satisfied. The corresponding single error comparison result is not satisfied; and at the end of this round of traversal, all the obtained prediction-threshold data pairs are brought into the preset error evaluation function to calculate the corresponding first evaluation value; wherein, the first correspondence table is a correspondence table reflecting the property-relative error range correspondence, which is composed of multiple first correspondence records; each first correspondence record corresponds to a specified molecular property; the first correspondence record includes the first property field and the first relative error range field; the first property field is used to store a corresponding specified molecular property, and the first relative error range field is used to store a corresponding relative error range; the single error comparison results include satisfied and not satisfied; the error evaluation function at least includes the RMSE function; and identifying whether the first evaluation value satisfies a preset first evaluation value range, and if so, setting the corresponding full comparison result as satisfied, and if not, setting the corresponding full comparison result as unsatisfied; wherein the full comparison result includes satisfied and unsatisfied; And all the obtained single error comparison results and the full comparison results are identified; if any of the single error comparison results is not satisfied or the full comparison result is not satisfied, the corresponding condition satisfaction state is set to not satisfying the constraint condition; if all the single error comparison results and the full comparison results are satisfied, the corresponding condition satisfaction state is set to satisfying the constraint condition.

3. The processing method of the molecular conformation generator with conditional constraints according to claim 1, characterized in that: The prediction vector C based on the property received at that time P and the conditional constraint vector C T Performing a round of parameter modulation processing on the encoder to obtain the corresponding modulation model parameters specifically includes: Step 51: the coding and modulation module converts the conditional constraint vector C T Each of the conditional property data that is 0 is recorded as the corresponding invalid property data; Step 52, and transform the property prediction vector C P The property prediction data corresponding to each of the invalid property data is reset to 0; Step 53, and complete the reset of the property prediction vector C P and the conditional constraint vector C T Bring in the preset model optimization objective function; Among them, the model optimization objective function is: ; θ is the model parameter of the encoder; L M is a preset model loss function, wherein the model loss function L M At least including L1 loss function, L2 loss function and cross entropy loss function; Step 54, based on a preset model parameter optimizer, a round of parameter modulation processing is performed on the model parameter θ of the encoder in a direction in which the model optimization objective function reaches a minimum value, and at the end of this round of parameter modulation processing, the model parameter θ obtained after this round of modulation is used as the corresponding modulation model parameter; Among them, the model parameter optimizer includes at least SGD optimizer, ADAM optimizer, RMSprop optimizer, AdamW optimizer, and LBGFS optimizer.

4. The processing method of the molecular conformation generator with conditional constraints according to claim 1, characterized in that: The performing three-dimensional conformation conversion processing on the first molecular sequence based on a preset chemical informatics tool to obtain a corresponding first conversion conformation specifically includes: Inputting the first molecular sequence into a first processing interface provided by the chemical informatics tool to perform molecular sequence standardization processing to obtain a corresponding first standard sequence; and inputting the first standard sequence into a second processing interface provided by the chemical informatics tool to perform molecular sequence three-dimensional conformation conversion processing to obtain the corresponding first conversion conformation; Among them, the first and second processing interfaces are two processing interfaces provided by the chemical informatics tool; the first processing interface is used to perform standardized sequence conversion on the molecular sequence input into the interface according to a preset standardized molecular sequence arrangement rule and output the obtained standardized sequence as the interface processing result; the first processing interface outputs the same standardized sequence for different SMIELS sequence expressions of the same molecule; the second processing interface is used to generate a corresponding three-dimensional molecular conformation according to the molecular sequence input into the interface and output the obtained three-dimensional molecular conformation as the interface processing result.

5. A device for executing the processing method of the molecular conformation generator with conditional constraints according to any one of claims 1 to 4, characterized in that: The device comprises: a model building module, a model initialization module, a user data receiving module, a data preprocessing module and a batch generation task processing module; The model building module is used to build a molecular conformation generator based on multiple Uni-Mol models and a property prediction model that has completed model training; wherein the molecular conformation generator is used to generate a molecular conformation X according to the input molecular conformation X. ini and the conditional constraint vector C T Perform new molecular conformation generation processing and output the corresponding generated molecular conformation X D ; The model initialization module is used to initialize the model parameters of all the Uni-Mol models in the molecular conformation generator based on the preset base model parameters; wherein the base model parameters are the model parameters of a base Uni-Mol model that has completed model pre-training; The user data receiving module is used to receive a first molecular sequence, a first molecular property vector and a first number of generated molecules input by a user; wherein the first molecular sequence is a one-dimensional SMIELS sequence; the first molecular property vector includes a plurality of first property data; each of the first property data corresponds to a specified molecular property; the specified molecular properties include at least melting point, solubility, density, acid / alkaline type, stability score and reactivity score; when each of the first property data is 0, it indicates that when generating a new molecular conformation, the specified molecular property corresponding to the current property data is not constrained, and when the first property data is not 0, it indicates that when generating a new molecular conformation, the corresponding specified molecular property is constrained with the current property data as the corresponding conditional property data; The data preprocessing module is used to perform three-dimensional conformation conversion processing on the first molecular sequence based on a preset chemical informatics tool to obtain a corresponding first conversion conformation; wherein the chemical informatics tool includes at least OpenBabel and RDKit; The batch generation task processing module is used to perform batch molecule generation task processing according to the first conversion conformation, the first molecular property vector, the first number of generated molecules and the molecular conformation generator to obtain a corresponding first batch molecule task report and feedback it to the user.

6. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is used to couple with the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 4; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Method and device for generating small molecule compound based on conditional diffusion model

    CN118230853A

  • Conditional molecule generation method and device based on attribute values

    CN118506901A