Molecular de novo design method and system based on physical spectrum structure activity rule

Through the molecular de novo design method based on the physical "spectral structure effect" rule, the problem of dimensional disasters in molecular generation is solved, efficient and accurate molecular generation is achieved, and the need for exploration of chemical space is reduced.

CN120089238APending Publication Date: 2025-06-03ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242622.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing molecular generation method directly generates molecular three-dimensional structures from one-dimensional user needs, which has a dimensional disaster problem.

Method used

Through the molecular de novo design method based on the physical "spectral structure effect" rule, the molecular properties required by the user are first mapped to the spectral data, and then map the spectral data to the molecular smiles data. The molecules that meet the conditions are generated through molecular infrared spectroscopy to avoid complex mapping directly from one-dimensional to three-dimensional.

Benefits of technology

It reduces the impact of dimensional disasters, improves the efficiency and accuracy of molecular generation, reduces the need for exploration of huge chemical space, and achieves the rapid generation of molecules that meet user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089238A_ABST
    Figure CN120089238A_ABST
Patent Text Reader

Abstract

The invention provides a molecular de novo design method based on a physical spectrum structure activity rule, and belongs to the field of molecular generation, the method comprises the following steps: obtaining molecular smiles expression, and respectively calculating molecular properties and molecular infrared spectrum according to the molecular smiles expression; training a preset one-dimensional Diffusion model based on the molecular property and the molecular infrared spectrum to obtain a trained Diffusion model, training a preset Transform model based on the molecular smiles representation and the molecular infrared spectrum to obtain a trained Transform model, and integrating the trained Diffusion model and the trained Transform model to obtain a molecular prediction model; inputting the property of a to-be-designed molecule into the molecule prediction model to obtain a molecule smiles expression; screening the molecular smiles representation to obtain a reasonable molecular smiles representation, and generating and optimizing a molecular conformation based on the reasonable molecular smiles representation to obtain a to-be-designed molecule; the invention further provides a molecular de novo design system. Molecules meeting conditions are generated through a molecular infrared spectrum, and a large dimension crossing problem is decomposed into two relatively simple dimension conversion problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of molecular generation, and particularly to a method and system for de novo molecular design based on physical "spectrum-structure-activity" rules. Background Art

[0002] With the continuous development of fields such as chemistry, biology, and materials science, an in-depth understanding of the relationship between molecular structure and properties has become crucial. As an emerging generative model, the Diffusion model has received extensive attention in the field of molecular research at home and abroad. Researchers have continuously improved and optimized the Diffusion model to enable it to more effectively generate molecules with specific properties. For example, by adjusting the parameters and architecture of the Diffusion model, precise control of the molecular structure can be achieved, and molecules with specific properties and functions can be generated. A research team from the Technion - Israel Institute of Technology and the University of Venice in Italy proposed a guided diffusion model for inverse molecular design: GaUDI, which combines an equivariant graph neural network for property prediction and a generative diffusion model. GaUDI demonstrated improved conditional design, generating molecules with optimal properties, even exceeding the original distribution, and proposing molecules better than those in the dataset.

[0003] Research scholars have achieved a series of results in the improvement and application of artificial intelligence algorithms. In view of the characteristics of molecular research, the Diffusion model and others are optimized and improved to enhance their performance in tasks such as molecular generation and property prediction. By improving the diffusion process and sampling strategy of the Diffusion model, the convergence speed of the model is accelerated, and the quality of the generated molecules is improved. For example, a paper titled "3d Equivariant Diffusion For Target-Aware Molecule Generation And Affinity Prediction" published by the team of Professor Jianzhu Ma from Tsinghua University in ICLR 2023 developed a non-autoregressive, rotation- and translation-invariant, pocket-conditioned molecular diffusion generative model TargetDiff by optimizing the internal structure of the Diffusion model. This model can generate molecules with more realistic 3D structures and better affinity for protein targets, showing obvious advantages compared to the benchmark model.

[0004] At present, all molecular generation methods directly generate the three-dimensional structure, molecular graph or SMILES of molecules based on the Diffusion model. However, this method has obvious defects. Starting from the one-dimensional user requirements and directly generating molecular information with three-dimensional structure through a deep learning model, there is a hidden danger of the curse of dimensionality. Generating three-dimensional information by crossing one dimension from one dimension, even if the generated molecular set contains molecules that meet the user's requirements, it is necessary to screen out molecules that meet the conditions through huge computing power analysis from a huge and complex chemical space. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: how to solve the problem of the curse of dimensionality existing in the direct generation of the three-dimensional structure of molecules from one-dimensional user requirements in the existing molecular generation methods.

[0006] The present invention solves the above technical problem through the following technical solutions: a method for de novo design of molecules based on the physical "spectrum-structure-activity" rule, the method comprising:

[0007] Obtain the SMILES representation of the molecule, and calculate the molecular properties and the molecular infrared spectrum respectively according to the SMILES representation of the molecule;

[0008] Train a preset one-dimensional Diffusion model based on the molecular properties and the molecular infrared spectrum to obtain a trained Diffusion model, train a preset Transformer model based on the SMILES representation of the molecule and the molecular infrared spectrum to obtain a trained Transformer model, and integrate the trained Diffusion model and the trained Transformer model to obtain a molecular prediction model;

[0009] Input the properties of the molecule to be designed into the molecular prediction model to obtain the SMILES representation of the molecule;

[0010] Screen the SMILES representation of the molecule to obtain a reasonable SMILES representation of the molecule, generate and optimize the molecular conformation based on the reasonable SMILES representation of the molecule to obtain the molecule to be designed.

[0011] The present invention first inputs the information of several dimensions of the molecular properties required by the user into the Diffusion model to obtain spectral data, maps from the limited molecular property dimensions to the spectral dimensions, and the dimensions are relatively controllable. Then, the spectral data is input into the Transformer model to obtain molecular SMILES data, and molecules that meet the conditions are generated through the molecular infrared spectrum. This is also a mapping in a relatively low-dimensional space with certain physical and chemical significance, avoiding the complex mapping from one dimension directly to the three-dimensional structure of the molecule, and decomposing the large dimension crossing problem into two relatively simple dimension conversion problems, thereby reducing the impact of the curse of dimensionality.

[0012] Preferably, the process of calculating the molecular infrared spectrum according to the molecular SMILES representation includes:

[0013] Construct multiple three-dimensional molecular conformations for each molecule in the form of molecular SMILES through the RDKIT library, preliminarily optimize the molecular stability with the MMFF force field, select the stable molecular conformation from multiple three-dimensional molecular conformations, obtain the spatial coordinates of each atom in the molecule, and save them to the molecular conformation file;

[0014] Batch optimize the molecular conformations by the DFT method to obtain the optimized molecular conformations;

[0015] Calculate the intensities of the molecule at multiple frequencies according to the optimized molecular conformation to form a discrete molecular infrared spectrum, and save the discrete molecular infrared spectrum in the output log file of Gaussian software;

[0016] Extract the corresponding frequencies and intensities of the molecule from the output log file, expand the frequencies of the discrete molecular infrared spectrum to 0 - 4000 cm -1 , set the broadening of the molecular infrared spectrum to 10, normalize the intensities, and extract the molecular infrared spectrum data.

[0017] Preferably, the method for batch optimizing molecular conformations is:

[0018] Based on a one-key batch processing script of the slurm scheduling system, submit the task queue in one key, optimize the molecular conformation by the DFT method, the functional used for optimizing the molecular conformation is B3LYP, and the basis set used is TZVP;

[0019] Based on the one-key export and normalization of molecular data of Gaussian software, obtain the optimized molecular conformation file.

[0020] Preferably, the process of training a preset one-dimensional Diffusion model based on molecular properties and molecular infrared spectra includes:

[0021] Modify the feature extraction layer of the existing Diffusion model to one-dimensional convolution for capturing local temporal features. Replace all Conv2D layers in the Unet of the existing Diffusion model with Conv1D layers. In downsampling, change the max pooling method from MaxPool2D to MaxPool1D. In upsampling, change the transposed convolution ConvTranspose2D to ConvTranspose1D. First, expand the length through linear interpolation of Upsample, and then refine the features with Conv1D. Use fully connected layers and self-attention mechanisms in the skip connection part for feature extraction to obtain the preset one-dimensional Diffusion model;

[0022] Input the molecular properties and molecular infrared spectra into the preset one-dimensional Diffusion model, perform epoch iterations of training, and update the parameters of the preset one-dimensional Diffusion model through the backpropagation algorithm to obtain the trained Diffusion model.

[0023] Preferably, the process of training the preset Transformer model based on the molecular smiles representation and molecular infrared spectra includes:

[0024] Add a one-dimensional convolutional feature extraction layer in front of the preset Transformer model to extract the feature information in the input molecular infrared spectra;

[0025] Set up dictionary mapping. Input the molecular smiles representation and molecular infrared spectra into the preset Transformer model, use the molecular smiles representation as the label, perform epoch iterations of training, and update the parameters of the Transformer model through the backpropagation algorithm to obtain the trained Transformer model.

[0026] Preferably, the process of setting up dictionary mapping includes:

[0027] In the model data input, for the molecular smiles string in the label, through dictionary mapping, map each character into a number, and the mapping of all characters in a molecular smiles forms a one-dimensional vector and inputs it into the Transformer model;

[0028] In the model prediction output, calculate the probability of each position in the vector output by the model through softmax, form a character prediction probability branch tree, take the top-n smiles, and through dictionary mapping, convert the autoregressively generated one-dimensional numerical vector into a molecular smiles representation.

[0029] Preferably, the process of screening the molecular SMILES representation to obtain a reasonable molecular SMILES representation includes:

[0030] Screen out the molecular SMILES representations that conform to the grammar from the molecular SMILES representations generated by the molecular prediction model according to the grammar of the molecular SMILES representation;

[0031] Screen out reasonable molecular SMILES representations from the molecular SMILES representations that conform to the grammar through the Python toolkit of RDKIT. The reasonable molecular SMILES representations meet the property requirements QED, PLogP, and SA of the prediction model input.

[0032] Preferably, the process of generating and optimizing the molecular conformation based on the reasonable molecular SMILES representation includes:

[0033] Use the Python toolkit of RDKIT, input the reasonable molecular SMILES representation, and generate n molecular conformations;

[0034] Use the MMFF force field optimization to perform initial energy minimization optimization on the molecular conformation to obtain the screened molecular conformation. If the MMFF force field optimization fails, use the UFF force field optimization and calculate the energy corresponding to the optimized molecular conformation;

[0035] Based on the screened molecular conformation, adopt the molecular optimization method of the Python toolkit of ASE+PSI4, or extract the molecular conformation file and perform custom molecular optimization to screen out the molecular conformation with the minimum energy.

[0036] Preferably, the molecular optimization method using the Python toolkit of ASE+PSI4 specifically includes:

[0037] Use the calc interface of ASE to set PSI4 as the calculation engine;

[0038] Set the basis set to "cc-pvdz" and the functional to "B3LYP" for the molecular optimization method;

[0039] Call the optimize method of ASE to perform molecular optimization. During the optimization process, calculate the energy and gradient and update the molecular structure until the convergence condition is met;

[0040] After the optimization is completed, save the optimized molecular structure and calculation results.

[0041] The present invention also provides a molecular de novo design system based on the physical "spectrum-structure-activity" rule. The system includes:

[0042] A data acquisition module for obtaining the SMILES representation of molecules and calculating molecular properties and molecular infrared spectra respectively based on the SMILES representation of molecules;

[0043] A model construction module for training a preset one-dimensional Diffusion model based on molecular properties and molecular infrared spectra to obtain a trained Diffusion model, training a preset Transformer model based on the SMILES representation of molecules and molecular infrared spectra to obtain a trained Transformer model, and integrating the trained Diffusion model and the trained Transformer model to obtain a molecular prediction model;

[0044] A prediction module for inputting the properties of the molecule to be designed into the molecular prediction model to obtain the SMILES representation of the molecule;

[0045] A generation module for screening the SMILES representation of molecules to obtain a reasonable SMILES representation of molecules, generating and optimizing molecular conformations based on the reasonable SMILES representation of molecules to obtain the molecule to be designed.

[0046] The advantages provided by the present invention are as follows:

[0047] (1) In the present invention, the information of several dimensions of the molecular properties required by the user is first input into the Diffusion model to obtain spectral data, mapping from the limited molecular property dimensions to the spectral dimensions, and the dimensions are relatively controllable. Then, the spectral data is input into the Transformer model to obtain the SMILES data of the molecule, and molecules that meet the conditions are generated through the molecular infrared spectrum. This is also a mapping in a relatively low-dimensional space with certain physical and chemical meanings, avoiding the complex mapping directly from one dimension to the three-dimensional structure of the molecule, and decomposing the large dimension crossing problem into two relatively simple dimension conversion problems, thereby reducing the impact of the curse of dimensionality.

[0048] (2) Traditional molecular tasks are to input some data such as molecular structures and predict the corresponding molecular properties. It is very difficult to screen out molecules with specific properties that meet the user's needs in the huge chemical space in this forward prediction manner. The present invention takes the specific molecular properties required by the user as the input, trains a molecular prediction model that can obtain the corresponding molecular structure at one time, and uses this trained knowledge for rapid inference and generation without having to re-explore the entire huge and complex chemical space.

[0049] (3) The present invention provides a de novo molecular design method developed based on the shell language and Python. By establishing a quantitative relationship between molecular properties, spectral data, and the molecular SMILES representation, it provides data-driven guidance for high-performance de novo molecular design, automatically processes and analyzes spectral data according to molecular properties, effectively improves the efficiency and accuracy of spectral data analysis, provides a molecular infrared spectrum generation and molecular SMILES prediction model with multiple trainings and stable effects, reduces the workload of users during the processing, and makes de novo molecular design more accurate and convenient. This learning method is used to guide the design and optimization of molecules, can quickly generate molecules that meet user needs, accelerate the development cycle of molecular materials and drugs, and reduce R & D costs. Compared with traditional methods, the present invention has the characteristics of high efficiency and easy deployment, and has broad application prospects in the fields of chemical synthesis, material research, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a flowchart of the de novo molecular design method based on the physical "spectrum-structure-activity" rule provided by an embodiment of the present invention;

[0051] Figure 2 is an architecture diagram of the preset one-dimensional Diffusion model in the de novo molecular design method based on the physical "spectrum-structure-activity" rule provided by an embodiment of the present invention;

[0052] Figure 3 is a flowchart of the Diffusion model mapping to obtain the molecular infrared spectrum in the de novo molecular design method based on the physical "spectrum-structure-activity" rule provided by an embodiment of the present invention;

[0053] Figure 4 is a flowchart of the molecular infrared spectrum inputting into the trained Transformer model to map and obtain the molecular SMILES representation in the de novo molecular design method based on the physical "spectrum-structure-activity" rule provided by an embodiment of the present invention;

[0054] Figure 5 is a flowchart of the screening of the molecular SMILES representation and the optimization of the conformational stability in the de novo molecular design method based on the physical "spectrum-structure-activity" rule provided by an embodiment of the present invention;

[0055] Figure 6 and Figure 7 is an effect diagram of the molecular spectrum sample generated in the de novo molecular design method based on the physical "spectrum-structure-activity" rule provided by an embodiment of the present invention;

[0056] Figure 8 is a sample diagram of the molecular conformation generated in the de novo molecular design method based on the physical "spectrum-structure-activity" rule provided by an embodiment of the present invention;

[0057] Figure 9 Schematic diagram of a molecular de novo design system based on physical "spectrum-structure-activity" rules provided by an embodiment of the present invention. Specific implementation manners

[0058] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely in conjunction with specific embodiments and with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0059] See Figure 1 , this embodiment provides a molecular de novo design method based on physical "spectrum-structure-activity" rules, including the following steps:

[0060] Step 1: Obtain the SMILES representation of a molecule, and calculate the molecular properties and molecular infrared spectrum respectively according to the SMILES representation of the molecule. Among them, the molecular properties are calculated according to the SMILES representation of the molecule, and the molecular properties include Quantitative Estimate of Drug-likeness (QED), Synthetic Accessibility (SA), and Penalized Logarithm of Partition Coefficient (PLogP).

[0061] The process of calculating the molecular infrared spectrum according to the SMILES representation of the molecule includes:

[0062] S11: Use the RDKIT library to construct multiple molecular three-dimensional conformations for each molecule with the data in the SMILES representation, and then preliminarily optimize the molecular stability with the MMFF force field. After comparison and screening, select the relatively stable conformation among the multiple molecular three-dimensional conformations, obtain the spatial coordinates of each atom in the corresponding molecule, and save them to the molecular conformation gjf file.

[0063] S12: Batch optimize the molecular conformation through the Density Functional Theory (DFT) in the Gaussian software to obtain the optimized molecular conformation, and the energy of the optimized molecular conformation reaches the minimum.

[0064] The specific manner of batch optimizing the molecular conformation is:

[0065] S121. One - key batch - processing script based on the slurm scheduling system, submitting the task queue in one - key, optimizing the molecular conformation through the DFT method. The functional used for optimizing the molecular conformation is B3LYP, and the basis set used is TZVP.

[0066] S122. One - key export and normalization of molecular data based on Gaussian software to obtain the optimized molecular conformation file.

[0067] S13. Continue to use Gaussian software to calculate the intensities of the molecule at some frequencies based on the optimized molecular conformation, forming a discrete molecular infrared spectrum, and saving the discrete molecular infrared spectrum in the output log file of Gaussian software.

[0068] Step S13 specifically includes:

[0069] S131. One - key batch - processing script based on the slurm scheduling system, submitting the task queue in one - key, calculating the intensities of the molecule at some frequencies. The functional used is B3LYP, and the basis set used is TZVP to obtain a discrete molecular infrared spectrum;

[0070] S132. One - key export and normalization of spectral data based on Gaussian software to obtain the molecular infrared spectrum file.

[0071] S14. Extract the corresponding frequencies and intensities of the molecule from the Gaussian software output log file, expand the frequencies of the discrete molecular infrared spectrum to 0 - 4000 cm -1 , set the broadening of the molecular infrared spectrum to 10, normalize the intensities, and use a python script to extract the one - dimensional molecular infrared spectrum data from the Gaussian software output file and save it to a csv table.

[0072] Step 2. Train the preset one - dimensional Diffusion model based on molecular properties and the molecular infrared spectrum to obtain the trained Diffusion model, train the preset Transformer model based on the molecular smiles representation and the molecular infrared spectrum to obtain the trained Transformer model, and integrate the trained Diffusion model and the trained Transformer model to obtain the molecular prediction model.

[0073] The molecular properties, the molecular smiles representation, and the molecular infrared spectrum are jointly saved in the csv data table. The process of training the preset Diffusion model based on molecular properties and the molecular infrared spectrum includes:

[0074] S211. Refer to Figure 2, the feature extraction layer of the existing Diffusion model is changed to a one-dimensional convolution (kernel_size = 5) to capture local temporal features. Similarly, all Conv2D layers in Unet are replaced with Conv1D layers to replace all two-dimensional convolution operations. In downsampling, the max pooling method is changed from MaxPool2D to MaxPool1D (kernel_size = 2, stride = 2); in upsampling, the transposed convolution ConvTranspose2D is changed to ConvTranspose1D. First, the length is extended by linear interpolation through Upsample(mode = 'linear'), and then the features are refined with Conv1D(kernel_size = 3). At the same time, a fully connected layer and a self-attention mechanism are used in the skip connection part for feature extraction to obtain the preset one-dimensional Diffusion model.

[0075] S212. Input the molecular properties and the molecular infrared spectrum into the preset one-dimensional Diffusion model, perform epoch iterations of training, and update the parameters of the preset one-dimensional Diffusion model through the backpropagation algorithm to obtain the trained Diffusion model;

[0076] During the training process, adjust the hyperparameters of the Diffusion model according to the performance during the training process to optimize the Diffusion model, specifically including: defining the optimizer, loss function, and learning rate decay strategy, and setting the hyperparameters; using the mean absolute error MAE as the evaluation metric of the model, setting the dimension and size of Unet in the Diffusion model; adopting the Adam optimization algorithm to adaptively adjust the learning rates of the hyperparameters of the model. Set to perform inference generation every 10 epochs, generate a one-dimensional spectrum, use the frequency of the one-dimensional spectrum as the abscissa and the intensity as the ordinate, see Figure 6 and Figure 7 , and generate a two-dimensional spectrum diagram to view the model effect.

[0077] The process of training the preset Transformer model based on the molecular smiles representation and the molecular infrared spectrum includes:

[0078] S221. Add a one-dimensional convolutional feature extraction layer in front of the preset Transformer model to extract the feature information in the input molecular infrared spectrum.

[0079] S222. Set up a dictionary mapping. Input the molecular SMILES representation and the molecular infrared spectrum into a pre-set Transformer model. Use the molecular SMILES representation as the label and perform iterative training for epoch times. Update the parameters of the Transformer model through the backpropagation algorithm to obtain the trained Transformer model. During the training process, adjust the hyperparameters of the model according to the performance during the training process to optimize the Transformer model. Define the optimizer, loss function, and learning rate decay strategy, and set the hyperparameters. Use the cross-entropy loss function and the SMILES complete prediction accuracy as the evaluation metrics of the model, and set the input and output parameter sizes of the feature extraction layer, Encoder, and Decoder in the Transformer model. The Transformer model uses the Adam optimization algorithm to adaptively adjust the learning rates of various hyperparameters. When validating the trained Transformer model, set early stopping. If the loss of the validation set is greater than the minimum loss for 10 consecutive epochs, stop the training and save the trained model.

[0080] Among them, the method of setting up the dictionary mapping includes:

[0081] In the model data input, for the molecular SMILES string in the label, through dictionary mapping, map each character into a number, and the mappings of all characters in a molecular SMILES form a one-dimensional vector and input it into the Transformer model.

[0082] In the model prediction output, calculate the probability of each position in the vector output by the model through softmax to form a character prediction probability branch tree. Take the top-n SMILES and convert the autoregressively generated one-dimensional numerical vector into a molecular SMILES representation through dictionary mapping.

[0083] Use the output of the trained Diffusion model as the input of the trained Transformer model, load the saved path of the trained model parameter file, and integrate the trained Diffusion model and the trained Transformer model to obtain a molecular prediction model.

[0084] Step 3. Input the properties of the molecule to be designed into the molecular prediction model to obtain the molecular SMILES representation; see Figure 3 and Figure 4, for the property QED, PLogP, and SA input molecular prediction models of the molecule to be designed, after being trained first, the Diffusion model maps to obtain the molecular infrared spectrum. The molecular infrared spectrum is input into the trained Transformer model to map to obtain the molecular smiles representation, and this molecular smiles representation conforms to the molecular properties required by the user.

[0085] In the present invention, the information of several dimensions of the molecular properties required by the user is first input into the Diffusion model to obtain spectral data. This step maps from the limited molecular property dimensions to the spectral dimension and is relatively controllable. Then, the spectral data is input into the Transformer model to obtain the molecular smiles data, and molecules that meet the conditions are generated through the molecular infrared spectrum. This is also a mapping in a relatively low-dimensional space with certain physical and chemical meanings, avoiding the complex mapping directly from one dimension to the three-dimensional molecular structure, and decomposing the large dimension crossing problem into two relatively simple dimension conversion problems, thereby reducing the impact of the curse of dimensionality.

[0086] Traditional molecular tasks are to input some data such as molecular structures and predict the corresponding molecular properties. It is very difficult to screen out molecules with specific properties that meet the user's needs in the huge chemical space in this forward prediction manner. The present invention takes the specific molecular properties required by the user as the input, trains a molecular prediction model that can obtain the corresponding molecular structure at one time, and uses this trained knowledge for rapid inference and generation without having to re-explore the entire huge and complex chemical space.

[0087] Step 4: Screen the molecular smiles representation to obtain a reasonable molecular smiles representation, generate and optimize the molecular conformation based on the reasonable molecular smiles representation to obtain the molecule to be designed.

[0088] Among them, the process of screening the molecular smiles representation to obtain a reasonable molecular smiles representation includes:

[0089] First, screen out the molecular smiles representations that conform to the grammar from the molecular smiles representations generated by the molecular prediction model according to the grammar of the molecular smiles representation, and then screen out the reasonable molecular smiles representations from the molecular smiles representations that conform to the grammar through the python toolkit of RDKIT. The reasonable molecular smiles representations conform to the three property requirements (QED, PLogP, SA) input by the prediction model.

[0090] The process of generating and optimizing the molecular conformation based on the reasonable molecular smiles representation includes:

[0091] S41. Use the Python toolkit of RDKIT, input the reasonable molecular SMILES representation, and generate n molecular conformations.

[0092] S42. Use the MMFF force field optimization to perform initial energy minimization optimization on the molecular conformations. If the MMFF force field optimization fails, use the UFF force field optimization and calculate the energy corresponding to the optimized molecular conformations.

[0093] S43. Screen out the conformation with the lowest energy among these molecular conformations for further molecular optimization in the following. Based on the screened molecular conformations, adopt the molecular optimization method of the Python toolkit of ASE+PSI4, or extract the molecular conformation file for custom molecular optimization, and screen out the molecular conformation with the lowest energy. The molecular optimization method of the Python toolkit of ASE+PSI4 specifically includes:

[0094] S431. Use the calc interface of ASE to set PSI4 as the calculation engine.

[0095] S432. During the optimization process, set the basis set to "cc-pvdz" and the functional to "B3LYP" for the molecular optimization method.

[0096] S433. Perform molecular optimization by calling the optimize method of ASE. During the optimization process, calculate the energy and gradient and update the molecular structure until the convergence condition is met.

[0097] S434. After the optimization is completed, save the optimized molecular structure and calculation results, see Figure 8 .

[0098] The present invention provides a molecular de novo design method based on the physical "spectrum-structure-activity" rule developed based on the shell language and Python. By establishing the quantitative relationship between molecular properties, spectral data, and molecular SMILES representation, it provides data-driven guidance for high-performance molecular de novo design, automatically processes and analyzes spectral data according to molecular properties, effectively improves the efficiency and accuracy of spectral data analysis, provides a molecular infrared spectrum generation and molecular SMILES prediction model with multiple trainings and stable effects, reduces the workload of users in the processing process, and makes molecular de novo design more accurate and convenient. This learning method is used to guide the design and optimization of molecules, can quickly generate and meet the molecules required by users, accelerate the development cycle of molecular materials and drugs, and reduce the R & D cost.

[0099] The machine learning model of the present invention has undergone extensive verification and testing, and the results show that it can generate molecules with specific properties according to requirements. Compared with traditional methods, the present invention has the characteristics of high efficiency and easy deployment. This method has broad application prospects in the fields of chemical synthesis, material research, etc.

[0100] See Figure 9 , the present invention also provides a de novo molecular design system based on physical "spectrum-structure-activity" rules, which can be used to execute the method of the present invention. For the details not disclosed in the system of the present invention, please refer to the method of the present invention and will not be elaborated here. The system includes:

[0101] A data acquisition module, which is used to obtain the SMILES representation of molecules, and calculate the molecular properties and molecular infrared spectra respectively according to the SMILES representation of molecules;

[0102] A model construction module, which is used to train a preset Diffusion model based on the molecular properties and molecular infrared spectra to obtain a trained Diffusion model, train a preset Transformer model based on the SMILES representation of molecules and molecular infrared spectra to obtain a trained Transformer model, and integrate the trained Diffusion model and the trained Transformer model to obtain a molecular prediction model;

[0103] A prediction module, which is used to input the properties of the molecule to be designed into the molecular prediction model to obtain the SMILES representation of the molecule;

[0104] A generation module, which is used to screen the SMILES representation of molecules to obtain a reasonable SMILES representation, generate and optimize the molecular conformation based on the reasonable SMILES representation to obtain the molecule to be designed.

[0105] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A molecular de novo design method based on the physical "spectrum structure-activity" rule, characterized by: Methods include: Get the molecular smiles representation, and calculate the molecular properties and molecular infrared spectrum according to the molecular smiles representation; The preset one-dimensional Diffusion model is trained based on the molecular properties and the molecular infrared spectrum to obtain the trained Diffusion model. The preset Transformer model is trained based on the molecular smiles representation and the molecular infrared spectrum to obtain the trained Transformer model. The trained Diffusion model and the trained Transformer model are integrated to obtain the molecular prediction model. Input the properties of the molecule to be designed into the molecular prediction model to obtain the molecular smiles representation; The molecular smiles representation is screened to obtain a reasonable molecular smiles representation, and the molecular conformation is generated and optimized based on the reasonable molecular smiles representation to obtain the molecule to be designed.

2. The molecular de novo design method based on physical "spectrum structure-activity" rules according to claim 1, characterized in that: The process of calculating the infrared spectrum of a molecule according to the molecular smiles representation includes: Through the RDKIT library, multiple molecular three-dimensional conformations are constructed for each molecule using molecular smiles, and the molecular stability is preliminarily optimized using the MMFF force field. A stable molecular conformation is selected from multiple molecular three-dimensional conformations to obtain the spatial coordinates of each atom in the molecule and save them in a molecular conformation file; Batch optimize molecular conformations through DFT method to obtain optimized molecular conformations; The intensity of the molecule at multiple frequencies is calculated according to the optimized molecular conformation to form a discrete molecular infrared spectrum, which is then saved in the output log file of the Gaussian software. Extract the corresponding frequency and intensity of the molecule from the output log file, and expand the frequency of the discrete molecular infrared spectrum to 0-4000cm by linear interpolation. -1 , set the molecular infrared spectrum broadening to 10, normalize the intensity, and extract the molecular infrared spectrum data.

3. The molecular de novo design method based on physical "spectrum structure-activity" rules according to claim 2 is characterized by: The method of batch optimization of molecular conformation is: One-click batch processing script based on slurm scheduling system, one-click submission of task queue, optimization of molecular conformation by DFT method, the functional used for optimization of molecular conformation is B3LYP, and the basis set used is TZVP; The molecular data based on Gaussian software is exported and normalized with one click to obtain the optimized molecular conformation file.

4. The molecular de novo design method based on physical "spectrum structure-activity" rules according to claim 1, characterized in that: The process of training the preset one-dimensional Diffusion model based on molecular properties and molecular infrared spectra includes: The feature extraction layer of the existing Diffusion model is changed to a one-dimensional convolution to capture local temporal features. All Conv2D layers in the existing Diffusion model Unet are replaced with Conv1D layers. In downsampling, the maximum pooling method is changed from MaxPool2D to MaxPool1D. In upsampling, the transposed convolution ConvTranspose2D is changed to ConvTranspose1D. The length is first extended by linear interpolation through Upsample, and then Conv1D is used to refine the features. In the skip link part, the fully connected layer and the self-attention mechanism are used for feature extraction to obtain the preset one-dimensional Diffusion model. The molecular properties and molecular infrared spectra are input into the preset one-dimensional Diffusion model, and epochs of iterative training are performed. The parameters of the preset one-dimensional Diffusion model are updated through the back propagation algorithm to obtain the trained Diffusion model.

5. The molecular de novo design method based on physical "spectrum structure-activity" rules according to claim 1, characterized in that: The process of training the preset Transformer model based on molecular smiles representation and molecular infrared spectrum includes: A one-dimensional convolution feature extraction layer is added before the preset Transformer model to extract feature information from the input molecular infrared spectrum; Set up dictionary mapping, input the molecular smile representation and molecular infrared spectrum into the preset Transformer model, use the molecular smile representation as the label, perform epoch iterative training, update the parameters of the Transformer model through the back-propagation algorithm, and obtain the trained Transformer model.

6. The molecular de novo design method based on physical "spectrum structure-activity" rules according to claim 5, characterized in that: The process of setting up dictionary mappings involves: In the model data input, for the molecular smiles string in the label, each character is mapped into a number through dictionary mapping, and all the character mappings in a molecular smile form a one-dimensional vector and are input into the Transformer model; In the model prediction output, the probability of each position in the vector output by the model is calculated through softmax to form a character prediction probability branch tree, take the top-n smiles, and convert the one-dimensional numerical vector generated by autoregression into a molecular smile representation through dictionary mapping.

7. The molecular de novo design method based on physical "spectrum structure-activity" rules according to claim 1, characterized in that: The process of screening the molecular smiles representation and obtaining a reasonable molecular smiles representation includes: According to the syntax of the molecular smiles representation, the molecular smiles representation that conforms to the syntax is selected from the molecular smiles representation generated by the molecular prediction model; Through the RDKIT python toolkit, reasonable molecular smile representations are screened out from the grammatical molecular smile representations. Reasonable molecular smile representations meet the property requirements QED, PLogP, and SA of the prediction model input.

8. The method for de novo molecular design based on physical "spectrum-structure-activity" rules according to claim 1, characterized in that: The process of generating and optimizing molecular conformations based on reasonable molecular smiles representations includes: Use the RDKIT python toolkit to input a reasonable molecular smiles representation and generate n molecular conformations; Use MMFF force field optimization to perform initial energy minimization optimization on the molecular conformation to obtain the screened molecular conformation. If the MMFF force field optimization fails, use UFF force field optimization to calculate the energy corresponding to the optimized molecular conformation. Based on the screened molecular conformations, the molecular optimization method of the python toolkit ASE+PSI4 is used, or the molecular conformation file is taken out for custom molecular optimization to screen out the molecular conformation with the minimum energy.

9. The method for de novo molecular design based on physical "spectrum structure-activity" rules according to claim 8, characterized in that: The molecular optimization method using the python toolkit ASE+PSI4 specifically includes: Use ASE's calc interface to set PSI4 as the calculation engine; Set the molecular optimization method with basis set to "cc-pvdz" and functional to "B3LYP"; Call the optimize method of ASE to perform molecular optimization. During the optimization process, the energy and gradient are calculated and the molecular structure is updated until the convergence conditions are met. After the optimization is completed, save the optimized molecular structure and calculation results.

10. A molecular de novo design system based on physical "spectrum structure-activity" rules, characterized by: The system includes: A data acquisition module is used to obtain the molecular smiles representation, and calculate the molecular properties and molecular infrared spectra according to the molecular smiles representation; A model building module is used to train a preset one-dimensional Diffusion model based on molecular properties and molecular infrared spectra to obtain a trained Diffusion model, train a preset Transformer model based on molecular smiles representation and molecular infrared spectra to obtain a trained Transformer model, and integrate the trained Diffusion model and the trained Transformer model to obtain a molecular prediction model; The prediction module is used to input the properties of the molecules to be designed into the molecular prediction model to obtain the molecular smiles representation; The generation module is used to screen the molecular smiles representation to obtain a reasonable molecular smiles representation, generate and optimize the molecular conformation based on the reasonable molecular smiles representation, and obtain the molecule to be designed.