Efficient screening method for chiral synthesis catalyst based on automatic machine learning
Catalyst screening is solved through the method based on automatic machine learning, and the problems of long cycles and high costs in traditional methods are realized, efficient and automated catalyst screening is achieved, cost and time are reduced, and screening accuracy is improved.
Patent Information
- Application Number
- CN202311004332.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-10
- Publication Date
- 2025-07-11
AI Technical Summary
The traditional chiral catalyst screening method has a long period, high cost and low flux, making it difficult to efficiently screen a large number of catalysts within a reasonable time.
Using the chiral synthetic catalyst screening method based on automatic machine learning, molecular feature vectors and model construction are carried out by collecting known catalyst data, and model construction is carried out using the AutoGluon toolkit for model construction and prediction screening of unknown catalysts.
Efficient and automated catalyst screening is achieved, reducing manual errors, reducing screening costs and time, and improving screening accuracy and throughput.
Smart Images

Figure CN120299550A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of application of artificial intelligence technology in chemoinformatics, and particularly relates to providing an efficient screening method for chiral synthesis catalysts based on automated machine learning. Background Art
[0002] With the rapid development of modern pharmaceutical science, the synthesis of small molecule drugs plays a crucial role in treating diseases and improving the quality of life of patients. Many drug molecules are chiral molecules, which have non-superimposable mirror isomers. The chirality difference of chiral drugs can have a significant impact on their interaction with biomolecules. Sometimes, only one chiral isomer is an effective drug, while the other chiral isomer may have no therapeutic effect at all or even cause toxic reactions. A classic example is the drug "Thalidomide", where one chiral isomer caused severe fetal developmental defects, while the other chiral isomer is widely used to treat various diseases. Therefore, the efficient control of the synthesis and preparation process of chiral drugs has become an important challenge in the field of drug synthesis today.
[0003] Chiral catalysts are a class of catalysts with high stereoselectivity, which can guide the formation of products with specific chirality in chemical reactions. Compared with traditional chiral synthesis methods, chiral catalysts can achieve higher chiral selectivity and higher yields, thus greatly improving the preparation efficiency of chiral drugs. The application of chiral catalysts in the synthesis of small molecule drugs has made remarkable progress in recent years and has been widely used in various drug synthesis reactions.
[0004] However, screening traditional experimental chiral catalysts not only has a long cycle and high cost, but also it is very difficult to efficiently screen a large number of chiral catalysts within a reasonable time. Therefore, it is necessary to design a practical data-driven new method for catalyst design and development to achieve the efficient screening of chiral synthesis catalysts for small molecule drugs, so as to reduce the development cost and development cycle. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide an efficient screening method for chiral synthesis catalysts based on automated machine learning, so as to solve the disadvantages of the traditional experimental screening method in the prior art, which not only has a long cycle, high cost, but also has a low screening throughput. The present invention uses a small amount of known catalyst data for accurate modeling, so as to achieve the efficient screening of unknown catalysts, thereby accelerating the progress of drug research and development and reducing the research and development cost.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] An efficient screening method for chiral synthesis catalysts based on automated machine learning, comprising the following steps:
[0008] S1: Collection and collation of catalyst molecular data:
[0009] Collect the molecular data of existing known catalysts, where the molecular data includes the chemical formula of the catalyst, the yield during synthesis of the catalyst, and the proportion of the dominant enantiomer, and perform normalization processing and collation on the molecular data; and collect and collate the chemical formula of the new catalyst molecules to be predicted.
[0010] S2: Vectorization of catalyst molecular features:
[0011] Calculate various molecular descriptors of the known catalysts and the new catalyst molecules to be predicted according to the chemical formula of the molecules, and splice the numerical values of the molecular descriptors into a one-dimensional feature vector; and organize the one-dimensional feature vectors of the molecules of the known catalysts and the data of their actual yields and the proportions of the dominant enantiomers into a table for subsequent machine learning modeling.
[0012] S3: Construction of an automated machine learning model:
[0013] Construct a model through an automated machine learning toolkit and save the model parameters as a file.
[0014] S4: Screening and verification of unknown catalyst molecules:
[0015] Use the constructed model to predict the new catalyst molecules to be predicted, and select the molecules with higher prediction scores for further experimental verification.
[0016] Furthermore, in step S1, it specifically includes: collecting the SMILES of the known catalysts in the public literature and the new catalyst molecules in the preliminary pre-experiments; collecting the numerical values of the yields and the proportions of the dominant enantiomers of the known catalysts, and normalizing these numerical values.
[0017] Further, the specific steps in step S2 include: using the rdkit toolkit to calculate the molecular descriptors of all the known catalysts and the new catalysts to be predicted. The molecular descriptors include molecular weight, number of radical electrons, and number of valence electrons, specifically including NumRadicalElectrons, SMR_VSA8, SlogP_VSA9, fr_azide, MolWt, HeavyAtomMolWt, ExactMolWt, NumValenceElectrons, fr_diazo, fr_isocyan, fr_isothiocyan, fr_nitroso, BertzCT, LabuteASA, MolMR, fr_prisulfonamd, fr_thiocyan, fr_aldehyde, fr_azo, fr_benzodiazepine, fr_dihydropyridine, fr_epoxide, fr_hdrzone, fr_quatN, fr_C_S, fr_SH, fr_alkyl_carbamate, fr_barbitur, fr_phos_acid, fr_phos_ester, fr_term_acetylene, fr_Imine, fr_amidine, fr_guanido, fr_hdrzine, fr_imide, fr_lactam, fr_lactone, fr_oxime, EState_VSA11, fr_N_O, fr_nitro, fr_nitro_arom, fr_nitro_arom_nonortho, MaxPartialCharge, MinPartialCharge, MaxAbsPartialCharge, MinAbsPartialCharge, Ipc, BCUT2D_MWHI, BCUT2D_MWLOW, BCUT2D_CHGHI, BCUT2D_CHGLO, BCUT2D_LOGPHI, BCUT2D_LOGPLOW, BCUT2D_MRHI, BCUT2D_MRLOW, and splicing the numerical values of the molecular descriptors into a one-dimensional feature vector; combining the feature vectors, actual yields, and enantiomeric excess ratio data of the known catalyst molecules and storing them in a csv file, where the one-dimensional feature vector is used as the input feature for model prediction, and the actual yield and enantiomeric excess ratio data are used as the output results of model prediction.
[0018] Further, in the step S3, it includes: using the AutoGluon automatic machine learning toolkit to construct the model with the sorted molecular data of the known catalyst, and storing the constructed model in the form of a file.
[0019] Further, in the step S4, it includes: using the model constructed in the step S3 to predict the unknown catalyst molecules, sorting them in descending order according to the predicted values, and selecting several molecules with higher rankings for further experimental verification.
[0020] Further, in the step S1, the known catalyst is the existing chiral catalyst in the synthesis of bedaquiline collected from the open literature, and the number of the collected known catalysts is 27; the new catalyst is the SMILES of the unknown catalyst molecules to be predicted collected from the open database, and the number of the new catalysts is 507.
[0021] Further, in the step S2, the values of the molecular descriptors are 8.50, 0.31, 8.50, 0.31, 0.48, 1.71, 2.57, 2.86, 2.14, 5.11, 4.35, 4.35, 3.43, 2.77, 2.77, 1.98, 1.98, 1.38, 1.38, 0.91, 0.91, -0.08, 5.06, 2.28, 1.14, 10.42, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 19.39, 6.04, 6.61, 5.11, 0, 0, 5.32, 0, 18.88, 13.15, 0, 0, 5.32, 0, 0, 0, 24.30, 0, 0, 12.84, 0, 0, 0, 32.26, 0, 0, 0, 12.65, 12.97, 6.42, 0, 0, 5.32, 5.11, 0, 0, 0, 11.66, 0, 0, 0.40, 2.38, 1.39, 0, 1, 7, 2, 2, 0, 1, 1, 0, 0, 0, 2, 2, 2, 1, 0, 1, 1, 1, -0.27, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0.24, 1, 1, 1, 0, 0, 1.38, 2.13, 0, 0, 0, 0, 0, 0, 0, 0, 0.89, 13.06, 1.89, 0.01, 12.41, 0, 0, 0, 0, 18.86, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, -5.99, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 061,5.11,0,0,5.32,0,18.88,13.15,0,0,5.32,0,0,0,24.30,0,0,12.84,0,0,0,32.26,0,0,0,12.65,12.97,6.42,0,0,5.32,5.11,0,0,0,11.66,0,0,0.40,2.38,1.39,0,1,7,2,2,0,1,1,0,0,0,2,2,2,1,0,1,1,1,-0.27,0,1,1,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.24,1,1,1,0,0,1.38,2.13,0,0,0,0,0,0,0,0,0,0.89,13.06,1.89,0.01,12.41,0,0,0,0,18.86,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,2,0,0,0,0,0,0,0,0,0,0,2,0,0,0,0,0,-5.99,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0}, The vector length of the one-dimensional feature vector is 302.
[0022] Further, in step S4, 507 of the new catalyst molecules are predicted and sorted in descending order according to the predicted values.
[0023] The advantages of the method of the present invention are as follows: It can use a small amount of data of known catalysts to perform machine learning modeling automatically and with high precision, so as to screen a large number of unknown catalyst molecules, thereby reducing the cycle and cost of catalyst screening.
[0024] The automated machine learning algorithm adopted in the method of the present invention can effectively avoid human errors, reduce cumbersome machine learning tasks, and complete high-precision modeling with only a small amount of data. The algorithm has a simple process and a wide range of applications. It can not only be used in the screening of chiral catalysts, but also in the screening of other ordinary catalysts, and even extended to the prediction of the properties of small drug molecules and other aspects.
[0025] Other advantages, objects and features of the present invention will be set forth to some extent in the following description, and to some extent, will be apparent to those skilled in the art based on the study of the following, or can be taught from the practice of the present invention. The objects and other advantages of the present invention can be achieved and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic flow chart of the present invention;
[0027] Figure 2 is a schematic diagram of the screening results of catalyst molecules for the chiral synthesis of bedaquiline as an example in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art fall within the scope defined by the appended claims of this application.
[0029] As Figure 1 shown, the present application provides an efficient screening method for chiral synthesis catalysts based on automated machine learning. Taking the catalyst for the chiral synthesis of bedaquiline, which is currently difficult to synthesize, has a low yield, and a low proportion of the dominant enantiomer, as an example, only 27 known molecules are used for automated machine learning modeling, and then 507 potential unknown catalysts are screened using this model. The molecules ranked among the top are shown in Figure 2 for subsequent experimental verification.
[0030] Specifically, it includes the following steps:
[0031] S1: Collection and arrangement of catalyst molecule data:
[0032] Collect the SMILES of known catalysts in the open literature and new catalysts in the preliminary pre-experiments; for subsequent molecular feature vectorization. Collect the yield of the known catalysts and the proportion values of the dominant enantiomers, and normalize these values for subsequent modeling.
[0033] The known catalysts are chiral catalysts in the existing bedaquiline synthesis collected from the open literature, and the number of the known catalysts collected is 27; the new catalysts are the SMILES of unknown catalyst molecules to be predicted collected from the open database, and the number of the new catalysts is 507.
[0034] S2: Catalyst molecule feature vectorization:
[0035] The molecular descriptors of all the said known catalysts and the said new catalysts to be predicted are calculated using the RDKit toolkit. The molecular descriptors include molecular weight, number of radical electrons, number of valence electrons, etc., specifically including NumRadicalElectrons, SMR_VSA8, SlogP_VSA9, fr_azide, MolWt, HeavyAtomMolWt, ExactMolWt, NumValenceElectrons, fr_diazo, fr_isocyan, fr_isothiocyan, fr_nitroso, BertzCT, LabuteASA, MolMR, fr_prisulfonamd, fr_thiocyan, fr_aldehyde, fr_azo, fr_benzodiazepine, fr_dihydropyridine, fr_epoxide, fr_hdrzone, fr_quatN, fr_C_S, fr_SH, fr_alkyl_carbamate, fr_barbitur, fr_phos_acid, fr_phos_ester, fr_term_acetylene, fr_Imine, fr_amidine, fr_guanido, fr_hdrzine, fr_imide, fr_lactam, fr_lactone, fr_oxime, EState_VSA11, fr_N_O, fr_nitro, fr_nitro_arom, fr_nitro_arom_nonortho, MaxPartialCharge, MinPartialCharge, MaxAbsPartialCharge, MinAbsPartialCharge, Ipc, BCUT2D_MWHI, BCUT2D_MWLOW, BCUT2D_CHGHI, BCUT2D_CHGLO, BCUT2D_LOGPHI, BCUT2D_LOGPLOW, BCUT2D_MRHI, BCUT2D_MRLOW, etc. The values of the molecular descriptors are 8.50, 0.31, 8.50, 0.31, 0.48, 1.71, 2.57, 2.86, 2.14, 5.11, 4.35, 4.35, 3.43, 2.77, 2.77, 1.98, 1.98, 1.38, 1.38, 0.91, 0.91, -0.08, 5.06, 2.28, 1.14, 10.42, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 19.39, 6.04, 6.61, 5.11, 0, 0, 5.32, 0, 18.88, 13.15, 0, 0, 5.32, 0, 0, 0, 24.30, 0, 0, 12.84, 0, 0, 0, 32.26, 0, 0, 0, 12.65, 12.97, 6.42, 0, 0, 5.32, 5.11, 0, 0, 0, 11.66, 0, 0, 0.40, 2.38, 1.39, 0, 1, 7, 2, 2, 0, 1, 1, 0, 0, 0, 2, 2, 2, 1, 0, 1, 1, 1, -0.27, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0.24, 1, 1, 1, 0, 0, 1.38, 2.13, 0, 0, 0, 0, 0, 0, 0, 0, 0.89, 13.06, 1.89, 0.01, 12.41, 0, 0, 0, 0, 18.86, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, -5.99, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, and concatenate the numerical values of the molecular descriptors into a one-dimensional feature vector {{8.50, 0.31, 8.50, 0.31, 0.48, 1.71, 2.57, 2.86, 2.14, 5.11, 4.35, 4.35, 3.43, 2.77, 2.77, 1.98, 1.98, 1.38, 1.38, 0.91, 0.91, -0.08, 5.06, 2.28, 1.14, 10.42, 0, 0, 0, 0, 0, 0, 0, 0, 19.39, 6.04, 6.61, 5.11, 0, 0, 5.32, 0, 18.88, 13.15, 0, 0, 5.32, 0, 0, 0, 24.30, 0, 0, 12.84, 0, 0, 0, 32.26, 0, 0, 0, 12.65, 12.97, 6.42, 0, 0, 5.32, 5.11, 0, 0, 0, 11.66, 0, 0, 0.40, 2.38, 1.39, 0, 1, 7, 2, 2, 0, 1, 1, 0, 0, 0, 2, 2, 2, 1, 0, 1, 1, 1, -0.{27,0,1,1,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.24,1,1,1,0,0,1.38,2.13,0,0,0,0,0,0,0,0,0,0.89,13.06,1.89,0.01,12.41,0,0,0,0,18.86,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,2,0,0,0,0,0,0,0,0,0,0,2,0,0,0,0,0,-5.99,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0}, The vector length of the one-dimensional feature vector is 302. The feature vector, actual yield, and enantiomeric excess ratio data of the known catalyst molecules are combined and stored in a csv file, where the one-dimensional feature vector is used as the input feature for model prediction, and the actual yield and enantiomeric excess ratio data are used as the output results of model prediction.
[0036] S3: Construction of the automated machine learning model:
[0037] Using the AutoGluon automated machine learning toolkit, construct the model with the organized molecular data of the known catalyst, and store the constructed model in the form of a file.
[0038] S4: Screening and verification of unknown catalyst molecules:
[0039] Using the model constructed in step S3, predict 507 predicted unknown catalyst molecules, sort them in descending order according to the predicted values, and select several molecules with higher rankings for further experimental verification. Select several molecules with higher rankings for further experimental verification.
[0040] The method of the present invention can effectively reduce the data for screening catalysts in experiments, without the need to screen all 507 unknown catalysts, significantly reducing the screening cost; at the same time, this method can further utilize the experimental values measured subsequently to further improve the accuracy of the model, thereby increasing the screening success rate.
[0041] Compared with traditional experimental screening methods, the method designed in this patent has the following advantages:
[0042] 1. This method uses an automated machine learning algorithm, effectively reducing human errors and providing modeling efficiency and accuracy;
[0043] 2. This method only requires a small amount of experimental data to build a model, reducing the cost of preliminary experiments;
[0044] 3. This method can quickly predict a large number of brand-new catalysts, greatly shortening the time for screening catalysts.
Claims
1. An efficient screening method for chiral synthesis catalysts based on automated machine learning, characterized in that, It includes the following steps: S1: Collection and arrangement of catalyst molecular data: Collect the molecular data of existing known catalysts, where the molecular data includes the chemical formula of the catalyst, the yield during synthesis of the catalyst, and the proportion of the dominant enantiomer, and perform normalization processing and arrangement on the molecular data; and collect and arrange the chemical formula of the new catalyst molecule to be predicted. S2: Vectorization of catalyst molecular features: Calculate various molecular descriptors of the known catalyst and the new catalyst molecule to be predicted according to the chemical formula of the molecule, and splice the numerical values of the molecular descriptors into a one-dimensional feature vector; and organize the one-dimensional feature vectors of the molecules of the known catalyst, as well as the data of their actual yields and the proportions of the dominant enantiomers, into a table for subsequent machine learning modeling. S3: Construction of an automated machine learning model: Construct a model through an automated machine learning toolkit and save the model parameters as a file. S4: Screening and verification of unknown catalyst molecules: Use the constructed model to predict the new catalyst molecules to be predicted, and select the molecules with higher prediction scores for further experimental verification.
2. The high-throughput screening method of chiral synthesis catalysts based on automated machine learning according to claim 1, wherein, Specifically included in step S1 are: collecting the SMILES of the molecules of the known catalysts in the public literature and the new catalysts in the preliminary pre-experiments; collecting the numerical values of the yields and the proportions of the dominant enantiomers of the known catalysts, and normalizing these numerical values.
3. The high-efficiency screening method of chiral synthesis catalysts based on automated machine learning according to claim 1, characterized in that, Specifically, the step S2 includes: calculating the molecular descriptors of all the known catalysts and the new catalysts to be predicted by using the rdkit toolkit. The molecular descriptors specifically include NumRadicalElectrons, SMR_VSA8, SlogP_VSA9, fr_azide, MolWt, HeavyAtomMolWt, ExactMolWt, NumValenceElectrons, fr_diazo, fr_isocyan, fr_isothiocyan, fr_nitroso, BertzCT, LabuteASA, MolMR, fr_prisulfonamd, fr_thiocyan, fr_aldehyde, fr_azo, fr_benzodiazepine, fr_dihydropyridine, fr_epoxide, fr_hdrzone, fr_quatN, fr_C_S, fr_SH, fr_alkyl_carbamate, fr_barbitur, fr_phos_acid, fr_phos_ester, fr_term_acetylene, fr_Imine, fr_amidine, fr_guanido, fr_hdrzine, fr_imide, fr_lactam, fr_lactone, fr_oxime, EState_VSA11, fr_N_O, fr_nitro, fr_nitro_arom, fr_nitro_arom_nonortho, MaxPartialCharge, MinPartialCharge, MaxAbsPartialCharge, MinAbsPartialCharge, Ipc, BCUT2D_MWHI, BCUT2D_MWLOW, BCUT2D_CHGHI, BCUT2D_CHGLO, BCUT2D_LOGPHI, BCUT2D_LOGPLOW, BCUT2D_MRHI, BCUT2D_MRLOW, etc., and splicing the numerical values of the molecular descriptors into a one-dimensional feature vector; combining the feature vectors, actual yield and enantiomeric excess ratio data of the known catalyst molecules and storing them in a csv file. The one-dimensional feature vector is used as the input feature for model prediction, and the actual yield and enantiomeric excess ratio data are used as the output results of model prediction.
4. The high-throughput screening method of chiral synthesis catalysts based on automated machine learning according to claim 1, characterized in that The step S3 includes: using the AutoGluon automated machine learning toolkit to construct the model with the sorted molecular data of the known catalysts, and storing the constructed model in the form of a file.
5. The high - efficiency screening method of chiral synthesis catalysts based on automated machine learning according to claim 1, characterized in that, Step S4 includes: using the model constructed in step S3 to predict the predicted unknown catalyst molecules, sorting them in descending order according to the predicted values, and selecting several molecules with higher rankings for further experimental verification. Select several molecules with higher rankings for further experimental verification.
6. The efficient screening method of chiral synthesis catalysts based on automated machine learning according to claim 2, characterized in that In step S1, the known catalysts are chiral catalysts in the existing bedaquiline synthesis collected from the open literature, and the number of the known catalysts collected is 27; the new catalysts are SMILES of unknown catalyst molecules to be predicted collected from the open database, and the number of the new catalysts is 507.
7. The efficient screening method of chiral synthesis catalysts based on automated machine learning according to claim 6, characterized in that, In the said step S2, the numerical values of the molecular descriptors are 8.50, 0.31, 8.50, 0.31, 0.48, 1.71, 2.57, 2.86, 2.14, 5.11, 4.35, 4.35, 3.43, 2.77, 2.77, 1.98, 1.98, 1.38, 1.38, 0.91, 0.91, -0.08, 5.06, 2.28, 1.14, 10.42, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 19.39, 6.04, 6.61, 5.11, 0, 0, 5.32, 0, 18.88, 13.15, 0, 0, 5.32, 0, 0, 0, 24.30, 0, 0, 12.84, 0, 0, 0, 32.26, 0, 0, 0, 12.65, 12.97, 6.42, 0, 0, 5.32, 5.11, 0, 0, 0, 11.66, 0, 0, 0.40, 2.38, 1.39, 0, 1, 7, 2, 2, 0, 1, 1, 0, 0, 0, 2, 2, 2, 1, 0, 1, 1, 1, -0.27, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0.24, 1, 1, 1, 0, 0, 1.38, 2.13, 0, 0, 0, 0, 0, 0, 0, 0, 0.89, 13.06, 1.89, 0.01, 12.41, 0, 0, 0, 0, 18.86, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 31.27, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, 0, 0, 0, 2, 0, 0, 0, 0, 0, -5.99, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, and the one-dimensional feature vector is {8.50, 0.31, 8.50, 0.31, 0.48, 1.71, 2.57, 2.86, 2.14, 5.11, 4.35, 4.35, 3.43, 2.77, 2.77, 1.98, 1.98, 1.38, 1.38, 0.91, 0.91, -0.08, 5.06, 2.28, 1.14, 10.42, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 19.39, 6.04, 6.61,5.11,0,0,5.32,0,18.88,13.15,0,0,5.32,0,0,0,24.30,0,0,12.84,0,0,0,32.26,0,0,0,12.65,12.97,6.42,0,0,5.32,5.11,0,0,0,11.66,0,0,0.40,2.38,1.39,0,1,7,2,2,0,1,1,0,0,0,2,2,2,1,0,1,1,1,-0.27,0,1,1,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0.24,1,1,1,0,0,1.38,2.13,0,0,0,0,0,0,0,0,0,0.89,13.06,1.89,0.01,12.41,0,0,0,0,18.86,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,0,0,0,0,0,31.27,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,2,0,0,0,0,0,0,0,0,0,0,2,0,0,0,0,0,-5.99,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0}, The vector length of the one-dimensional feature vector is 302.
8. The efficient screening method of chiral synthesis catalysts based on automated machine learning according to claim 6, characterized in that Step S4 is to predict the 507 new catalyst molecules and sort them in descending order according to the predicted values.