Intelligent Analysis Method for Organic Chemical Synthesis Based on Topological Machine Learning

Through topological machine learning combined with LightGBM algorithm, the topological features of three-dimensional chemical structures were extracted, and the problem of inaccurate yield prediction in the Buchwald-Hartwig amination reaction was solved, efficient and accurate yield prediction and reaction condition analysis were achieved, and the intelligent and green development of organic chemical synthesis was promoted.

CN115910225BActive Publication Date: 2025-07-11HENAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211425974.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-07-11
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

The prior art has insufficient structural information extracted by three-dimensional chemical descriptors in the Buchwald-Hartwig amination reaction, and the relationship between yield and reaction conditions is not effectively paid attention to, resulting in complex reaction conditions, time-consuming and costly, and it is difficult to meet the efficient synthesis needs of complex molecules.

Method used

The topological machine learning method is adopted to extract the topological features of three-dimensional chemical structures through topological data analysis, and model training is carried out in combination with the LightGBM algorithm. The grid search method is used to optimize parameters to achieve intelligent prediction of yield and deep mining of reaction conditions.

Benefits of technology

It improves the accuracy and efficiency of yield prediction, simplifies the operation process, provides more reliable decision-making information, greatly accelerates the chemical research and development process, and realizes green chemistry and intelligent chemistry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910225B_ABST
    Figure CN115910225B_ABST
Patent Text Reader

Abstract

The present invention proposes an intelligent analysis method for organic chemical synthesis based on topological machine learning, including: acquisition of topological features: extracting topological invariants from three-dimensional structure descriptors through topological data analysis to obtain topological features, and cascading the three-dimensional structure descriptors and topological features; intelligent prediction: training and predicting the cascaded features through the LightGBM algorithm, obtaining the optimal parameters of the LightGBM algorithm by using the grid search method to obtain the LightGBM model, and predicting the chemical reaction yield by using the LightGBM model; correlation analysis of yield and reaction conditions: performing clustering analysis on the cascaded features according to the chemical reaction yield through topological data analysis to explore the relationship between the yield and reaction conditions. The present invention can automatically and efficiently perform intelligent prediction on the organic chemical reaction yield, can deeply explore the internal relationship between the yield and reaction conditions, provide reliable decision-making information for users, and accelerate the chemical R & D process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of organic chemical synthesis based on applied mathematics and artificial intelligence, and in particular to an intelligent analysis method for organic chemical synthesis based on topological machine learning. Background Art

[0002] Coupling reactions are very important in organic chemical synthesis, and their products are widely used in pharmaceuticals, pesticides, natural products, and even advanced functional materials. The reaction of amino acids combining into proteins is also a coupling reaction. In the past few decades, transition-metal-catalyzed coupling reactions have developed rapidly. Among them, palladium (Pd)-catalyzed cross-coupling reactions are a major class of their applications. This type of reaction has high efficiency, good selectivity, and mild reaction conditions, and is an effective means of modern organic synthesis.

[0003] Some complex and important drug molecules and organic materials need to combine carbon atoms through chemical reactions during the manufacturing process. However, the carbon-carbon bonds in organic compounds are relatively stable, and it is difficult to directly undergo chemical reactions with other molecules. Although some methods can make carbon atoms more active, they also produce a large amount of by-products, which not only have low efficiency but may also cause certain pollution. Using a palladium catalyst can solve this problem to a certain extent. Palladium atoms will attract different carbon atoms, making the distance between carbon atoms closer and easier to combine, which is called "coupling". In this way, it is not necessary to activate carbon atoms to a very active degree, and the corresponding by-products will also be relatively few, and the reaction will be more precise and efficient. Richard F. Heck discovered as early as the 1970s that a palladium catalyst can achieve the connection between carbon atoms under relatively mild conditions. Subsequently, Ei-ichi Negishi and Akira Suzuki further developed the method of using palladium to catalyze carbon-carbon bond cross-coupling, further expanding the types of substrates and products of this type of chemical reaction. In 2010, Richard F. Heck, Ei-ichi Negishi, and Akira Suzuki won the Nobel Prize in Chemistry of that year for developing the "palladium-catalyzed cross-coupling method in organic synthesis". The birth of these coupling synthesis methods has unprecedentedly improved the ability and level of chemists to manipulate atoms and molecules.

[0004] Palladium-catalyzed coupling reactions can not only realize the bonding of carbon atoms with carbon atoms, but also the bonding of carbon atoms with other atoms. Among them, the most representative is the carbon-nitrogen (CN) coupling reaction of carbon atoms with nitrogen atoms. The formation of CN bonds is an important field in modern organic chemistry and biochemistry. Amines and their derivatives, nitrogen-containing heterocycles, etc. can be prepared by generating CN bonds. Many of them are compounds with biological and medicinal activities and important intermediates for the synthesis of some substances, such as antihypertensive drugs, drugs for the treatment of epilepsy, anesthetics, etc.; aromatic amines also play an important role in the fields of medicine and functional materials, such as drugs for the treatment of headaches, anti-leprosy drugs, important intermediates for dyes, semiconductor materials, functional materials, etc. Since it was reported in the early twentieth century, the UIImann reaction has become a classic method for constructing CN bonds, but the catalyst used in this method is metallic copper, the reaction is violent, and the functional group compatibility is poor, so it is difficult to meet the requirements of efficient synthesis of organic functional molecules, especially the modification of complex molecules. Since 1994, the development of palladium-catalyzed Buchwald-Hartwig coupling reaction has made significant progress in metal dosage, reaction conditions, substrate universality, etc. in CN bonding. However, the reaction conditions are complex, time-consuming, and costly. Traditional chemical experiments require a lot of manual trial and error, which makes it challenging to apply this reaction to complex drug-like molecules. Therefore, how to obtain higher reaction yields and corresponding reaction combinations with less cost and high efficiency has become the most concerned issue for researchers.

[0005] In recent years, with the increasing availability of big data, machine learning has gradually been applied to the field of chemistry as an efficient method. It has also brought new opportunities for the development of organic chemistry, and has achieved the prediction of catalyst activation performance, chemical reaction performance, compound properties, etc. Researchers form a certain expression of chemical information, namely descriptors, by performing a certain form of calculation, screening or encoding of information in the chemical system. In this way, research in the field of chemistry can be transformed into a data processing process, thereby reducing dependence on personnel to a certain extent. Machine learning can mine the correlation of massive data generated in chemical reaction experiments to help chemists make reasonable analysis and predictions. In addition, machine learning can also deeply explore the relationship between reaction products and reaction conditions, which helps to design the required chemical materials more efficiently, greatly accelerate the chemical research and development process, and realize green chemistry and intelligent chemistry. Summary of the invention

[0006] In view of the problems that in the current prediction of the yield of Buchwald-Hartwig amination reaction by machine learning, the structural information extracted relying on three-dimensional chemical descriptors is insufficient, and the relationship between the yield and reaction conditions is not concerned, etc., the present invention proposes an intelligent analysis method for organic chemical synthesis based on topological machine learning. This method can combine three-dimensional chemical structure information and topological information to automatically and efficiently make an intelligent prediction of the chemical reaction yield, and deeply explore the relationship between the yield and reaction conditions, facilitating the research of subsequent relevant researchers; the entire model training takes a short time, has high recognition accuracy and good robustness.

[0007] The technical solution of the present invention is realized as follows:

[0008] An intelligent analysis method for organic chemical synthesis based on topological machine learning, the steps are as follows:

[0009] Step 1: Obtaining topological features: Extract topological invariants from three-dimensional structure descriptors through topological data analysis to obtain topological features, and cascade the three-dimensional structure descriptors and topological features;

[0010] Step 2: Intelligent prediction: Train and predict the cascaded features through the LightGBM algorithm, obtain the best parameters of the LightGBM algorithm by using the grid search method to get the LightGBM model, and use the LightGBM model to predict the chemical reaction yield;

[0011] Step 3: Correlation analysis of yield and reaction conditions: According to the chemical reaction yield, perform clustering analysis on the cascaded features through topological data analysis to explore the relationship between the yield and reaction conditions.

[0012] The implementation method of Step 1 is as follows:

[0013] S1.1. Import the three-dimensional structure descriptor into topological data analysis to generate a persistence diagram, and then vectorize the persistence diagram through relevant methods to output topological features;

[0014] S1.2. Cascade the three-dimensional structure descriptor and topological features, and correspond the cascaded features with the yield one by one and divide them into a training set and a test set.

[0015] The specific calculation process of the topological structure in Step S1.1 is as follows:

[0016] S1.1.1. Import the three-dimensional descriptor information into the topological data analysis algorithm, and convert the topological information therein into a persistence diagram;

[0017] S1.1.2. Record the change of each topological invariant through the persistence diagram;

[0018] Among them, the persistence diagram represents the results of persistent homology analysis as pairs of birth times and death times. The horizontal axis represents the filtration value at the birth of the topological invariant, and the vertical axis represents the filtration value at the death of the topological invariant. Denote the birth position of each topological invariant on the filtration axis as b α and the death position of each topological invariant on the filtration axis as d α . Then p α =d α -b α represents the survival period of each topological invariant;

[0019] S1.1.3. Obtain topological features by vectorizing the persistence diagram: the actual number of persistent connected components H0, cyclic structures H1, and cavity structures H2, the average survival period of connected components H0, cyclic structures H1, and cavity structures H2, and the persistent entropy; among them, the persistent entropy D ={(b α ,d α )} α∈A . The persistent entropy D is calculated according to .

[0020] The implementation method of the LightGBM model is as follows: Import the data of the training set and test set obtained in step S1.2 into the LightGBM algorithm. Use the grid search method to permute and combine the possible values of multiple parameters in the LightGBM algorithm. By calculating the loss function value of each iteration in the LightGBM algorithm until the loss function value converges to the minimum, output the prediction result and the corresponding parameter value. Finally, select the parameter corresponding to the best prediction result and save the LightGBM model. The objective function of the LightGBM algorithm is:

[0021]

[0022] Among them, is the loss function on the linear space; i is the i-th sample; is the predicted value of the i-th sample x i ; is the k-th tree, and K is the number of trees; y i is the true value; f k (x i ) represents the score of each tree for the i-th sample x i .

[0023] The implementation method of step three is as follows:

[0024] S3.1. According to the statistical concept of quantiles, classify the chemical reaction yields into two categories: low yields and high yields;

[0025] S3.2. Import the cascaded features obtained in step S1.2 into topological data analysis. The user adjusts the interval between adjacent filtering value intervals and the overlapping interval according to the data characteristics, and sets the number of intervals for the single-link clustering histogram to obtain the optimal clustering result.

[0026] S3.3. According to the clustering result in step S3.2, analyze the reaction conditions in each cluster of samples, and then make a comparative analysis to obtain the reaction conditions corresponding to the high yield.

[0027] The implementation method of step S3.2 is as follows:

[0028] S3.2.1. Calculate a filtering value for each data point using the centrality index L-infinity of the distance matrix:

[0029]

[0030] where d is the original data, len(d) represents the sample size, n represents the number of features, d[j] represents the j-th sample, and d[j][0] represents the first feature of the j-th sample.

[0031] S3.2.2. Divide the data points into different filtering value intervals in ascending order of the filtering value L-infinity; there is an overlapping area between adjacent filtering value intervals, where the interval between adjacent filtering value intervals is N and the overlapping interval is P.

[0032] S3.2.3. Use single-link clustering to cluster the data in each filtering value interval.

[0033] S3.2.4. Put the small classes obtained by clustering each filtering value interval together, and each small class is represented by a circle; if there are the same original data points between two classes, add an edge between them.

[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] 1) The present invention extracts the topological information of the three-dimensional descriptor through topological data analysis (TDA), and then inputs the three-dimensional descriptor and topological features in cascade, so that the prediction accuracy of the model is greatly improved; in addition, through topological data analysis (TDA), the relationship between the yield and the reaction conditions is deeply explored, so as to provide more reliable decision-making information for users.

[0036] 2) The LightGBM adopted in the present invention reduces the number of data samples and features during calculation, making it faster and more efficient. The grid search method can simultaneously search for possible values of a certain parameter or multiple parameters, thus easily obtaining the optimal parameters of the model; the combination of LightGBM and the grid search method for predicting organic chemical synthesis is more accurate and efficient.

[0037] 3) The operation of the present invention is simple and easy to implement, and the analysis results are relatively accurate, greatly accelerating the chemical R & D process, facilitating the use of relevant users, and realizing green chemistry and intelligent chemistry. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 It is the reaction formula of the chemical reaction and the structural formula of related variables in the embodiment of the present invention;

[0040] Figure 2 It is the two-dimensional structure diagram of a certain reaction condition combination and the corresponding three-dimensional structure diagram;

[0041] Figure 3 It is the flow chart of the present invention;

[0042] Figure 4 It is the correlation analysis result of the yield and reaction conditions obtained by different methods; among them, (a) K-Means, (b) PCA, (c) t-SNE, (d) UMPA, (e) TDA.

[0043] In the figure: Equation: Buchwald-Hartwig amination reaction and the variable selection range in the Buchwald-Hartwig amination reaction, Aryl: halide, Additive: additive, Base: base, L (Ligand): ligand. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0045] AsFigure 3 As shown in the figure, an embodiment of the present invention provides an intelligent analysis method for organic chemical synthesis based on topological machine learning, including obtaining topological features, intelligent prediction, and correlation analysis of yield and reaction conditions; the specific steps are as follows:

[0046] Step 1: Obtaining topological features: Extract topological invariants from three-dimensional structure descriptors through Topological Data Analysis (TDA) to obtain topological features, and concatenate the three-dimensional structure descriptors with the topological features; through the persistent homology tool in topological data analysis, convert the persistence diagram of topological invariants (topological invariants are topological structures in topological spaces whose properties do not change with specific types of continuous changes, such as connected components H0, cyclic structures H1, and cavity structures H2, etc.) into a structured feature represented by a quantified vector, so as to extract the topological information of the three-dimensional descriptor; specifically including:

[0047] S1.1: Import the three-dimensional structure descriptor into the topological data analysis (TDA) algorithm, so as to convert the topological information therein into a persistence diagram, and then vectorize the persistence diagram through relevant methods to output topological features.

[0048] S1.1.1: Import the three-dimensional descriptor information into the topological data analysis algorithm to convert the topological information therein into a persistence diagram; the data acquisition of the three-dimensional descriptor is to convert the two-dimensional structure diagram of the relevant chemical structure drawn into a three-dimensional structure diagram, and then use relevant software to calculate its three-dimensional structure descriptor. As Figure 1 shown in this embodiment, when drawing the two-dimensional structure diagram, arrange and combine the reaction variables in the Buchwald-Hartwig amination reaction (including halides, ligands, bases, and additives); and use ChemOffice software to convert each combination of the drawn two-dimensional structure diagram (as Figure 2 shown) into a three-dimensional structure diagram (as Figure 2 shown) combination. There are 15 kinds of halides, 4 kinds of ligands, 3 kinds of bases, and 23 kinds of additives, and the corresponding permutations and combinations are 4140 kinds; excluding invalid reactions, finally 3960 valid reactions are obtained, and these reactions correspond to their reaction yields one by one. Then use the RDKit toolkit to calculate and output the three-dimensional structure descriptor of each three-dimensional structure diagram in Python.

[0049] S1.1.3. Obtain topological features by vectorizing the persistence diagram: For some construction methods of topological features, features can be constructed for the homology group connection components H0, cyclic structures H1, and cavity structures H2 of different dimensions respectively. Of course, appropriate selection also needs to be made according to the data characteristics. For example, the actual continuously existing quantities such as the connection components H0, cyclic structures H1, and cavity structures H2 can be calculated, the average lifetimes of the connection components H0, cyclic structures H1, and cavity structures H2, etc. can be calculated, and distances based on Persisstence Landscape, Bottleneck distance, Wasserstein distance, persistent entropy, and their amplitudes can also be introduced; the persistent entropy D = {(b α ,d α )} α∈A , according to where etc. to "vectorize" the persistence diagram.

[0050] S1.2. Concatenate the three-dimensional structure descriptors and topological features of all reaction combinations obtained by calculation, and divide the concatenated features into a training set (70%) and a test set (30%) after corresponding them one by one with the yields, and correspond them with the corresponding reaction yields, so as to facilitate in-sample and out-of-sample prediction by the LightGBM algorithm.

[0051] Step 2: Intelligent prediction: Train and predict the concatenated features through the distributed gradient boosting tree model LightGBM algorithm, use the grid search method to obtain the optimal parameters of the LightGBM algorithm to obtain the LightGBM model, and use the LightGBM model to predict the chemical reaction yield; the grid search method is embedded in this algorithm, permute and combine the possible values of multiple parameters, and select the optimal parameters through the model prediction results; specifically including:

[0052] S2.1. Import the training set and test set data into the LightGBM algorithm, use the grid search method to permute and combine the possible values of multiple parameters in the LightGBM algorithm, calculate the loss function value of each iteration in the LightGBM algorithm until the loss function value converges to the minimum or converges to a certain number of times, then output the prediction result and the corresponding parameter value, and finally select the parameters corresponding to the best prediction result and save the model.

[0053] This embodiment automatically obtains the optimal parameters by using the grid search method, mainly because there are many parameters in the LightGBM algorithm, and the grid search method can specify a certain parameter value for exhaustive search or permute and combine the possible values of multiple parameters, so as to find the optimal parameters, which is more efficient than manual parameter tuning.

[0054] The objective function of the LightGBM algorithm is:

[0055]

[0056] Among them, is the loss function on the linear space; i is the i-th sample; is the predicted value of the i-th sample x i : is the k-th tree, and K is the number of trees; y i is the true value; f k (x i ) represents the score of each tree for the i-th sample x i .

[0057] Since the objective function in the LightGBM algorithm can be freely selected as long as it satisfies first-order differentiability, the objective function of the LightGBM algorithm in the present invention selects the squared loss function:

[0058] Finally, the model will return the predicted value when the loss function reaches the minimum, and judge the prediction effect of the model through evaluation indicators.

[0059] S2.2. Conduct out-of-sample prediction to prove the effectiveness of the model; out-of-sample prediction means predicting data outside the model prediction samples. If the out-of-sample prediction is effective, it can prove that the model selected in the present invention can predict the chemical reaction yield, and determine the combination of reactants and reaction conditions to provide the reaction combination corresponding to the highest yield.

[0060] Step three: Correlation analysis of yield and reaction conditions: According to the chemical reaction yield, use topological data analysis to perform clustering analysis on the cascaded features to explore the relationship between the yield and reaction conditions. Specifically, it includes:

[0061] S3.1. According to the statistical concept of quantiles, divide the chemical reaction yield into two categories: low yield (less than the 0.5 quantile, that is, less than the sample median) and high yield (greater than the 0.5 quantile, that is, greater than the sample median).

[0062] In statistics, the quantile is the point at which a batch of data is divided by the probability. The meaning of the quantile represents the proportion of the data subset less than a certain value in the entire sample set after the data set is arranged from small to large, providing a good basis for discovering data outliers and observing the data distribution.

[0063] S3.2. Import the cascaded features obtained in step S1.2 into topological data analysis (TDA). The user adjusts the interval (N) between adjacent filtering value intervals and the overlapping interval (P) according to the data characteristics, and sets the appropriate number of intervals (K) for the single-link clustering histogram to obtain the best clustering result.

[0064] The specific calculation process of step S3.2 includes:

[0065] S3.2.1. Calculate a filtering value for each data point using a filtering function. In the present invention, the centrality index of the distance matrix is selected, L-infinity (the value of L-infinity is the distance from this point to the farthest point from it):

[0066]

[0067] where d is the original data, len(d) represents the sample size, n represents the number of features, d[j] represents the j-th sample, and d[j][0] represents the first feature of the j-th sample;

[0068] S3.2.2. Divide the data points into different filtering value intervals from small to large according to the filtering value L-infinity; there is an overlapping area set for adjacent filtering value intervals, that is, the points in the overlapping area belong to two intervals at the same time. That is, here a set of N equal-length intervals is determined by two resolution parameters (N intervals and P overlapping percentages), and the overlapping percentage of these intervals is P of the interval length.

[0069] S3.2.3. Cluster the data in each filtering value interval using single-link clustering; cluster the data in each interval. Here, the present invention uses single-link clustering to cluster each group. Let N' be the number of points in the bin. First, construct a single-link dendrogram for the data in the bin and record the threshold of each transition in the clustering. The present invention selects an integer K' and constructs a K'-interval histogram of these transition values. Use the last threshold before the first gap in the histogram for clustering. The larger the value of K', the more clusters are generated, and the smaller the value of K', the fewer clusters are generated.

[0070] S3.2.4. Put together the small classes obtained by clustering each filtering value interval. Each small class is represented by a circle of different sizes. The size of the circle represents the amount of the contained sample size. The larger the circle, the more samples are contained in this class; the depth of the circle color represents the mean value of the sample labels in this class (here it is the mean value of the yield). The darker the circle color, the larger the label mean value; if there are the same original data points between two classes (this is the reason why the intervals need to overlap each other), then add an edge between them. The closer the distance between the two circles, the closer the internal samples are related, or the more similar these two classes are.

[0071] S3.3. According to the clustering results in step S3.2, analyze the reaction conditions in each cluster of samples, and then compare and analyze to obtain the reaction conditions corresponding to high yield.

[0072] Topological Data Analysis (TDA) performs better than other clustering algorithms when dealing with high-dimensional data because its analysis process does not cause information loss and is considered stable for missing and noisy samples. It utilizes the latest developments in applied topology to identify the shape features of the dataset, which can not only show the relationships between clusters and variables but also obtain a higher-level understanding of the high-dimensional data structure.

[0073] Simulation experiment:

[0074] The system of the present invention is further demonstrated by simulation experiments. Taking the Buchwald-Hartwig amination reaction in organic chemical synthesis as an example (the chemical reaction formula is as Figure 1 shown), it is selected as the user data and imported into the chemical reaction.

[0075] (1) Intelligent prediction

[0076] Table 1 Prediction effects with and without topological features

[0077]

[0078] As shown in Table 1, taking the goodness of fit R 2 , root mean square error RMSE, mean absolute error, and running time as evaluation indicators, the prediction effects with and without topological features are compared, as well as the prediction effects of LightGBM and random forest. As shown in the above table, the prediction results of the cascade of three-dimensional descriptors and topological features are better than those of the single three-dimensional descriptors, and the prediction results of LightGBM are better than those of random forest.

[0079] (2) Correlation analysis of yield and reaction conditions

[0080] As Figure 4 shown, obviously, methods such as K-Means, PCA, TSNE, and UMAP cannot effectively classify the Buchwald-Hartwig coupling reaction data. While TDA can obtain more detailed sample stratification results of clustering analysis methods such as K-Means, PCA, TSNE, and UMAP. Furthermore, the relationship between yield and reaction conditions can be further analyzed, thus providing corresponding decision-making information for researchers.

[0081] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent analysis method for organic chemical synthesis based on topological machine learning, characterized in that, The steps are as follows: Step 1: Acquisition of topological features: Extraction of topological invariants of three-dimensional structure descriptors through topological data analysis. Obtaining topological features and concatenating the three-dimensional structure descriptor with the topological features; The implementation method of step one is: S1.1, import the three-dimensional structure descriptor into the topological data analysis to generate the persistence map, then vectorize the persistence map through the relevant method and output the topological features; The specific calculation process of the topological structure is: S1.1.1, importing the three-dimensional descriptor information into the topological data analysis algorithm, and converting the topological information therein into a persistence graph; S1.1.

2. Record the changes of each topological invariant through a persistence graph; Among them, the persistence diagram represents the results of persistent homology analysis as pairs of birth times and death times. The horizontal axis represents the filtration value at the birth of the topological invariant, and the vertical axis represents the filtration value at the death of the topological invariant. Denote it as b α Record the position where each topological invariant is born on the filtration axis, denoted as d α Record the position where each topological invariant dies on the filtration axis, then p α = d α - b α represents the survival period of each topological invariant; S1.1.

3. Obtain topological features by vectorizing the persistence diagram: the true persistent numbers of connected components H0, cyclic structures H1, and cavity structures H2, the average lifetimes of connected components H0, cyclic structures H1, and cavity structures H2, and the persistence entropy; where the persistence entropy D = {(b α , d α )} α∈A , and the persistence entropy D is calculated according to . S1.2, cascading the three-dimensional structure descriptor and the topological features, and dividing the cascaded features into a training set and a test set after corresponding one-to-one with the yield; Step 2: Intelligent prediction: Train and predict the cascaded features through the LightGBM algorithm, use the grid search method to obtain the optimal parameters of the LightGBM algorithm to obtain the LightGBM model, and use the LightGBM model to predict the chemical reaction yield; Step 3: Correlation analysis between yield and reaction conditions: Based on the chemical reaction yield, topological data analysis is used to perform cluster analysis on the features after cascading to explore the relationship between yield and reaction conditions.

2. The intelligent analysis method for organic chemical synthesis based on topological machine learning according to claim 1, characterized in that The implementation method of the LightGBM model is as follows: import the data of the training set and the test set obtained in step S1.2 into the LightGBM algorithm, use the grid search method to arrange and combine the possible values ​​of multiple parameters in the LightGBM algorithm, calculate the loss function value of each iteration in the LightGBM algorithm until the loss function value converges to the minimum, output the prediction result and the corresponding parameter value, and finally select the parameters corresponding to the best prediction result and save the LightGBM model.

3. The intelligent analysis method for organic chemical synthesis based on topological machine learning according to claim 2, wherein The objective function of the LightGBM algorithm is: Among them, is a loss function on a linear space; i is the i-th sample; is the predicted value of the i-th sample x i : k = 1, 2, ..., K is the k-th tree, where K is the number of trees; y i is the true value; f k (x i ) represents the score of each tree for the i-th sample x i .

4. The intelligent analysis method for organic chemical synthesis based on topological machine learning according to claim 1, wherein The implementation method of step three is: S3.

1. Based on the statistical concept of quantiles, chemical reaction yields are divided into low yield and high yield categories; S3.2, import the cascaded features obtained in step S1.2 into the topological data analysis. The user adjusts the interval and overlapping interval of adjacent filter value intervals according to the data characteristics, and sets the number of single-link clustering histogram intervals to obtain the best clustering result; S3.

3. According to the clustering results in step S3.2, the reaction conditions in each cluster of samples are analyzed, and then compared and analyzed to obtain the reaction conditions corresponding to the high yield.

5. The intelligent analysis method for organic chemical synthesis based on topological machine learning according to claim 4, wherein The implementation method of step S3.2 is: S3.2.

1. Use the distance matrix centrality index L-infinity to calculate a filter value for each data point: Where d is the original data, len(d) represents the sample size, n represents the number of features, d[j] represents the jth sample, and d[j][0] represents the first feature of the jth sample; S3.2.

2. Divide the data points into different filtering value intervals in ascending order of the filtering value L-infinity; there is an overlapping area set for adjacent filtering value intervals, where the interval between adjacent filtering value intervals is N and the overlapping interval is P; S3.2.

3. Use single-link clustering to cluster the data in each filtering value interval; S3.2.

4. Put together the small classes obtained by clustering each filtering value interval, and each small class is represented by a circle; if there are the same original data points between two classes, add an edge between them.

Citation Information

Patent Citations

  • Metal organic framework material structure characteristic rapid evaluation method based on machine learning

    CN112382352A

  • XGBoost-based chemical reaction yield intelligent prediction and analysis method in small sample environment

    CN113517033A