A prediction method for the quality of catalyst ligands in acetylene hydrochlorination based on support vector machine algorithm
Through the support vector machine algorithm training model, topological expansion and ECFP fingerprint technology are used to solve the problem of time-consuming and labor-consuming traditional methods in the design and prediction of mercury-free catalysts, and efficient and accurate catalyst screening is achieved.
Patent Information
- Application Number
- CN202411670052.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-21
AI Technical Summary
Traditional methods are time-consuming and labor-intensive in designing and predicting mercury-free catalysts, and it is difficult to effectively screen highly efficient catalysts.
The support vector machine algorithm (SVM) is used to train the model, convert the ligand into binary form through topological expansion, and data input is used using 64-bit ECFP fingerprints to find the best hyperplane for classification, so as to achieve prediction of new samples.
Efficient design and prediction of mercury-free catalysts in the case of small sample sizes have high accuracy, which can greatly shorten the time for screening highly efficient catalysts.
Smart Images

Figure CN119541731B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the quality of acetylene hydrochlorination reaction catalyst ligands based on a support vector machine algorithm, in particular to a method combining catalysis and support vector machine algorithm knowledge design and realizing a set of artificial intelligence-assisted screening of the quality of acetylene hydrochlorination reaction catalyst ligands, belonging to the technical field of materials informatics. Background Art
[0002] Vinyl chloride is an important chemical raw material used to polymerize polyvinyl chloride (PVC). Polyvinyl chloride is widely used in plastic products and is the third largest plastic in the world after polyethylene and polypropylene. As a monomer of polyvinyl chloride, how to improve the efficiency and output of vinyl chloride production is the most important. At present, the most commonly used method is the hydrochlorination of acetylene to prepare vinyl chloride. In this chemical production process, the traditional practice is to use a catalyst containing HgCl2 to hydrochlorinate acetylene. Hg, as a heavy metal element, can cause great harm to the human body and the natural environment. In addition to this environmental problem, finding a catalyst that is equally efficient and clean and pollution-free has become the most important issue at present. At present, the research on mercury-free catalysts is mainly based on the following four directions: catalysts based on ruthenium, platinum, gold and copper. Based on these four types of metals as substrates, designing a class of efficient hydrochlorination catalysts is the main research direction at present.
[0003] The design and prediction of catalysts is also the most difficult process before the catalyst is adopted. The performance of catalysts is affected by many factors. Screening efficient catalysts through experiments is the most traditional and widely used method. However, experimental methods usually consume a lot of time and financial costs, which is also a common problem of traditional design and prediction methods. At present, with the rapid development of artificial intelligence, artificial intelligence has gradually been widely used in various fields. This patent chooses to train a model based on the support vector machine (SVM algorithm) algorithm with the help of artificial intelligence to facilitate the design and prediction of high-performance catalysts for acetylene hydrochlorination reactions.
[0004] The main research direction is to use machine learning for prediction, and the algorithm used is the support vector machine algorithm (SVM). The support vector machine algorithm (SVM) is a particularly powerful and flexible supervised learning model that can analyze data for classification and regression. Its usual algorithm complexity scales polynomially with the dimension of the data space and the number of data points. The model is trained by incorporating factors such as ligands into the range of parameters that need to be considered. In this patent, the support vector machine algorithm can efficiently design and predict mercury-free catalysts with a small sample size. And the accuracy is high, and the mercury-free catalysts calculated by the model also show high performance. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a new method for predicting the central ligand screening of mercury-free catalysts for acetylene hydrochlorination by using a support vector machine algorithm (SVM). The method first converts the ligand into a binary form by topological expansion, and then inputs different ligands in an existing database in the form of 64-bit ECFP fingerprints, and separates the data points by finding the best hyperplane, thereby realizing the classification of new samples. In the screening of central ligands of catalysts, SVM can infer the effect of unknown catalysts based on the characteristics and performance of known catalysts. The method can calculate the ECFP fingerprint of the relevant ligand based on the SMILES string to perform specialized prediction on the conversion rate, thereby improving the accuracy of the prediction.
[0006] In order to solve the technical problem of the present invention, the present invention is implemented by the following technical scheme: a method for predicting the quality of acetylene hydrochlorination catalyst ligand based on support vector machine algorithm, the method comprising the following steps:
[0007] (1) Extracting datasets from public literature and classifying them to train the model,
[0008] (2) Extract the composite SMILES string of the ligand molecule and convert it into a tensor matrix.
[0009] (3) Input the tensor matrix into the model based on the support vector machine algorithm for regression training, using the RBF kernel function and a penalty coefficient C equal to 0.9;
[0010] (4) The SMILES string of the unknown performance ligand is input to obtain the prediction of the ligand modification results.
[0011] Preferably, the catalyst ligands and the catalyst performance after ligand modification are used to train the model: by collecting many compounds to naturally form a corpus to identify SMILES strings, using ECFP (Extended Connectivity Fingerprint) technology to collect information and perform regression training.
[0012] Preferably, the ligand features (SMILES strings) are used to construct a tensor matrix and input into the model, and the trained LSVM model is used to convert the SMILES strings of the ligands into tensors to output prediction results.
[0013] Preferably, the method comprises the following steps:
[0014] Step 1: Establish a data set and collect information on known catalyst ligands from the literature. The molecular structure of the ligand and the corresponding composite SMILES string and conversion rate are recorded in the database;
[0015] Step 2: Convert the molecular structure of the ligand into an ECFP fingerprint
[0016] The composite SMILES string of each ligand is converted into a 64-bit binary vector called "ECFP fingerprint". For the SMILES string of known ligands, a search method is used to map each functional group to a molecular fingerprint composed of 0 and 1 from the RDKit open source library, and then combine them into a 64-bit molecular fingerprint. Different sequences of 0 and 1 represent different functional groups, thus reflecting the influence of different ligand structures.
[0017] Step 3: Construct the feature tensor matrix
[0018] Each ligand collected has its own 64-bit molecular fingerprint. The fingerprint data of all ligands are integrated into a tensor matrix. Each row of the matrix represents a ligand and each column is a feature bit.
[0019] Step 4: Build and tune the model architecture
[0020] Input the complete feature matrix into the support vector machine model for preliminary training. The support vector machine model will automatically calculate and fit; try two kernel functions (linear kernel, polynomial kernel and RBF kernel) respectively, and use the cross-validation method to optimize the selection of hyperparameters (penalty coefficient C);
[0021] Step 5: Model Evaluation
[0022] The simplified feature matrix is input into the support vector machine model and training continues. The target variable is the conversion rate. Through model training, the model gradually learns the relationship between ligand structure, experimental conditions and conversion rate. After training, the model is evaluated using the validation set. Assume that the conversion rate of ligand A is predicted to be 78%, while the actual conversion rate is 80%, with an error of 2%;
[0023] Through model training, the parameters were determined as follows: the optimal model was finally obtained using the RBF kernel and the penalty coefficient C was equal to 0.9, and the model was evaluated using the validation set; the probability that the difference between the predicted conversion rate and the actual conversion rate of different ligands was less than 7% reached more than 95%;
[0024] Step 6: Predict the conversion rate of unknown ligands
[0025] After training, the model can be used to predict the conversion rate of new ligands. For example, if there is a ligand X, RDKit is used to generate a 64-bit fingerprint of ligand X, which is then used in the model for prediction. The model outputs the predicted conversion rate improvement. This output result can help determine whether ligand X has a good catalytic effect under the reaction conditions.
[0026] The ligands are screened by using the prediction method of the quality of acetylene hydrochlorination catalyst ligands based on the support vector machine algorithm. As a modifier for acetylene hydrochlorination catalyst.
[0027] A computer-readable storage medium has a computer program, and the computer program can run the method for predicting the quality of acetylene hydrochlorination reaction catalyst ligand based on a support vector machine algorithm.
[0028] A device for predicting the quality of ligands of acetylene hydrochlorination reaction catalyst based on a support vector machine algorithm, wherein the device is equipped with the computer-readable storage medium.
[0029] The specific steps of the method for predicting the conversion rate of acetylene hydrochlorination by adding unknown ligands are as follows: (1) Collect historical experimental data, including ligands added to various catalysts and the corresponding catalytic effects, determine the characteristics related to catalytic performance, and establish a database based on this. (2) Select appropriate kernel functions (linear kernel and RBF kernel) according to the distribution characteristics of the data, and use the cross-validation method to select appropriate hyperparameter optimization to improve the accuracy of the model. (3) Input the 64-bit ECFP fingerprint of the unknown ligand to train the SVM model on the training set, and improve the accuracy of the model by continuously adjusting the parameters. (4) Evaluate the performance of the model on the test set. Once the model is verified, it can be used to predict new catalyst ligands to screen out potential candidates.
[0030] First, we collected a large amount of catalyst data by collecting historical experimental data and checking relevant literature. Each catalyst needs to include the ligands added to the catalyst and the catalytic effect, and determine the characteristics related to the catalytic performance, such as the electronic properties of the ligand, spatial configuration and other related information. The structural characteristics of the ligand are quantified and encoded. A database is then established.
[0031] Furthermore, by sorting out the information in the database and deleting incomplete data records, the features of different scales are converted to a unified scale, and the feature data of different dimensions are standardized or normalized so that they have the same scale, avoiding the impact of the model performance due to large differences in data magnitude and structure. The accuracy of the model can be improved by continuously adjusting the parameters (penalty parameters, kernel function).
[0032] Furthermore, according to the distribution characteristics of the data, appropriate kernel functions (linear kernel and RBF kernel) are selected. The cross-validation method is used to select appropriate hyperparameter values for optimization to improve the accuracy of the model. When selecting the kernel function, it depends on the distribution characteristics of the data and the background of the problem. The linear kernel and RBF kernel are tried. In actual application verification, the model accuracy is obtained by comparing the results through hyperparameter optimization, and the optimal choice is determined. At the same time, cross-validation is added to help avoid overfitting, and the model is made generalizable by testing the model on different training sets and validation sets.
[0033] Furthermore, the 64-bit ECFP extended connectivity fingerprint and conversion rate of the unknown ligand are input to train the SVM model on the training set, and the accuracy of the model is improved by continuously adjusting the parameters. ECFP uses a topological algorithm to extract the neighborhood information of atoms or groups in the molecule and constructs it into a series of binary bit representations. After each round of iteration, the new feature values will be mapped to a high-dimensional space, so that similar atomic neighborhood information will be mapped to the same binary bit. The final fingerprint is constructed from these hash values, and the result is a binary vector of fixed length. For example, a 64-bit ECFP fingerprint contains 64 binary bits, and each position represents a topological feature.
[0034] The performance of the model was evaluated on the test set. The model was validated and used to predict new catalyst ligands. The established model was evaluated by inputting the 64-bit ECFP fingerprint of the known ligands, with conversion as the target result quantity.
[0035] Beneficial effects:
[0036] The present invention uses a support vector machine algorithm to train a model. The main advantage of the support vector machine is that it can fit and predict when the flux is small. This type of model can greatly reduce the workload and has a high accuracy rate. We have used this algorithm to design and predict a number of high-performance mercury-free hydrochlorination catalysts, and they have shown high accuracy and catalytic performance.
[0037] Existing catalyst conversion rate prediction methods are usually based on solving the Schrödinger equation, which is too complicated and requires a lot of computing resources. This model uses a common screening method for small models of artificial intelligence, blurring and black-boxing the intermediate process, highlighting the input and output. For the completed training model, the prediction of a single ligand only takes tens of milliseconds, greatly saving computing costs.
[0038] Many traditional catalyst performance prediction methods often use simple molecular descriptors or features to represent ligands. This method is usually difficult to effectively capture the complexity of the molecular structure, resulting in weak prediction capabilities of the model. The present invention extracts the structural features of the ligand through ECFP fingerprints and converts them into 64-bit binary vectors. This method can not only accurately capture the local structural features of the ligand, but also further reflect the higher-order atomic and group interactions in the ligand through the representation of tensor matrices. With this efficient representation method, the model can understand the complex structural information of the ligand more comprehensively and meticulously, thereby improving the accuracy of catalyst performance prediction and the expressiveness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The present invention will be further described below in conjunction with the accompanying drawings.
[0040] Figure 1 It is a schematic diagram of the method flow of the present invention DETAILED DESCRIPTION
[0041] The technical solution of the present invention is further explained below through embodiments in combination with the accompanying drawings.
[0042] Example 1
[0043] Step 1: Data collection and database establishment
[0044] A large amount of catalyst data is collected through historical experimental data. Each catalyst needs to include the ligands added in the preparation of various catalysts and their catalytic effects, determine the characteristics related to catalytic performance, including the electronic properties of the ligands, spatial configuration and other related information, and thus establish a database.
[0045] Step 2: Preprocess the database and construct the tensor matrix
[0046] First, we sorted out the information in the database and removed incomplete records. Then, in order to let the model "recognize" the structure of the ligand, we converted the composite SMILES string of each ligand into a 256-bit binary vector called "ECFP fingerprint". For the SMILES string of the known ligand, we used a search method to map each functional group to a molecular fingerprint composed of 0 and 1 from the RDKit open source library, and combined them into a 64-bit molecular fingerprint. Different sequences of 0 and 1 represent different functional groups, thus reflecting the influence of different ligand structures.
[0047] For example, the composite SMILES string of ligand A is C=CBr, which is retrieved in the RDKit library and converted into the following 64-bit binary fingerprint: [1,1,0,1,0,0,1,...]. Each bit in this string of data represents a detail of the molecular structure, such as whether it contains certain specific chemical bonds or groups. For subsequent model training, this fingerprint becomes the "identity card" of the ligand.
[0048] Assume that each ligand collected has its own 64-bit molecular fingerprint. Integrate the fingerprint data of all ligands into a tensor matrix. Each row of the matrix represents a ligand, and each column is a feature bit. For example, the first two rows can be like this:
[0049]
[0050] Each bit in this string of data represents a detail of the molecular structure, such as whether it contains certain chemical bonds, groups, etc. Then the data is recorded and a database is established.
[0051] Step 3: Build and train the model
[0052] Two kernel functions (linear kernel, polynomial kernel and RBF kernel) were tried, and the cross-validation method was used to optimize the selection of hyperparameters (mainly trying the penalty coefficient C).
[0053] The unknown ligand 64-bit ECFP fingerprint and the conversion efficiency of the catalyst were further input to train the SVM model on the training set, and the accuracy of the model was improved by continuously adjusting the parameters. After that, through model training on the test set, the model gradually learned the relationship between the ligand structure and the conversion rate.
[0054] Step 4: Model Evaluation
[0055] In model training, the model gradually learns the relationship between ligand structure and conversion rate. Two different training methods are adopted successively: 1) The dataset of known acetylene hydrochlorination catalyst ligands is divided into K (5 < K < 10) equal (or nearly equal) subsets. Then, each time one subset is selected as the test set, and the remaining K - 1 subsets are used as the training set. After further iterating K times, K model evaluation results are obtained, and finally the average value of the K evaluations is calculated. 2) Using the model training results determined by the above samples, each sample is used as a test set separately, and the remaining samples are used as the training set. After cross-validation, the model can reduce overfitting and improve the robustness of model evaluation.
[0056] Finally, the optimal model is obtained when using the RBF kernel and the penalty coefficient C is equal to 0.9, and the difference in the predicted catalyst conversion rate is less than 7% in up to 95% of cases.
[0057] Step 5: Result analysis and application
[0058] After completing the above model establishment and adjusting the parameters to a good evaluation, in order to predict the pros and cons of the ligand action of unknown acetylene hydrochlorination catalysts, more than 8000 ligands were collected, converted into SMLILES strings, and brought into the model for prediction. About 100 ligands with excellent prediction effects were obtained. After removing factors such as toxicity, danger, and price, 2 were selected from the remaining ligands for further prediction of their performance in experiments. For the remaining catalyst ligands, no further experiments are required, greatly shortening the time for selecting and designing excellent catalyst ligands.
[0059] More than 8000 ligands were collected, converted into SMLILES strings, with the HCl ratio set to 1.15 and the reaction temperature set to 180 °C, and brought into the model for prediction. About 100 ligands with excellent prediction effects were obtained. After removing factors such as toxicity, danger, and price, 2 were selected from the remaining ligands for verification.
[0060] Step 6: Catalyst preparation
[0061] The ligands with good predicted effects under the reaction conditions given in Step 7 are used to prepare a catalyst with a central metal content of 1% Ru by the impregnation method. The performance of the catalyst is tested by experiments under the same reaction conditions. At the same time, a catalyst with a central metal content of 1% Ru without ligand is prepared under the same conditions as a control to determine the improvement rate after modification.
[0062] Step 7: Experimental testing
[0063] Experimental tests were conducted on the catalyst prepared with the ligand (CAS No.: 173035-10-4) with a predicted conversion rate of 74%, and the conclusion was drawn that the catalyst conversion rate was 81% under the optimal conditions after adding the ligand, with a difference of 7%, and the effect was good.
[0064] Experimental tests were conducted on the catalyst prepared with the ligand (CAS No.: 4368-51-8) with a predicted conversion rate of 80%, and the conclusion was drawn that the catalyst conversion rate was 83% under the optimal conditions after adding the ligand, with a difference of 3%, and the effect was good.
[0065] The present invention is not limited to the specific technical solutions described in the above embodiments, and all technical solutions formed by equivalent replacement are within the protection scope required by the present invention.
[0066] The present invention is not limited to the specific technical solutions described in the above embodiments, and all technical solutions formed by equivalent replacement are within the protection scope required by the present invention.
Claims
1. A method for predicting the quality of acetylene hydrochlorination catalyst ligands based on support vector machine algorithm, characterized in that: The method comprises the following steps: (1) Extracting data sets from public literature and classifying them to train the model, (2) Extract the composite SMILES string of the ligand molecule and convert it into a tensor matrix. (3) Input the tensor matrix into the model based on the support vector machine algorithm for regression training, using the RBF kernel function and a penalty coefficient C equal to 0.9; (4) The SMILES string of the unknown performance ligand is input into the model to obtain the prediction of the ligand modification results; The SMILES string constructs a tensor matrix and inputs it into a support vector machine model, and uses the trained support vector machine model to output a prediction result; The method comprises the following steps: Step 1: Establish a data set and collect information on known catalyst ligands from the literature. The molecular structure of the ligand and the corresponding composite SMILES string and conversion rate are recorded in the database; Step 2: Convert the molecular structure of the ligand into an ECFP fingerprint The composite SMILES string of each ligand is converted into a 64-bit binary vector called "ECFP fingerprint". For the SMILES string of known ligands, a search method is used to map each functional group into a molecular fingerprint composed of 0 and 1 from the RDKit open source library, and then combine them into a 64-bit molecular fingerprint. Different sequences of 0 and 1 represent different functional groups, thus reflecting the influence of different ligand structures. Step 3: Construct the feature tensor matrix Each ligand collected has its own 64-bit molecular fingerprint. The fingerprint data of all ligands are integrated into a tensor matrix. Each row of the matrix represents a ligand and each column is a feature bit. Step 4: Build and tune the model architecture Input the complete tensor matrix into the support vector machine model for preliminary training; the support vector machine model will automatically calculate and fit; try the linear kernel, polynomial kernel and RBF kernel functions respectively, and use the cross-validation method to optimize the selection of the penalty coefficient C; Step 5: Model Evaluation The simplified tensor matrix is input into the support vector machine model and training continues. The target variable is the conversion rate. Through model training, the model gradually learns the relationship between ligand structure, experimental conditions and conversion rate. After training is completed, the model is evaluated using the validation set; Through model training, the optimal model was finally obtained using the RBF kernel and a penalty coefficient C equal to 0.
9. The model was evaluated using the validation set. The probability that the difference between the predicted conversion rate and the actual conversion rate of different ligands was less than 7% reached more than 95%; Step 6: Predict the conversion rate of unknown ligands After training, the model can be used to predict the conversion rate of new ligands.
2. The method for predicting the quality of acetylene hydrochlorination catalyst ligands based on support vector machine algorithm according to claim 1, characterized in that: Screening out ligands As a modifier for acetylene hydrochlorination catalyst.
3. A computer-readable storage medium, characterized in that: The storage medium has a computer program, and the computer program can run the method for predicting the quality of acetylene hydrochlorination reaction catalyst ligand based on support vector machine algorithm as described in claim 1.
4. A prediction device for the quality of acetylene hydrochlorination catalyst ligands based on support vector machine algorithm, characterized in that: The device is equipped with the computer-readable storage medium according to claim 3.
Citation Information
Patent Citations
Drug chemical reaction type prediction method based on multi-level information fusion
CN115810404A
Screening method of copper-based catalyst for synthesizing vinyl chloride through acetylene hydrochlorination
CN118553330A