A prediction method for the advantages and disadvantages of ligands of acetylene hydrochlorination reaction catalysts based on the random forest algorithm
The catalyst ligand database was constructed through random forest algorithm and ECFP fingerprint technology, which solved the problems of high design cost of mercury-free catalysts and insufficient prediction accuracy, and achieved efficient screening of high-performance catalyst ligands, reducing experimental costs and improving prediction accuracy.
Patent Information
- Application Number
- CN202411670055.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The prior art has high cost and high experimental costs when designing mercury-free catalysts, and traditional models fail to effectively capture the interaction between catalyst ligands and experimental environment variables, resulting in insufficient prediction accuracy.
The random forest algorithm is used to combine ECFP fingerprint technology to construct a catalyst ligand database, and through feature extraction and feature screening, a prediction model is established, and the comprehensive impact of ligand molecular structure and experimental conditions is trained to predict the conversion rate of the catalyst.
It improves the prediction accuracy of the advantages and disadvantages of catalyst ligands, reduces experimental costs, can show excellent prediction effects under sparse data sets, and screens out high-performance mercury-free catalysts.
Smart Images

Figure CN119541733B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the advantages and disadvantages of ligands of a catalyst for acetylene hydrochlorination reaction based on random forest, and in particular, an artificial intelligence-assisted screening for the advantages and disadvantages of ligands of a catalyst for acetylene hydrochlorination reaction is designed and implemented by combining knowledge of catalysis and random forest, belonging to the technical field of materials informatics. Background Art
[0002] Vinyl chloride is an important chemical raw material, which can be polymerized into polyvinyl chloride (PVC). Polyvinyl chloride is widely used in plastic products and is one of the most widely used plastics in the world. As a monomer of polyvinyl chloride, it is particularly important to synthesize vinyl chloride efficiently and in large quantities. Currently, the most commonly used method is the acetylene hydrochlorination reaction to prepare vinyl chloride, and the most commonly used catalyst contains mercury (Hg). As a heavy metal element, Hg has great harm to the human body and the environment. Therefore, finding a catalyst that is equally efficient and clean and pollution-free has become the main issue at present. At present, the main research directions of mercury-free catalysts are ruthenium-based, platinum-based, gold-based, and copper-based. Based on these four types of metals as substrates, designing a class of efficient hydrochlorination catalysts is the main research direction at present.
[0003] The design of catalysts usually requires a large number of experiments, and the costs of manpower and material resources are relatively high. With the rapid development of artificial intelligence, artificial intelligence has been widely used in various fields. Using artificial intelligence can greatly reduce the time cost and reduce unnecessary trial-and-error processes, which is an efficient and fast tool. Currently, it is mainly through training artificial intelligence models to see the advantages of artificial intelligence in terms of time and cost, and it is decided to design and predict mercury-free catalysts with the assistance of artificial intelligence.
[0004] Using artificial intelligence learning for catalyst design is the main research direction, and the mainly used algorithm is the random forest algorithm. The advantages of the random forest algorithm are mainly its high compatibility with errors and better prediction effects than linear regression in some fields. Factors such as reaction conditions are included in the scope of consideration for model training. The random forest model is mainly used to predict the influence of ligands of mercury-free catalysts on the performance of catalytic reactions. The present invention shows excellent prediction effects and prediction values under sparse dataset training and out-of-sample prediction. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to propose a prediction method for the advantages and disadvantages of ligands of acetylene hydrochlorination reaction catalysts based on the random forest algorithm. An integrated learning algorithm is used to establish a model from existing papers and experimental data to obtain a prediction model, so as to predict the conversion rate of catalysts with unknown ligands added, and guide the synthesis of catalysts. The method first constructs a ligand database by chemical ligands obtained from reading various relevant literatures. Then, a novel prediction model for the conversion rate of catalysts with unknown ligands added to acetylene hydrochlorination is proposed, which combines a feature extractor and the random forest algorithm. The model can calculate the ECFP fingerprint according to the SMILES string of the relevant ligand, and can make a specialized prediction of the conversion rate in combination with a specific experimental environment, thereby improving the accuracy of the prediction.
[0006] To solve the technical problems of the present invention, the present invention is implemented by the following technical solutions: A prediction method for the advantages and disadvantages of ligands of acetylene hydrochlorination reaction catalysts based on the random forest algorithm, the method comprising the following steps:
[0007] (1) Collect and extract the data in the literature as a data set for training the model; catalyst ligands, reaction conditions and ligand properties are used to train the model; the reaction conditions are temperature and HCl ratio;
[0008] (2) Extract the composite SMILES string of the ligand molecule and construct a tensor matrix, that is, extract the 64-bit ECFP fingerprint of the ligand, separate it bit by bit and then combine it to construct a tensor matrix; convert the experimental conditions of temperature and HCl ratio into vectors and add them to the feature matrix; the complete feature matrix contains not only the molecular structure information of the ligand, but also different experimental conditions;
[0009] (3) Input the complete feature matrix into the model based on the random forest algorithm for scoring training; each column in the tensor matrix is a feature, set the importance attribute of each feature to the model performance, obtain the importance parameters of each feature by initially training the random forest model using the database data, arrange them in descending order, and remove the features with lower importance in proportion to reduce the dimension and prevent overfitting;
[0010] (4) Input the simplified feature matrix into the random forest for scoring training of the model and input the SMILES string of the ligand with unknown performance to obtain the result of the advantages and disadvantages of the ligand.
[0011] Preferably, the catalyst ligands, reaction conditions and ligand properties are used to train the model: Each atom and chemical bond in the SMILES string is a word. By collecting many compounds, a corpus is naturally formed. The ECFP (Extended Connectivity Fingerprint) technology is used to collect information, and combined with the reaction conditions, the scoring weights are trained.
[0012] Preferably, the ligand complex SMILES string and reaction condition embedding model are used to construct a tensor matrix. The trained LDT model is used to embed the SMILES string of the ligand into the tensor and output a score.
[0013] Preferably, the method includes the following steps:
[0014] Step 1: Establish a dataset. Collect information on known catalyst ligands from the literature. Record the molecular structure of the ligand, the corresponding complex SMILES string, temperature, HCl ratio, and conversion rate in the database.
[0015] Step 2: Convert the molecular structure of the ligand into ECFP fingerprints
[0016] Convert the complex SMILES string of each ligand into a 64-bit binary vector called "ECFP fingerprint". For the SMILES strings of known ligands, search in the RDKit open-source library to map each functional group to a molecular fingerprint composed of 0s and 1s, and combine them into a 64-bit molecular fingerprint. Different sequences of 0s and 1s represent different functional groups, thus reflecting the influence of different ligand structures.
[0017] Step 3: Construct a feature tensor matrix
[0018] Each ligand collected has its own 64-bit molecular fingerprint. Integrate the fingerprint data of all ligands into a tensor matrix. Each row of the matrix represents a ligand, and each column is a feature bit.
[0019] Convert the experimental conditions of temperature and HCl ratio into vectors and add them to the feature matrix.
[0020] Step 4: Establish and adjust the model architecture
[0021] Input the complete feature matrix into the random forest model for preliminary training. The random forest model will automatically calculate the importance of each feature. These scores represent the contribution of each feature to the prediction.
[0022] By adjusting the number of trees, optimize and adjust by trying the number of trees (5, 10, 15, 20), the minimum number of samples for leaf merging (1, 2, 3), and the minimum gain for splitting (0, 0.05, 0.1). After model training, it is found that the contribution of some features is extremely low. Among the 64-bit fingerprints, 7 bits have a low contribution. Remove these unimportant features and only retain the features with greater influence. This feature screening can effectively reduce noise and prevent model overfitting.
[0023] Step 5: Model evaluation
[0024] Input the simplified feature matrix into the random forest model for further training. The target variable is the conversion rate. Through model training, the model gradually learns the relationship between ligand structure, experimental conditions, and conversion rate. After training, use the validation set to evaluate the model. Suppose the predicted conversion rate of ligand A is 78%, while the actual value is 80%, and the error is 2%.
[0025] Through model training, the determined parameters are as follows: number of trees: 10, minimum number of samples for leaf merging: 1, minimum gain for splitting: 0.05. After training, use the validation set to evaluate the model. The probability that the difference between the predicted conversion rate and the actual conversion rate of different ligands by the final model is less than 5% reaches more than 50%.
[0026] Step 6: Predict the conversion rate of unknown ligands
[0027] After training, this model can be used to predict the conversion rate of new ligands, ligand X. Generate the 64-bit fingerprint of ligand X using RDKit, and then merge it with the corresponding experimental conditions into the feature matrix. After inputting the feature vectors of the temperature and HCl ratio of the experimental conditions of X into the model, the predicted conversion rate improvement rate output by the model. This output result can help determine whether ligand X has good catalytic effects under these reaction conditions.
[0028] The prediction method for the quality of acetylene hydrochlorination reaction catalyst ligands based on the random forest algorithm screens out ligands as modifiers for acetylene hydrochlorination reaction catalysts.
[0029] The storage medium has a computer program that can run the prediction method for the quality of acetylene hydrochlorination reaction catalyst ligands based on the random forest algorithm as claimed in claim 1.
[0030] A prediction device for the quality of acetylene hydrochlorination reaction catalyst ligands based on the random forest algorithm, where the device is equipped with the computer-readable storage medium as claimed in claim 7.
[0031] Method for catalyzing the addition of an unknown ligand and predicting the conversion rate of acetylene hydrochlorination, the specific steps of the method are as follows: (1) Construct a chemical ligand database to train the model; (2) Extract the 64-bit ECFP fingerprint of the ligand, separate it bit by bit and then combine it to construct a tensor matrix; Convert the experimental conditions of temperature and HCl ratio into vectors and add them to the feature matrix; The complete feature matrix not only contains the molecular structure information of the ligand, but also contains different experimental conditions; (3) Input the complete feature matrix as the initial feature vector into the random forest, each column in the matrix is a feature, set the importance attribute of each feature to the model performance, obtain the importance parameters of each feature by initially training the random forest model using the database data, sort them in descending order, and remove the features with lower importance according to a ratio to reduce the dimension and prevent overfitting; (4) Input the simplified feature matrix into the random forest for model training and evaluate its performance.
[0032] First, by consulting various relevant literatures, a large amount of chemical ligand data is collected. Each ligand needs to include its molecular structure, properties, and its conversion rate in the acetylene hydrochlorination reaction and other relevant information. The information of the ligand is obtained through literature investigation, laboratory data collection, and the public chemical database PubChem.
[0033] Furthermore, the chemical ligand is characterized by the ECFP (Extended Connectivity Fingerprint) technology. ECFP can effectively represent the topological information of the molecule and convert the chemical structure of the ligand into a fixed-length vector representation. Use the chemical informatics tool RDKit to extract the 64-bit ECFP fingerprint of each ligand. The neighborhood information of atoms or groups in the molecule is extracted through an iterative algorithm to construct a series of binary bit representations, thereby generating a fixed-length fingerprint. Convert the extracted 64-bit ECFP fingerprint into a tensor matrix. The ECFP fingerprint of each ligand is a 64-dimensional binary vector. When constructing the ECFP tensor, the fingerprint of each ligand will become a row vector of a matrix. When there are multiple ligands, these row vectors will form a two-dimensional tensor, where each row represents the fingerprint of a ligand.
[0034] Furthermore, use the tensor matrix obtained in the previous step as the feature input and the conversion rate as the target output to train a random forest model. As the training progresses, the random forest model will assign an importance score to each input feature, and this score reflects the contribution degree of each feature to the model performance. Through these scores, it can be identified which features contribute more to predicting the conversion rate and which features contribute less to the model. Sort the feature importance and remove the features with lower importance according to a ratio, which will help reduce the negative impact of irrelevant or redundant features on the model, thereby reducing the complexity of the model and the risk of overfitting.
[0035] Furthermore, select the temperature and HCL ratio that are closely related to the acetylene hydrochlorination reaction. These variables will affect the reaction performance of the catalyst, so it is necessary to consider these factors in the model. Convert the environmental variables into a tensor matrix, and use the value of each item as a column in the matrix to construct a two-dimensional tensor of environmental variables. Combine the tensor matrix of environmental variables and the dimension-reduced ECFP tensor matrix column by column to obtain a complete feature matrix.
[0036] Furthermore, use the complete feature matrix obtained in the previous step as the input, and the conversion rate as the target variable to input into the random forest model for training, and further adjust the training parameters.
[0037] Beneficial effects:
[0038] The present invention uses the random forest algorithm to train a model. The advantage of the random forest is mainly excellent in dealing with noise points, which can reduce the influence of errors caused by some experiments. This type of model can greatly reduce the workload, and in the future development, with the increase of the data set, the prediction accuracy will still improve. We have designed and predicted multiple high-performance mercury-free hydrochlorination catalysts using this algorithm, and they show high accuracy and catalytic performance.
[0039] Existing methods for predicting the conversion rate of catalysts usually focus on data features from a single source, such as modeling only based on the molecular structure information of ligands or a single experimental environmental variable. Such models may ignore the interaction between chemical ligands and experimental environmental variables and their comprehensive impact on the catalytic reaction performance. The present invention adopts a method of combining the ECFP fingerprint of ligands with experimental environmental variables. In this way, the model not only considers the molecular structure of ligands during the catalyst preparation process, but also can make dynamic predictions according to changes in experimental conditions, thus more comprehensively capturing the influencing factors of catalyst performance. This multi-feature fusion greatly improves the prediction accuracy of the model, especially when facing unknown ligands and changing experimental conditions.
[0040] Many traditional methods use simple molecular descriptors or features to represent ligands. These methods usually fail to effectively capture the complexity of ligand structures, resulting in poor model performance. The present invention uses the ECFP fingerprint to extract features of chemical ligands and converts them into 64-bit binary vectors. By converting the ECFP fingerprint of each ligand into a tensor matrix and inputting it into the model, the present invention can efficiently express the complex information of the molecular structure. This tensor matrix form can capture higher-order correlations in the molecule, thereby improving the model's representation ability and prediction accuracy for ligands. Description of the drawings
[0041] The present invention will be further described below in conjunction with the accompanying drawings.
[0042] Figure 1 is a schematic diagram of the method flow of the present invention Specific embodiments
[0043] The technical solution of the present invention will be further explained below through examples in conjunction with the accompanying drawings.
[0044] Example 1
[0045] Step 1: Establish a data set
[0046] First, collect information on known catalyst ligands. For example, a catalyst prepared with a ligand A was found to have an acetylene hydrochlorination reaction at 230 °C with an HCl ratio of 1.15 and a conversion rate of 80%. Record this information (molecular structure, SMILES string, experimental conditions, conversion rate) in the database. Considering the influence of different batches of activated carbon, we set the target value as the difference in conversion rates between adding the ligand and not adding the ligand, and calculate all the target values.
[0047] Step 2: Convert the molecular structure of the ligand into an ECFP fingerprint
[0048] To enable the model to "understand" the structure of the ligand, first obtain the SMILES string of the ligand from the website https: / / pubchem.ncbi.nlm.nih.gov / . Further, encode the SMILES string of each ligand into a 64-bit binary vector. Use the chemical tool RDKit to complete this step and generate the ECFP fingerprint of each ligand. For example, the SMILES string of ligand A is C=CCl, and RDKit converts it into the following 64-bit binary fingerprint: [1, 0, 0, 1, 1, 0, 1,...]. Each bit in this string of data represents a detailed information of the molecular structure, such as whether it contains certain specific chemical bonds, groups, etc. For subsequent model training, this fingerprint becomes the "identity identifier" of the ligand.
[0049] Step 3: Construct a feature tensor matrix
[0050] Suppose 100 ligands are collected, and each ligand has its own 64-bit fingerprint. Integrate the fingerprint data of all ligands into a tensor matrix. Each row of the matrix represents a ligand, and each column is a feature bit. For example, the first two rows can be like this:
[0051] | Ligand number | Bit 1 | Bit 2 | Bit 3 | Bit 4 | Bit 5 |... |
[0052] |----------|-----|-----|-----|-----|-----|-----|
[0053] | A | 1 | 0 | 0 | 1 | 1 | ... |
[0054] | B | 0 | 1 | 1 | 0 | 1 | ... |
[0055] Under this matrix structure, the molecular structure characteristics of all ligands are clear at a glance, which is convenient for subsequent input into the model.
[0056] Step 4: Add experimental conditions
[0057] The conversion rate of the acetylene hydrochlorination reaction is not only affected by the ligand structure, but also closely related to experimental conditions (such as temperature, HCl ratio). To improve the prediction accuracy, experimental conditions such as temperature and HCl ratio are converted into vectors and added to the feature matrix. Suppose the experimental temperature of ligand A is 230 °C and the HCl ratio is 1.15. These values can be encoded as additional feature columns, for example:
[0058] | Ligand number | Bit 1 | Bit 2 | ... | Temperature | HCl ratio |
[0059] |----------|-----|-----|-----|------|---------|
[0060] | A | 1 | 0 | ... | 230 | 1.15 |
[0061] | B | 0 | 1 | ... | 250 | 1.05 |
[0062] In this way, the feature matrix not only contains the molecular structure information of the ligands, but also contains different experimental conditions, and the model can take into account the reaction performance differences brought about by environmental changes accordingly.
[0063] Step 5: Feature screening
[0064] Input the complete feature matrix into the random forest model for preliminary training. The random forest model will automatically calculate the importance of each feature, and these scores represent the contribution of each feature to the prediction.
[0065] By adjusting the number of trees, optimize and adjust by trying the number of trees (5, 10, 15, 20), the minimum number of samples for leaf merging (1, 2, 3) and the minimum gain for splitting (0, 0.05, 0.1) respectively.
[0066] After model training, it is found that the contributions of certain features are extremely low. Among the 64-bit fingerprints, 7 bits have relatively low contributions. Remove these unimportant features and only retain the features with greater influence. This feature screening can effectively reduce noise and prevent model overfitting.
[0067] Step 6: Model Training and Evaluation
[0068] Input the simplified feature matrix into the random forest model for continued training. The target variable is the conversion rate. Through model training, the model gradually learns the relationship between ligand structure, experimental conditions, and conversion rate. After training is completed, use the validation set to evaluate the model. Suppose the predicted conversion rate of ligand A is 78%, while the actual value is 80%, with an error of 2%. Through this evaluation, the model parameters can be adjusted. After attempts, the optimal parameters are determined as the number of trees: 10, the minimum number of samples for leaf merging: 1, and the minimum splitting gain: 0.05 to ensure the prediction accuracy of the model for different ligands.
[0069] Step 7: Predict the Conversion Rate of Unknown Ligands
[0070] After training is completed, the model can be used to predict the conversion rate of new ligands. Suppose there is a new ligand X with its SMILES string being C#CCl. Generate the 64-bit fingerprint of ligand X using RDKit and then merge it with the experimental conditions into the feature matrix. For example, assume the experimental conditions of X are a temperature of 320 °C and an HCl ratio of 1.10. After inputting its feature vector into the model, the predicted conversion rate increase output by the model is 20%. This result can help determine whether ligand X has good catalytic effects under these reaction conditions and whether it is worth further verification in experiments.
[0071] Over 8000 ligands were collected from https: / / www.chemsrc.com / and converted into SMLILES strings. The HCl ratio was set to 1.15 and the reaction temperature was set to 180 degrees Celsius, and they were input into the model for prediction. Approximately 100 ligands were found to have excellent prediction effects. After removing factors such as toxicity, danger, and price, 2 ligands were selected from the remaining ligands for verification.
[0072] Step 8: Preparation of Catalysts
[0073] Use the impregnation method to prepare a catalyst with a central metal content of 1% Ru from the ligands with good predicted effects under the reaction conditions given in Step 7. Use the same reaction conditions to experimentally test the performance of the catalyst. At the same time, prepare a catalyst with a central metal content of 1% Ru without adding ligands under the same conditions as a control to determine the improvement rate after modification.
[0074] Step 9: Experimental Testing
[0075] Experimental tests were carried out on the catalyst prepared from the ligand (CAS No.: 3109-63-5) with a predicted increase in the conversion rate of 16% and a blank control. The conclusion was that the conversion rate of the catalyst with the ligand added was 88%, that without the ligand added was 63%, and the increase was 25%, showing good results.
[0076] Experimental tests were carried out on the catalyst prepared from the ligand (CAS No.: 65039-09-0) with a predicted increase in the conversion rate of 40% and a blank control. The conclusion was that the conversion rate of the catalyst with the ligand added was 95%, that without the ligand added was 63%, and the increase was 32%, showing good results.
[0077] The present invention is not limited to the specific technical solutions described in the above embodiments. Any technical solution formed by equivalent replacement is within the scope of protection required by the present invention.
Claims
1. A prediction method for the advantages and disadvantages of catalysts ligands for acetylene hydrochlorination reaction based on the random forest algorithm, characterized in that, The method includes the following steps: Step 1: Establish a data set. The catalyst is a catalyst with a central metal content of 1% Ru. Information on known catalyst ligands, the molecular structure of the ligands and their corresponding composite SMILES strings, the temperature of the reaction conditions, the HCl ratio of the reaction conditions, and the conversion rate are recorded in the database; Step 2: Convert the molecular structure of the ligand into an ECFP fingerprint Convert the composite SMILES string of each ligand into a 64-bit binary vector called "ECFP fingerprint"; for the SMILES strings of known ligands, search the RDKit open-source library to convert each functional group into a molecular fingerprint composed of 0s and 1s, and combine them into a 64-bit molecular fingerprint; different compositions of 0s and 1s represent different functional groups, thus reflecting the influence of different ligand structures; Step 3: Construct a feature tensor matrix Each ligand collected has its own 64-bit molecular fingerprint. Integrate the fingerprint data of all ligands into a tensor matrix. Each row of the matrix represents a ligand, and each column is a feature bit; Convert the reaction condition temperature and HCl ratio into vectors and add them to the feature matrix; Step 4: Establish and adjust the model architecture Input the complete feature matrix into the random forest model for preliminary training. The random forest model will automatically calculate the importance scores of each feature. These scores represent the contribution of each feature to the prediction. By adjusting the number of trees, optimize and adjust by trying the number of trees 5, 10, 15, 20, the minimum number of samples for leaf merging 1, 2, 3, and the minimum gain for splitting 0, 0.05, 0.
1. After model training, it is found that the contribution of some features is extremely low. Among the 64-bit fingerprints, 7 bits have a low contribution. Remove these unimportant features and only retain the features with greater influence. This feature screening can effectively reduce noise and prevent model overfitting; Step 5: Model evaluation Input the simplified feature matrix into the random forest model for continued training. The target variable is the conversion rate; through model training, the model gradually learns the relationship between ligand structure, reaction conditions and conversion rate; After training is completed, use the validation set to evaluate the model; Through model training, determine the parameters as follows: number of trees: 10, minimum number of samples for leaf merging: 1, minimum gain for splitting: 0.
05. After training is completed, use the validation set to evaluate the model; Step 6: Predict the conversion rate of unknown ligands After training is completed, this model can be used to predict the conversion rate of a new ligand X; generate the 64-bit fingerprint of ligand X using RDKit, convert the reaction condition temperature and HCl ratio into vectors and merge them into the feature matrix. After inputting into the model, the model outputs the predicted conversion rate.
2. The prediction method for the advantages and disadvantages of the ligands of the acetylene hydrochlorination reaction catalyst based on the random forest algorithm according to claim 1, wherein: Use the impregnation method to prepare a catalyst with a central metal content of 1% Ru and select the ligand , As a modifier for the acetylene hydrochlorination reaction catalyst.