Amine anti-aging agent performance prediction method based on machine learning and molecular simulation
By combining machine learning and molecular simulation, a model for p-phenylenediamine antioxidants was constructed, which solved the problems of high cost, long cycle and easy error of traditional experimental methods, and realized efficient and low cost performance prediction and design of antioxidants.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PETROLEUM & CHEMICAL CORP
- Filing Date
- 2026-01-14
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional experimental methods for screening high-performance p-phenylenediamine antioxidants are costly, time-consuming, cumbersome, and prone to errors, making it difficult to efficiently screen high-performance antioxidants.
A model of p-phenylenediamine antioxidants was constructed by combining machine learning and molecular simulation. The relationship between molecular structure and performance was analyzed through machine learning, a database was established and the model was optimized to predict the performance of antioxidants.
It reduces the time and money costs of traditional experiments, improves the efficiency of antioxidant performance prediction, reduces errors, and guides the design of efficient molecular structures.
Smart Images

Figure CN121963950A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of performance evaluation technology for p-phenylenediamine antioxidants, and in particular to a method for predicting the performance of amine antioxidants based on machine learning and molecular simulation. Background Technology
[0002] Rubber, plastics, lubricants, and other materials play vital roles in daily life and industrial production. However, these materials readily react with oxygen in the air, leading to increased oxidation over time and ultimately rendering them unusable. Antioxidants, commonly used additives in rubber and plastic processing, effectively inhibit oxidation. Among them, p-phenylenediamine antioxidants exhibit excellent anti-aging properties at medium and high temperatures, and their production technology and cost control are more mature, making them widely used in various industrial materials.
[0003] Current research on designing high-performance p-phenylenediamine antioxidants mainly focuses on experimental exploration. Researchers largely employ traditional trial-and-error methods to test and evaluate the anti-aging performance of these antioxidants. Specifically, this can be summarized as follows: based on the hypothetical molecular structure of the p-phenylenediamine antioxidant, a synthetic experimental apparatus is built; reactants and catalysts are added; reaction conditions such as temperature and pressure are adjusted; and the product is separated and purified. The synthesized product is analyzed to determine the target product structure. Then, rubber experiments are conducted, from raw rubber to compound rubber to vulcanized rubber. Rubber samples are prepared according to the experimental requirements, and macroscopic properties such as tensile properties, aging life, abrasion resistance, and dynamic properties are tested.
[0004] However, when using this experimental method to study the anti-aging properties of dozens of p-phenylenediamine antioxidants, the above process must be repeated to prepare the antioxidants and the rubber materials again, indirectly resulting in high financial and time costs. Therefore, the method of screening high-performance p-phenylenediamine antioxidants through experiments is highly unreliable, cumbersome, time-consuming, and labor-intensive. Furthermore, improper operation may lead to experimental errors, making the screening of high-performance p-phenylenediamine antioxidants difficult. Summary of the Invention
[0005] This application provides a method for predicting the performance of amine antioxidants based on machine learning and molecular simulation. It uses molecular simulation modeling to change the type of p-phenylenediamine antioxidant, and uses machine learning methods to analyze the quantitative relationship between the molecular structure of p-phenylenediamine antioxidants and their anti-aging performance. This reduces the time cost and error of traditional trial-and-error methods for synthesizing p-phenylenediamine antioxidants, guides the design of high-performance p-phenylenediamine antioxidant molecular structures, and fills the gap in predicting the performance of p-phenylenediamine antioxidants using machine learning and molecular simulation.
[0006] The above-mentioned objective of this application is achieved through the following technical solution: a method for predicting the performance of amine antioxidants based on machine learning and molecular simulation, comprising the following steps:
[0007] Multiple models of p-phenylenediamine antioxidants were constructed and optimized.
[0008] Calculate the anti-aging performance parameters of each p-phenylenediamine antioxidant model structure;
[0009] Build a machine learning model and train the machine learning model using the optimized structural model and performance parameters as training data;
[0010] The performance of p-phenylenediamine antioxidants was predicted using a trained machine learning model, and the structure-activity relationship between the structure and properties of p-phenylenediamine antioxidants was obtained.
[0011] Furthermore, the p-phenylenediamine antioxidant model includes a 6PPD structure.
[0012] Furthermore, the constructed p-phenylenediamine antioxidant model includes the initial antioxidant 6PPD model. Other models are based on antioxidant 6PPD, with substituents R introduced at different positions in the two benzene rings.
[0013] Furthermore, in the step of constructing the model of p-phenylenediamine antioxidants, a topological index is introduced to distinguish isomers, and the valence linkage index of the antioxidant molecule is calculated using the following formula:
[0014]
[0015] in, Let be the m-th order valence connectivity index, where m is the order of the valence connectivity index of the antioxidant molecule, p is the path type of the valence connectivity index, and n is the number of m-th order p-type subgraphs in the molecule. This represents the valence connectivity of each atom in the molecule.
[0016] Furthermore, in the optimization step of the constructed model, the structure of the constructed model is kept in a state of minimum energy.
[0017] Furthermore, in the optimization step of the constructed model, the density functional method is used to optimize the constructed model.
[0018] Furthermore, it also includes the following steps:
[0019] A database was constructed using the optimized p-phenylenediamine antioxidant model and performance parameters. Based on the molecular structure characteristics of the p-phenylenediamine antioxidant model, its structure was decomposed into functional group fragments and statistically analyzed, transforming the chemical structure into a feature matrix.
[0020] Furthermore, after obtaining the anti-aging performance parameters, the data undergoes dimensionality reduction optimization. This optimization is performed using a maximum-minimum normalization method, with the following formula:
[0021] ;
[0022] Where x' is the normalized data, and x is the original data of the molecular structure characteristic parameters and anti-aging performance parameters of p-phenylenediamine antioxidants. min and x max These are the minimum and maximum values of the corresponding features in the original data, respectively.
[0023] Furthermore, the machine learning models mentioned include neural network models and random forest models.
[0024] Furthermore, a fully connected BP network structure was selected for the neural network model, with 3-5 layers and 3-7 hidden neurons. The sigmoid function was chosen as the neuron activation function, the epochs were set to 1000, and the mean squared error (MSE) loss function was selected as shown below:
[0025]
[0026] Where n is the number of samples, These are the predictions from the neural network. This is the target value of the neural network, and the Levenberg-Marquardt algorithm is used for optimization.
[0027] The random forest model uses a classification and regression tree algorithm structure that generates only binary decision trees. The number of decision trees is selected between 1 and 200, the minimum number of samples per leaf node is selected between 1 and 10, and the minimum number of samples per leaf node split is selected between 1 and 15. The random forest model is optimized by performing a gridded search on the random forest model under cross-validation, considering the number of decision trees, the minimum number of samples per leaf node, and the minimum number of samples per leaf node split. Based on the gridded search results, the optimal parameters of the random forest model are determined.
[0028] In summary, the beneficial effects of this application are as follows:
[0029] 1. By combining machine learning and molecular simulation, the traditional trial-and-error method for exploring the structure of high-performance p-phenylenediamine antioxidants is transformed into a simulation process, which saves experimental costs, shortens the design cycle, and improves design efficiency.
[0030] 2. The structure-activity relationship of p-phenylenediamine antioxidants was analyzed by combining machine learning and molecular simulation, which filled the gap in the field of machine learning-guided efficient molecular structure design of p-phenylenediamine antioxidants and improved the testing efficiency of the anti-aging performance of p-phenylenediamine antioxidants.
[0031] 3. When establishing the database, introduce a valence connectivity index to distinguish isomers of p-phenylenediamine antioxidants, so as to avoid antioxidant performance prediction errors caused by duplicate input data.
[0032] 4. Data dimensionality reduction and optimization processes eliminate the influence of different dimensions between parameters, thereby improving the convergence speed and optimization efficiency of machine learning models. Attached Figure Description
[0033] Figure 1 This is the structural model of the antioxidant 6PPD in this invention.
[0034] Figure 2 This is a graph showing the relationship between the predicted and actual values of the neural network model in this invention.
[0035] Figure 3 This invention presents the effects of different functional groups and structural parameters in the molecular structure of p-phenylenediamine antioxidants on their anti-aging performance, as predicted by the random forest model.
[0036] Figure 4 This is a flowchart illustrating the method for predicting the performance of amine antioxidants based on machine learning and molecular simulation in this invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this invention more prominent and clear, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The specific embodiments in this detailed description are only used to explain the content of this invention and are not intended to limit the scope of application or promotion of this invention.
[0038] This invention provides a method for predicting the performance of amine antioxidants based on machine learning and molecular simulation. It predicts the anti-aging performance parameters of p-phenylenediamine antioxidants based on the molecular structural characteristics of these antioxidants. This transforms the traditional trial-and-error method for synthesizing antioxidants and the rubber testing process into a machine learning and simulation process, significantly reducing time costs and improving the efficiency of performance prediction and molecular structure design for p-phenylenediamine antioxidants. Furthermore, this method fills a gap in the field of machine learning-guided molecular structure design for amine antioxidants.
[0039] This invention discloses a method for predicting the performance of amine antioxidants based on machine learning and molecular simulation, which specifically includes the following steps:
[0040] S1, Constructing multiple p-phenylenediamine antioxidant models: Models were constructed using Materials Studio software, such as... Figure 1 The initial model of antioxidant 6PPD shown was optimized using the Forcite module to keep the molecular structure of antioxidant 6PPD in a state of minimum energy.
[0041] S2, Simulation Calculation: In this embodiment, the DMol3 module of Materials Studio software is used to further optimize the geometry of the antioxidant 6PPD molecule in the lowest energy state under three functional functions: LDA-PWC, GGA-BP, and B3LYP.
[0042] Furthermore, based on the geometric optimization results, the energy parameters and correction parameters related to the anti-aging performance of the antioxidant were extracted from the results. The anti-aging performance parameters of the antioxidant 6PPD under three different functional functions were calculated and compared with the parameters in the literature experiments. The GGA-BP functional was selected as the optimal one.
[0043] Other p-phenylenediamine antioxidants use antioxidant 6PPD as a matrix, introducing substituents R at different positions on the two benzene rings. Preferred R values are C1-C8 alkyl groups, alkyl groups containing ether bonds, etc., thus forming the basis for subsequent structure-activity relationship studies of p-phenylenediamine antioxidants. The molecules of these other p-phenylenediamine antioxidants also undergo step S1 to maintain their molecular structure in the lowest energy state, followed by DMol3 optimization in step S2, calculating the anti-aging performance parameters for each p-phenylenediamine antioxidant.
[0044] S3, Database Establishment: In step S2, the method constructed various different molecular structures of p-phenylenediamine antioxidants. Based on the characteristics of these antioxidant molecular structures, their structures were decomposed into functional group fragments and statistically analyzed, transforming the chemical structure into a feature matrix. Combined with the anti-aging performance parameters of the p-phenylenediamine antioxidants calculated in step S2, each molecular structure corresponds to one anti-aging performance parameter, thus forming a data set. All data sets are combined to form a machine learning database.
[0045] Furthermore, during the construction of the molecular structure of p-phenylenediamine antioxidants, new parameters need to be introduced to avoid errors introduced by isomers into the machine learning model. This embodiment preferably uses a topological index to distinguish isomers. The valence linkage index of the antioxidant molecule is calculated using the following formula:
[0046]
[0047] in, Let be the m-th order valence connectivity index, where m is the order of the valence connectivity index of the antioxidant molecule, p is the path type of the valence connectivity index, and n is the number of m-th order p-type subgraphs in the molecule. This represents the valence connectivity of each atom in the molecule. As the formula shows, the higher the order of the topological index, the more comprehensive the information about the antioxidant molecule's structure it covers. In this embodiment, the preferred value for m is 2 to 6.
[0048] The molecular structure of p-phenylenediamine antioxidants was decomposed using the method described above to obtain complete structural characteristic parameters. Combined with the anti-aging performance parameters of the antioxidants obtained from the simulation calculations in S2, the following maximum-minimum normalization method was used to perform dimensionality reduction and optimization on the data, constructing a database for subsequent machine learning.
[0049]
[0050] Where x' is the normalized data, and x is the original data of the molecular structure characteristic parameters and anti-aging performance parameters of p-phenylenediamine antioxidants. min and x max These are the minimum and maximum values of the corresponding features in the original data, respectively.
[0051] S4. Based on the features of the constructed database, machine learning models, including neural network models and random forest models, are built for subsequent structure-activity relationship analysis between the molecular structure and properties of p-phenylenediamine antioxidants.
[0052] The neural network model uses a fully connected BP network structure, with 3-5 layers to avoid overfitting. Based on the features in the S3 database, the number of neurons in the input and output layers is determined, and the number of neurons in the hidden layers is preferably between 3 and 7, based on empirical formulas. The sigmoid function is chosen as the neuron activation function, the epochs are set to 1000, and the mean squared error (MSE) loss function is selected as shown below:
[0053]
[0054] Where n is the number of samples, These are the predictions from the neural network. It is the target value of the neural network, and the optimization algorithm adopts the Levenberg-Marquardt algorithm.
[0055] Random forest uses a classification and regression tree algorithm structure that generates only binary decision trees. Based on the database features in S3, the number of neurons in the input and output layers is determined. The number of decision trees is preferably between 1 and 200. The maximum depth of the decision trees is not limited. The minimum number of samples for leaf nodes is preferably between 1 and 10. The minimum number of samples for leaf node splits is preferably between 1 and 15.
[0056] S5, after completing the database setup in S3 and the machine learning model framework in S4, optimizes the neural network model and random forest model to determine the optimal parameters for the two models to fit the database. First, the database is divided into two sub-databases, A1 and A2, according to two ratios: 7:3 for the training set and 8:2 for the validation set.
[0057] In this embodiment of the neural network model, the A1 database is fixed, and the number of neural network layers and the number of hidden layer neurons are changed individually. By comparing the MSE results of different combinations, the preferred number of neural network layers and the number of hidden layer neurons are selected. Under the preferred parameters, the same optimization process is performed on the A2 database. The mean relative error (ARE) is defined as the final error of the neural network model's prediction result. By comparing the ARE and its range distribution of the two sub-databases A1 and A2, the preferred ratio of training set to validation set is determined.
[0058]
[0059] Similarly, the random forest model described in the embodiments is optimized. Under 5x cross-validation, a gridded search is performed on the random forest model between the number of decision trees (1~200), the minimum number of samples per leaf node (1~10), and the minimum number of samples per leaf node split (1~15). The step size for the number of decision trees is 5, and the step sizes for the minimum number of samples per leaf node and the minimum number of samples per leaf node split are 1. Based on the gridded search results, the optimal parameters of the random forest model are determined.
[0060] S6, In Example S5, the optimal parameters for the neural network model and the random forest model were determined. Based on these optimal parameters, the two machine learning models were modified to obtain the final determined model. Further, the database from S3 was input into the final determined neural network model to obtain the structure-activity relationship between the molecular structure and anti-aging properties of p-phenylenediamine antioxidants. The relationship between the predicted values and actual values of the neural network model is as follows: Figure 2 As shown in the diagram. According to this neural network model, changing the input value (i.e., the structure of p-phenylenediamine antioxidants) will yield the corresponding output value (i.e., the anti-aging performance parameters of p-phenylenediamine antioxidants), thereby enabling the prediction of the performance of p-phenylenediamine antioxidants.
[0061] Furthermore, the database in S3 was input into the finalized random forest model to obtain the influence of different functional groups and structural parameters in the molecular structure of p-phenylenediamine antioxidants on their anti-aging performance. The results are as follows: Figure 3 As shown in the figure. According to the results of this random forest model, introducing groups with large influence factors into the antioxidant 6PPD can effectively improve the antioxidant's anti-aging performance. This model plays an important role in guiding the molecular structure design of high-performance antioxidants.
[0062] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the inventive concept of this application, and these all fall within the protection scope of this application.
Claims
1. A method for predicting the performance of amine antioxidants based on machine learning and molecular simulation, characterized in that, Includes the following steps: Multiple models of p-phenylenediamine antioxidants were constructed and optimized. Calculate the anti-aging performance parameters of each p-phenylenediamine antioxidant model structure; Build a machine learning model and train the machine learning model using the optimized structural model and performance parameters as training data; The performance of p-phenylenediamine antioxidants was predicted using a trained machine learning model, and the structure-activity relationship between the structure and properties of p-phenylenediamine antioxidants was obtained.
2. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 1, characterized in that, The p-phenylenediamine antioxidant model includes a 6PPD structure.
3. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 2, characterized in that, The constructed p-phenylenediamine antioxidant model includes the initial antioxidant 6PPD model. Other models are based on antioxidant 6PPD, with substituents R introduced at different positions in the two benzene rings.
4. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 1, characterized in that, In the process of constructing the model of p-phenylenediamine antioxidants, a topological index is introduced to distinguish isomers. The valence linkage index of the antioxidant molecule is calculated using the following formula: in, Let be the m-th order valence connectivity index, where m is the order of the valence connectivity index of the antioxidant molecule, p is the path type of the valence connectivity index, and n is the number of m-th order p-type subgraphs in the molecule. This represents the valence connectivity of each atom in the molecule.
5. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 1, characterized in that, In the optimization step of the constructed model, the structure of the constructed model is kept in a state of minimum energy.
6. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 1, characterized in that, In the optimization step of the constructed model, the density functional method is used to optimize the constructed model.
7. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 1, characterized in that, After obtaining the anti-aging performance parameters, the data undergoes dimensionality reduction optimization. This optimization is performed using a maximum-minimum normalization method, with the following formula: ; Where x' is the normalized data, and x is the original data of the molecular structure characteristic parameters and anti-aging performance parameters of p-phenylenediamine antioxidants. min and x max These are the minimum and maximum values of the corresponding features in the original data, respectively.
8. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 1, characterized in that, It also includes the following steps: A database was constructed using the optimized p-phenylenediamine antioxidant model and performance parameters. Based on the molecular structure characteristics of the p-phenylenediamine antioxidant model, its structure was decomposed into functional group fragments and statistically analyzed, transforming the chemical structure into a feature matrix.
9. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 8, characterized in that, The machine learning models mentioned include neural network models and random forest models.
10. The method for predicting the performance of amine antioxidants based on machine learning and molecular simulation according to claim 9, characterized in that, The neural network model uses a fully connected BP network structure with 3-5 layers and 3-7 hidden neurons. The sigmoid function is chosen as the neuron activation function, the epochs are set to 1000, and the mean squared error (MSE) loss function is selected as shown below: Where n is the number of samples, These are the predictions from the neural network. This is the target value of the neural network, and the Levenberg-Marquardt algorithm is used for optimization. The random forest model uses a classification and regression tree algorithm structure that generates only binary decision trees. The number of decision trees is selected between 1 and 200, the minimum number of samples per leaf node is selected between 1 and 10, and the minimum number of samples per leaf node split is selected between 1 and 15. The random forest model is optimized by performing a gridded search on the random forest model under cross-validation, considering the number of decision trees, the minimum number of samples per leaf node, and the minimum number of samples per leaf node split. Based on the gridded search results, the optimal parameters of the random forest model are determined.